Review DoAsIDo - Heungwoo/research GitHub Wiki
Paper: "Do as I Do: Dexterous Manipulation Data from Everyday Human Videos" — arXiv 2606.19333 (Jun 17 2026) Authors: Bhawna Paliwal, Haritheja Etukuru, William Liang, Pieter Abbeel, Nur Muhammad Mahi Shafiullah, Jitendra Malik · UC Berkeley / NYU What it is: a device-free pipeline that turns ordinary monocular RGB human videos into robot-executable dexterous-hand trajectories — the cheapest L1→L4 path in the Dexterous-Hand Data Pyramid.
Companions: Dexterous-Hand Data Pyramid · AnyDexRT (the retargeting-bridge counterpart) · Dexterous Manipulation · Human Video → Robot Transfer.
- Everyday video → dexterous robot data, no gloves or teleop. Do As I Do reconstructs 4D hand-object dynamics from a single monocular RGB clip (internet, egocentric, or generated) and dynamics-aware retargets it into actions for a real dexterous hand.
- Reconstruction stack: HaWoR (hand tracking) + SAM 3D (object mesh/pose) + a guided-diffusion object tracker (fix shape at an anchor frame, track pose via flow matching at inference) + centroid/gravity alignment (GeoCalib).
- Dynamics-aware retargeting by sampling-based (MPPI-style) optimization in MuJoCo Warp, robustified for noisy references with warmup steps, random force perturbation, and a transition reward.
- Results: retargeting success 25% → 71% on reconstructed in-the-wild video and 72% → 81% on clean OakInk2 MoCap (vs annealed-sampling alone); object tracking preferred over FoundationPose 67% of the time. Ships 500 human-verified trajectories deployed on a 22-DoF Sharpa Wave hand across 10 real tasks.
- It removes the capture device from the bottom of the pyramid. Where DexUMI needs a sensorized glove and teleop needs a rig, Do As I Do works from any RGB video — the most abundant (L1) data — and carries it through the retargeting bridge (L4) with physics verification (L5). This is the device-free extreme of the human-video fork.
- Dynamics-aware, not kinematic-only. Retargeting inside a physics sim with force perturbation and contact-transition penalties directly targets the weakness of naive human→robot mapping (dropped force / broken grasps) — the exact L4 failure mode the data pyramid §4 flags.
- Honest about the yield. The pipeline's quality filter is severe (see §4), which is itself the useful signal: internet video is abundant but most clips don't survive faithful reconstruction — quantifying how leaky the L1→L4 path really is.
flowchart LR
V[Monocular RGB video<br/>internet · ego · generated] --> H[HaWoR<br/>hand tracking]
V --> O[SAM 3D<br/>object mesh + pose]
O --> T[Guided-diffusion object tracker<br/>fix shape @ anchor, track pose via flow matching]
H --> A[Align hand+object<br/>centroid opt + GeoCalib gravity]
T --> A
A --> R[Dynamics-aware retargeting<br/>MPPI-style optim in MuJoCo Warp<br/>+ warmup · force perturbation · transition reward]
R --> D[Robot-executable trajectory<br/>22-DoF Sharpa Wave + dual UR3e @ 50 Hz]
- Reconstruction: HaWoR hands + SAM-3D object geometry; the object tracker holds shape fixed at an anchor frame and recovers per-frame pose entirely at inference via flow matching with adaptive guidance; hand and object are aligned across scales by centroid optimization and gravity alignment.
- Retargeting: sampling-based optimization in MuJoCo Warp (200 Hz sim). Three robustifications for noisy references: warmup (hold the object while the hand settles — the single largest gain, 0.25→0.66 on reconstructed data), random force perturbation (encourages robust grasps), and a transition reward penalizing failed contact transitions.
- Hardware: 22-DoF Sharpa Wave hand on dual UR3e arms, commanded at 50 Hz.
Retargeting success (all components vs annealed sampling alone):
| Source | Baseline | Full method |
|---|---|---|
| Reconstructed in-the-wild video | 25% | 71% |
| OakInk2 (clean MoCap) | 72% | 81% |
- Reconstruction quality: on 150 in-the-wild videos, human raters prefer the object tracking over FoundationPose 67% of the time; SOTA F-5/F-10/Chamfer on DexYCB/HOI4D.
- Real-world: 10 tasks — whisking, pouring, dusting, squeezing, tamping, erasing, stirring, hammering, spreading, picking.
- Dataset produced: 500 high-quality, human-verified trajectories — 53% internet, 31% egocentric, 16% generated video.
- Yield reality-check: from 2,000 100DOH clips, only 83 (4%) survived the reconstruction pass before the final 500 validated trajectories were assembled.
Significance. Do As I Do is the strongest 2026 demonstration that device-free everyday video can become deployable dexterous-hand data, and its physics-in-the-loop retargeting is a concrete recipe for the pyramid's load-bearing L4 bridge. It also quantifies the leakiness of the human-video base (4% survival), which most L1-scaling papers gloss over.
Limitations (authors').
- Assumes rigid objects and semi-accurate monocular metric depth — fails otherwise.
- Monocular hand-object distance ambiguity is inherent.
- No environmental reasoning — obstacles, articulated objects out of scope.
- Sim is an upper bound — physics approximation caps real-world transfer.
- Low yield — heavy quality filtering means raw internet scale ≠ usable-trajectory scale.
- Paper: arXiv 2606.19333
- Pyramid placement: L1 (web/ego video) → L4 (dynamics-aware retarget) → L5 (sim verify) — Dexterous-Hand Data Pyramid
- Counterpart bridge: AnyDexRT (calibration-free retargeting) · glove alternative: DexUMI · scaling law: EgoScale
- Dexterous Manipulation · Human Video → Robot Transfer