Review DoAsIDo - Heungwoo/research GitHub Wiki

In-Depth Review — Do As I Do: Dexterous Manipulation Data from Everyday Human Videos

Paper: "Do as I Do: Dexterous Manipulation Data from Everyday Human Videos" — arXiv 2606.19333 (Jun 17 2026) Authors: Bhawna Paliwal, Haritheja Etukuru, William Liang, Pieter Abbeel, Nur Muhammad Mahi Shafiullah, Jitendra Malik · UC Berkeley / NYU What it is: a device-free pipeline that turns ordinary monocular RGB human videos into robot-executable dexterous-hand trajectories — the cheapest L1→L4 path in the Dexterous-Hand Data Pyramid.

Companions: Dexterous-Hand Data Pyramid · AnyDexRT (the retargeting-bridge counterpart) · Dexterous Manipulation · Human Video → Robot Transfer.


1. TL;DR

  1. Everyday video → dexterous robot data, no gloves or teleop. Do As I Do reconstructs 4D hand-object dynamics from a single monocular RGB clip (internet, egocentric, or generated) and dynamics-aware retargets it into actions for a real dexterous hand.
  2. Reconstruction stack: HaWoR (hand tracking) + SAM 3D (object mesh/pose) + a guided-diffusion object tracker (fix shape at an anchor frame, track pose via flow matching at inference) + centroid/gravity alignment (GeoCalib).
  3. Dynamics-aware retargeting by sampling-based (MPPI-style) optimization in MuJoCo Warp, robustified for noisy references with warmup steps, random force perturbation, and a transition reward.
  4. Results: retargeting success 25% → 71% on reconstructed in-the-wild video and 72% → 81% on clean OakInk2 MoCap (vs annealed-sampling alone); object tracking preferred over FoundationPose 67% of the time. Ships 500 human-verified trajectories deployed on a 22-DoF Sharpa Wave hand across 10 real tasks.

2. Why it matters

  • It removes the capture device from the bottom of the pyramid. Where DexUMI needs a sensorized glove and teleop needs a rig, Do As I Do works from any RGB video — the most abundant (L1) data — and carries it through the retargeting bridge (L4) with physics verification (L5). This is the device-free extreme of the human-video fork.
  • Dynamics-aware, not kinematic-only. Retargeting inside a physics sim with force perturbation and contact-transition penalties directly targets the weakness of naive human→robot mapping (dropped force / broken grasps) — the exact L4 failure mode the data pyramid §4 flags.
  • Honest about the yield. The pipeline's quality filter is severe (see §4), which is itself the useful signal: internet video is abundant but most clips don't survive faithful reconstruction — quantifying how leaky the L1→L4 path really is.

3. Method

flowchart LR
  V[Monocular RGB video<br/>internet · ego · generated] --> H[HaWoR<br/>hand tracking]
  V --> O[SAM 3D<br/>object mesh + pose]
  O --> T[Guided-diffusion object tracker<br/>fix shape @ anchor, track pose via flow matching]
  H --> A[Align hand+object<br/>centroid opt + GeoCalib gravity]
  T --> A
  A --> R[Dynamics-aware retargeting<br/>MPPI-style optim in MuJoCo Warp<br/>+ warmup · force perturbation · transition reward]
  R --> D[Robot-executable trajectory<br/>22-DoF Sharpa Wave + dual UR3e @ 50 Hz]
Loading
  • Reconstruction: HaWoR hands + SAM-3D object geometry; the object tracker holds shape fixed at an anchor frame and recovers per-frame pose entirely at inference via flow matching with adaptive guidance; hand and object are aligned across scales by centroid optimization and gravity alignment.
  • Retargeting: sampling-based optimization in MuJoCo Warp (200 Hz sim). Three robustifications for noisy references: warmup (hold the object while the hand settles — the single largest gain, 0.25→0.66 on reconstructed data), random force perturbation (encourages robust grasps), and a transition reward penalizing failed contact transitions.
  • Hardware: 22-DoF Sharpa Wave hand on dual UR3e arms, commanded at 50 Hz.

4. Results (paper-reported)

Retargeting success (all components vs annealed sampling alone):

Source Baseline Full method
Reconstructed in-the-wild video 25% 71%
OakInk2 (clean MoCap) 72% 81%
  • Reconstruction quality: on 150 in-the-wild videos, human raters prefer the object tracking over FoundationPose 67% of the time; SOTA F-5/F-10/Chamfer on DexYCB/HOI4D.
  • Real-world: 10 tasks — whisking, pouring, dusting, squeezing, tamping, erasing, stirring, hammering, spreading, picking.
  • Dataset produced: 500 high-quality, human-verified trajectories — 53% internet, 31% egocentric, 16% generated video.
  • Yield reality-check: from 2,000 100DOH clips, only 83 (4%) survived the reconstruction pass before the final 500 validated trajectories were assembled.

5. Significance & limitations

Significance. Do As I Do is the strongest 2026 demonstration that device-free everyday video can become deployable dexterous-hand data, and its physics-in-the-loop retargeting is a concrete recipe for the pyramid's load-bearing L4 bridge. It also quantifies the leakiness of the human-video base (4% survival), which most L1-scaling papers gloss over.

Limitations (authors').

  1. Assumes rigid objects and semi-accurate monocular metric depth — fails otherwise.
  2. Monocular hand-object distance ambiguity is inherent.
  3. No environmental reasoning — obstacles, articulated objects out of scope.
  4. Sim is an upper bound — physics approximation caps real-world transfer.
  5. Low yield — heavy quality filtering means raw internet scale ≠ usable-trajectory scale.

6. Links

← Back to Reviews · Home

⚠️ **GitHub.com Fallback** ⚠️