RSS 2026 DexImit - Heungwoo/research GitHub Wiki
DexImit: Learning Bimanual Dexterous Manipulation from Monocular Human Videos
Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Manipulation 1 · paper #3 Authors: Juncheng Mu, Sizhe Yang, Yiming Bao, Hojin Bae, Tianming Wei, Linning Xu, Boyi Li, Huazhe Xu, Jiangmiao Pang arXiv: 2602.10105 · program page
Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Figure 1 teaser: on the left, generated or in-the-wild human manipulation videos (pouring water, placing fruits, stacking cups, cutting, etc.); on the right, a gallery of bimanual dexterous-hand manipulations synthesized by DexImit from those videos, spanning tool-using, long-horizon, and fine-grained tasks.
Problem
Bimanual dexterous manipulation is starved for data: teleoperating multi-fingered hands is hard and hardware is expensive, so collecting large-scale demonstrations is far more costly than for simple jaw-grippers. Human manipulation videos (from the Internet or text-to-video models) carry the needed manipulation knowledge, but the embodiment gap between human and robot hands makes direct pretraining on them ineffective, and prior video-to-robot reconstruction pipelines rely on absolute depth or strict reconstruction accuracy.
Method
DexImit converts monocular human videos into physically plausible robot data via a four-stage pipeline: (1) depth-free 4D reconstruction of hand–object interactions at near-metric scale — Qwen3-VL for video understanding, Grounded SAM2 segmentation, SpatialTracker v2 depth with a hand-size-prior scale factor, SAM3D image-to-3D object generation, Wilor hand meshes, and a tracking variant of FoundationPose for 6D object poses, all mapped into a table-anchored world frame; (2) subtask decomposition plus an Action-Centric Scheduling Algorithm that queues pregrasp/grasp/motion/release subactions across arbitrary horizons and bimanual concurrency; (3) force-closure-based grasp synthesis ranked by proximity to the reconstructed human hand pose, with keyframe-based motion planning; (4) comprehensive augmentation of object pose and scale, camera pose, and visual observations (e.g., point-cloud noise) for zero-shot sim-to-real transfer. Policies are trained on the generated data with DP3.
Results
In a reconstruction study over 100 short-horizon tasks, the chosen ST2+FoundationPose++ combination reaches an 82% object-trajectory reconstruction success rate vs. 76% for ST2+PCR, 45% DA3+PCR, 38% TA+RANSAC, 32% VGGT+PCR, and 11% TA+PCR (Table I). On six simulated data-quality tasks (Table II), DexImit tops both re-implemented baselines: 100% on Put Cup, Grapefruit, Fruits, and Pour, 78% on the long-horizon Pot task, and 52% on Stack Six Cups, where RigVid and DexMan both fail on the long-horizon/fine-grained settings. Zero-shot real-world deployment uses two UR5e arms with XHands and an Azure Kinect on four meta-tasks; ablations show removing scale augmentation, regenerating grasps per scale, or dropping point-cloud noise augmentation each significantly degrades success.
Significance
A depth-free, fully automated route from ordinary (or generated) human videos to training-ready bimanual dexterous data, directly attacking the data-scarcity bottleneck highlighted across the Manipulation sessions. Related threads: Review-Dexterous-Manipulation · Review-Human-Video-Transfer.
← Back to RSS 2026 survey · RSS-2026-Papers · Home