CoRL 2026 HuRo - Heungwoo/research GitHub Wiki
CoRL 2026 — HuRo: Robotizing Human Videos for Scalable VLA Pretraining
Venue: CoRL 2026 (Austin, TX, Nov 9–12) · RLWRLD + Yonsei. Paper: arXiv 2609.10706. Representative of: robotizing human video at scale — turn egocentric human video into robot-aligned obs+actions for VLA pretraining. Companions: Human Video → Robot Transfer · Egocentric Video Pre-Training · CoRL 2026 survey.

1. Problem
Human video is abundant and diverse, but a large embodiment gap — in both observations (human arms/hands vs. robot) and actions — blocks its use as direct supervision for VLA policies. Prior work either robotizes video only in task-matched settings, or aligns observations and actions independently. HuRo asks whether heterogeneous egocentric human videos, with varying annotation levels, can be turned into a scalable source of robot-aligned supervision.
2. Method
A robotization pipeline converts human videos into robot-aligned observations and action trajectories, while inferring the intermediate signals that raw sources lack.
- Motion retargeting: MANO-based hand-pose estimation extracts hand motion; fingertip positions and hand structure are mapped to robot joint trajectories via PyRoKi optimization, with a chunk-level world→robot-base transform and hand-motion + ego-view objectives under kinematic regularization.
- Visual robotization: SAM2 segments and ProPainter inpaints the human arm out of frames, then an Isaac Sim–rendered robot is overlaid using estimated camera intrinsics and the retargeted joint configuration.
- Inferring missing signals: camera intrinsics (droidcalib/AnyCalib), hand poses (HAWOR), camera trajectories (masked DROID-SLAM + MoGe-2 metric scaling), and language instructions (VLM captioning) are recovered per source.
3. Results
- Scale: ~630K robotized episodes / 142M frames (~1,317 h @ 30 fps) from five sources — EgoDex (56%), EgoVerse (27%), Ego4D (10%), Ego10K (6%), EPIC-Kitchens (2%).
- Real-world (ALLEX robot), 0%→100% pretraining scale: overall completion 51.5% → 80.3%; in-distribution 68.1% → 88.4%; OOD (spatial + visual shifts) 34.9% → 72.2%.
- Ablations: visual robotization drives OOD robustness (72.2% with robot overlay vs. 55.7% without, ~16.5 pts); end-to-end visual+action pretraining beats visual-only transfer on Diverse Pick-and-Place (61.1% ID / 50.0% OOD vs. 44.4% / 27.8%).
4. Why it matters
HuRo treats human video not just as a visual prior but as a source of both robot-aligned pixels and executable action supervision, unifying observation and action transfer in one heterogeneous-source pipeline — a route to scaling VLA pretraining beyond costly teleoperation.
Limitations (reviewer): (1) robotized-observation fidelity is bounded by reconstruction and rendering quality; (2) no force/tactile supervision, limiting contact-rich manipulation; (3) kinematic retargeting ignores self-collision — only 55.2% of trajectories are free of non-grasp self-contact.
5. Links
- arXiv 2609.10706
- Survey: CoRL 2026 · Related: Human Video → Robot Transfer · Egocentric Video Pre-Training
← Back to CoRL 2026 survey · Home