CoRL 2025 Visual Imitation Humanoid - Heungwoo/research GitHub Wiki
Venue: CoRL 2025 (Best Student Paper) ยท Authors: Allshire, Choi, Zhang, McAllister, A. Zhang, C. M. Kim, Darrell, Abbeel, Malik, Kanazawa โ UC Berkeley (BAIR) Project: VideoMimic ยท arXiv 2505.03729 ยท https://www.videomimic.net Category: Humanoid / Whole-Body Trend tag: Human video โ humanoid skills
flowchart LR
V[Monocular RGB videos<br/>everyday human activities] --> R[4D reconstruction<br/>human SMPL + scene mesh<br/>joint metric-scale optim.]
R --> RT[Retarget to Unitree G1]
RT --> SIM[PPO tracking in sim<br/>MoCap pretrain โ scene-conditioned]
SIM --> DIST[DAgger distill โ<br/>under-conditioned student<br/>heightmap + root commands]
DIST --> REAL[Real Unitree G1<br/>Fast-LIO2 + terrain map]
Humanoid control has been trained mostly on teleoperation or curated motion-capture. Internet-scale everyday human video is the obvious next data source, but retargeting to a humanoid and closing the sim-to-real gap at scale hadn't been shown convincingly.
Real-to-sim-to-real pipeline on the Unitree G1 (23-DoF, deployed at 50 Hz with low joint gains, Kpโ75).
Real-to-sim reconstruction (4D human + scene from monocular RGB): Grounded-SAM2 (tracking) โ VIMO (SMPL human mesh recovery) + ViTPose (2D keypoints) + BSTRO (foot contact); scene geometry via MegaSaM / MonST3R (depth + camera pose), GeoCalib (gravity alignment), and NKSR mesh generation. Human trajectory and scene scale are jointly optimized (Levenberg-Marquardt) to recover metric scale, then the SMPL motion is retargeted to the G1 (optional SMPL scale-adaptation to G1 proportions for climbing/dynamic skills).
Sim training โ 4 stages with PPO + DAgger distillation:
- MoCap pre-training โ RL on clean motion-capture to establish robust motor priors.
- Scene-conditioned tracking โ inject a local heightmap and track video-reconstructed motions over terrain.
- Distillation (DAgger) โ distill the teacher into a student that drops target-joint-angle conditioning and instead follows only root commands.
- Under-conditioned RL fine-tuning โ final RL pass on the reduced-observation policy.
The deployed student policy observes only proprioception (5-frame history), a local 11ร11 heightmap (0.1 m spacing around the torso), and root commands (desired x-y offset + yaw in the robot frame) โ i.e. it is conditioned on the environment, not a target trajectory. Onboard perception uses Fast-LIO2 SLAM + probabilistic terrain mapping. Sim-to-real relies on relaxed termination tolerances and domain randomization (mass, friction, latency, sensor noise).
CoRL 2025 Best Student Paper. A single context-conditioned policy performs staircase ascent/descent, sitting/standing from chairs and benches, and terrain traversal โ without task-specific reward engineering. Key ablation: removing the MoCap pre-training stage significantly degrades learning, since noisy video references and unstable initialization need a clean-motion foundation first. Authors release a dataset of 123 curated everyday-activity monocular videos.
First credible demonstration at scale that humanoid skills can be unlocked from everyday internet video. Slots into the CoRL 2025 trend of human video as the default cross-embodiment data source (DexUMI, UniSkill, ImMimic, X-Sim) and foreshadows ICLR 2026's Human-Video Pretraining and WholeBodyVLA.
- Paper (arXiv): https://arxiv.org/abs/2505.03729
- Project page: https://www.videomimic.net
- Code: https://github.com/hongsukchoi/VideoMimic
- CoRL 2025 awards: https://2025.corl.org/program/awards
- DexUMI ยท X-Sim ยท Survey (CoRL 2025)
โ Back to CoRL-2025