RSS 2026 Unlocking In the Wild Loco Manipulation with Robot Free - Heungwoo/research GitHub Wiki
Unlocking In-the-Wild Loco-Manipulation with Robot-Free Egocentric Demonstration
Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Imitation learning 3 · paper #204 Authors: Modi Shi, Shijiapeng, Jin Chen, Haoran Jiang, Tianyu Li, Ping Luo, Di Huang, Hongyang Li, Li Chen arXiv: 2602.10106 · program page
Summary compiled from the arXiv paper (v2, titled "EgoHumanoid: Unlocking In-the-Wild Loco-Manipulation with Robot-Free Egocentric Demonstration"); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Figure 1: in-the-wild egocentric human data (left; diverse objects, viewpoints, environments — outdoor, home, supermarket) and lab-bound teleoperated robot data (right) are bridged by view alignment and action alignment for VLA co-training (center); deployment (bottom) generalizes without wild robot data, with the generalization score climbing +51% over the robot-only baseline as human data scales to 300 demos.
Problem
Humanoid loco-manipulation is data-hungry, but teleoperation confines collection to labs (hardware, safety, mocap logistics), while abundant egocentric human demonstrations carry a severe embodiment gap — different morphology, camera height/viewpoint, and walking dynamics — that is amplified by whole-body movement.
Method
EgoHumanoid co-trains a VLA policy (finetuned π0.5, egocentric RGB + language in, unified actions out, proprioception deliberately omitted) on teleoperated Unitree G1 data plus robot-free human demonstrations captured with a portable PICO VR rig (headset + 5 motion trackers + head-mounted ZED X Mini; ~2x faster than teleoperation per demo). The alignment pipeline has two parts. View alignment: MoGe depth estimation, reprojection of the 3D point map into the robot camera frame, and latent-diffusion inpainting of disoccluded regions. Action alignment: a unified action space of 6-DoF delta end-effector poses for the upper body (pelvis-centric, Savitzky–Golay smoothed, SO(3) tangent-space filtering, 100→20 Hz), discrete constant-velocity locomotion primitives derived from the human pelvis trajectory, and a binary gripper state inferred from finger-curvature thresholds; mini-batch human:robot sampling ratios are tuned per task.
Results
On four real tasks (Pillow Placement, Trash Disposal, Toy Transfer, Cart Stowing; 20 trials per setting, 100 robot + 300 human episodes per task), co-training reaches 78% average score in-domain vs 59% robot-only, and 82% vs 31% in generalization scenes seen only in human data — the 51% gain. Subtask analysis: navigation transfers almost fully from human data alone (100% on navigation-dominated sub-steps), coarse manipulation transfers well, but precision-critical phases need robot data (Cart Stowing s2: human-only 5%, robot-only 15%, co-training 60%). Performance scales consistently with human demos from 0 to 300; coarse tasks favor 1:2 robot:human sampling, fine manipulation 2:1.
Significance
First evidence that in-the-wild egocentric human video-action data can drive whole-body humanoid loco-manipulation, extending the human-data co-training recipe beyond tabletop arms — a direct data point for Review-Human-Video-Transfer and the humanoid VLA thread in Review-Humanoid-VLA.
← Back to RSS 2026 survey · RSS-2026-Papers · Home