RSS 2026 X Loco - Heungwoo/research GitHub Wiki
X-Loco: Towards Generalist Humanoid Locomotion Control via Synergetic Policy Distillation
Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Humanoids · paper #22 Authors: Dewei Wang, Xinmiao Wang, Chenyun Zhang, Jiyuan Shi, Yingnan Zhao, Chenjia Bai, Xuelong Li arXiv: 2603.03733 · program page
Summary compiled from the arXiv paper (v2); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Figure 1: the Unitree G1 running the single X-Loco policy through its skill spectrum — stair/box climbing, forward rolls, fall recovery back to standing, and traversal of complex terrains — driven only by velocity commands, proprioception, and depth vision (no reference motions).
Problem
Humanoid controllers are fragmented by skill: exteroceptive terrain traversal, fall recovery, and whole-body coordination (WBC) are each solved by separate policies with conflicting reward designs and dynamics, so a robot cannot, e.g., autonomously resume walking after a fall. Motion-tracking and teleoperation approaches cover diverse behaviors but need reference motions or human input, sacrificing autonomy.
Method
X-Loco trains three privileged oracle specialists sharing a unified action space — upright locomotion, fall recovery, and whole-body coordination (rolling, box climbing) — then distills them into one vision-based generalist student via synergetic policy distillation. Case-Adaptive Specialist Selection (CASS) queries the most relevant specialist for action guidance based on robot state and terrain; Specialist Annealing Rollout (SAR) mixes specialist-driven rollouts into training with a hysteresis-scheduled ratio (decay step 1e-4 per iteration when distillation MSE < 0.005, paused above 0.010); Stochastic Fall Injection (SFI) applies external disturbances during walking/climbing (conditioned on vulnerable scenarios like high-speed turns) so the policy learns locomotion-to-recovery transitions. The student uses an MoE architecture (2 experts) and a custom NVIDIA Warp parallel ray-casting depth-rendering pipeline; training is in IsaacLab, deployment on a Unitree G1 at 50 Hz over a 500 Hz PD loop.
Results
In simulation (Table I), X-Loco averages 0.939 success on Upright Locomotion (vs. 0.799 PPO, 0.550 AHC; MoRE 0.921), 0.871 on WBC where MoRE/AHC/PPO fail entirely (BeyondMimic reaches 0.958 but cannot locomote or recover), and a perfect 1.000 on Recovery — matching the recovery specialist and recovering about 94.8% of the locomotion specialist's success rate. Ablations show CASS, SAR, SFI, and the MoE head each matter: without MoE success rates degrade across tasks, and SFI yields a higher post-fall recovery success rate (Table IV). Real-world G1 deployments demonstrate seamless transitions among the three core skill regimes, including resuming climbing after an induced fall on hybrid terrain.
Significance
First vision-based humanoid controller unifying locomotion, whole-body coordination, and fall recovery under pure velocity commands — a distillation recipe for composing conflicting skills that complements the humanoid-foundation-policy thread. Related threads: Review-Humanoid-VLA · Review-System-0-1-2.
← Back to RSS 2026 survey · RSS-2026-Papers · Home