ICLR 2026 Lang To Loco - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Category: Humanoid Control — Language-Conditioned Trend tag: Multimodal / motion latents Authors: Zhe Li, Yangyang Wei, Boan Zhu, Yibo Peng, Tao Huang, et al. (equal contribution); Cheng Chi (project leader); corresponding Shanghang Zhang, Chang Xu. BAAI + HIT + HKUST + SJTU + PKU + U Sydney.
flowchart LR
subgraph Stage1[Stage 1 · Continuous Autoregressive Motion Generator]
Text[Task instruction T] --> LaMP[LaMP text transformer · Li 2024d]
Mocap[MotionMillion humanml + kungfu<br/>50,378 sequences pretraining] --> ME[Causal motion encoder]
ME --> CT[Causal transformer · masked AR<br/>cos-schedule mask ratio γ τ]
LaMP --> CT
CT --> AdaLN[AdaLN] --> Diff[Diffusion model · MLP 16-layer · 5.84s]
Diff --> Lref[Motion latent l_ref · 64-d]
end
subgraph Stage2[Stage 2 · Hybrid Teacher-Student RL]
Priv[Privileged sim info<br/>root vel · global joints<br/>friction · motor strength<br/>reference motion 69-d] --> MoE["5-expert MoE teacher policy<br/>Actor 512-256-128 · PPO"]
Lref --> Stud[Student diffusion policy<br/>4-MLP 256·256·256<br/>AdaLN conditioning]
Prop[Proprioception 75 × 10 history] --> Stud
MoE -->|DAgger optimal a^t| Stud
Stud --> Act[23-dim joint targets · PD @ 50 Hz]
Act --> Robot[Unitree G1 humanoid · Jetson Orin NX<br/>500 Hz low-level control]
end
CAS[Causal Adaptive Sampling<br/>α u = γ^u · weight failure antecedents] -.- MoE
Standard language-to-humanoid pipelines decode human motion from language, retarget it to robot morphology, then track with a physics-based controller. The paper's three concrete failures of this pipeline:
- Cumulative errors across decoding → retargeting → tracking.
- High latency from sequential stages (paper measures 17.85 s end-to-end with PHC-1000).
- Loose coupling — each stage optimized in isolation, no end-to-end gradient flow from semantics to action.
Existing fixes (RLPF, Serifi et al., LangWBC) tweak local stages but leave the overall pipeline fragile. RoboGhost's fix: make the motion latent a first-class conditioning signal, skip decoding and retargeting altogether.
Architecture. Causal autoencoder + continuous masked autoregressive (MAR) transformer with causal attention masks (replacing bidirectional). Differs from prior MAR works (MARDM, OmniMotion) by avoiding token shuffling/batch-token prediction and using causal masking to mitigate low-rank approximation limitations.
Masking schedule. Cosine ratio γ(τ) = cos(πτ/2), τ ~ U(0, 1). γ(τ)·N tokens are randomly masked.
Text encoder. LaMP (Li et al., 2024d) text transformer extracts language features.
Diffusion head on top of transformer outputs. The transformer-predicted latent representations condition a diffusion model that produces refined latents lref. Default backbone: 16-layer MLP (chosen over 4-layer DiT for latency — see ablations). Velocity-prediction objective (after SiT/MARDM) instead of noise prediction; improves dynamics consistency.
Teacher policy (PPO oracle in IsaacGym):
- Inputs: privileged sim info (root vel, global joints, friction, motor strength) + reference motion + proprioception.
- Outputs: 23-D joint targets for PD control.
- Mixture-of-Experts with 5 experts; gating network produces expert distribution; final action a = Σ pi·ai. Ablation in Figure 9 confirms 5 experts is optimal.
- Actor MLP [512, 256, 128]; AdamW; lr 1e-4; β=(0.9, 0.999); γ=0.99; clip 0.2; entropy coef 0.005; value loss coef 1; 5 update epochs; 4 minibatches; max grad norm 1; batch size 4096.
Student diffusion policy (DAgger with diffusion):
- No privileged info; takes 75-dim proprioception × 10 timesteps history + 64-D motion latent lref as input.
- No retargeted reference motion — this is the paper's main bet.
- 4-layer MLP backbone with hidden size 256; AdaLN injects motion latent into denoiser.
- x0-prediction strategy. Loss: ‖a − ât‖² where a is reconstructed clean action.
- DDIM sampling at inference for low latency.
Failures in long-horizon motion typically have antecedent causes (s steps prior — a misstep, collision). The motion sequence is divided into K equal-length intervals. After a failed rollout terminates at interval kt:
Δpi = α(t − i)·p, α(u) = γu, γ ∈ (0, 1), for i ∈ [t−s, t]; else 0
p′i ← pi + Δpi, then renormalized. Restart frames sampled from Multinomial(p′1, ..., p′K). Best parameters: λ=0.8, p=0.005 (Fig 8).
- Task rewards (Table 11): root velocity (10), root velocity direction (6), root angular velocity (1), keypoint position (10), feet position (12), DoF position (6), DoF velocity (6).
- Penalties: DoF position limits (-10), torque limits (-5), termination (-200).
- Regularization: DoF acceleration (-3e-7), action rate (-0.5), feet air time (10), feet contact force (-0.003), stumble (-2), waist roll-pitch error (-0.5), ankle action (-0.3).
- Termination curriculum (after ASAP): start at 1.5 m tracking-error tolerance and progressively tighten.
- Reference State Initialization (RSI): uniform sampling of starting phase τ ∈ [0, 1] across the reference motion.
Friction U(0.5, 2.2); P-gain U(0.75, 1.25)·default; control delay U(20, 40) ms; push robot every 8 s with vxy=0.5 m/s.
- IsaacGym for RL training → MuJoCo zero-shot transfer → Unitree G1 humanoid (Jetson Orin NX).
- Policy at 50 Hz, low-level controller at 500 Hz.
- Communication via LCM with 18–30 ms latency.
MotionMillion (Fan et al., 2025): pretrain on full humanml + kungfu subsets (50,378 sequences); after stability filtering (CoM-CoP distance threshold) and policy filter (tracking error ≤0.6), the curated set has 3,261 humanml sequences and 200 kungfu sequences. Train two separate policies for the two subsets (large domain gap). Index-0 sequences only (no mirrored variants); 8:2 train/test split.
| Model | R@1 / R@2 / R@3 (HumanML3D) | FID ↓ | MM-Dist ↓ | Diversity → 27.49 |
|---|---|---|---|---|
| Ground Truth | 0.702 / 0.864 / 0.914 | 0.002 | 15.151 | 27.492 |
| MDM | 0.523 / 0.692 / 0.764 | 23.454 | 17.423 | 26.325 |
| MLD | 0.546 / 0.730 / 0.792 | 18.236 | 16.638 | 26.352 |
| T2M-GPT | 0.606 / 0.774 / 0.838 | 12.475 | 16.812 | 27.275 |
| MotionGPT | 0.456 / 0.598 / 0.628 | 14.375 | 16.892 | 27.114 |
| MoMask | 0.621 / 0.784 / 0.846 | 12.232 | 16.138 | 27.127 |
| AttT2M | 0.592 / 0.765 / 0.834 | 15.428 | 15.726 | 26.674 |
| MotionStreamer | 0.631 / 0.802 / 0.859 | 11.790 | 16.081 | 27.284 |
| Ours-DDPM | 0.639 / 0.808 / 0.867 | 11.706 | 15.772 | 27.230 |
| Ours-SiT | 0.641 / 0.812 / 0.870 | 11.743 | 15.663 | 27.307 |
On the HumanML (MotionMillion) subset, Ours-SiT reaches 0.646 R@1, FID 11.716, MM-Dist 15.603, Diversity 26.471.
| Method | IsaacGym Succ ↑ | Empjpe ↓ | Empkpe ↓ | MuJoCo Succ ↑ | Empjpe ↓ | Empkpe ↓ |
|---|---|---|---|---|---|---|
| HumanML (MotionMillion) | ||||||
| Baseline (Exbody2-style) | 0.92 | 0.23 | 0.19 | 0.64 | 0.34 | 0.31 |
| Ours-DDPM | 0.97 | 0.12 | 0.09 | 0.74 | 0.24 | 0.20 |
| Ours-SiT | 0.98 | 0.14 | 0.08 | 0.72 | 0.26 | 0.23 |
| Kungfu (MotionMillion) | ||||||
| Baseline | 0.66 | 0.43 | 0.37 | 0.51 | 0.58 | 0.52 |
| Ours-DDPM | 0.72 | 0.34 | 0.31 | 0.57 | 0.54 | 0.50 |
| Ours-SiT | 0.71 | 0.36 | 0.32 | 0.55 | 0.53 | 0.48 |
+5–6% success on HumanML, +6 on Kungfu — agile motions still hardest.
| Method (HumanML test set) | IsaacGym Succ | Empjpe | Empkpe | MuJoCo Succ | Empjpe | Empkpe | Time Cost (s) |
|---|---|---|---|---|---|---|---|
| Ours-Explicit (PHC 1000 iter) | 0.93 | 0.21 | 0.17 | 0.66 | 0.32 | 0.27 | 17.85 |
| Ours-Implicit (latent-driven) | 0.97 | 0.12 | 0.09 | 0.74 | 0.24 | 0.20 | 5.84 |
+4 IsaacGym SR, +8 MuJoCo SR, latency cut 67% by skipping retargeting.
| Method | Time (s) | Succ | Empjpe | Empkpe |
|---|---|---|---|---|
| PHC-100 | 1.63 | 0.81 | 0.45 | 0.40 |
| PHC-500 | 6.09 | 0.88 | 0.31 | 0.25 |
| PHC-800 | 9.87 | 0.91 | 0.25 | 0.21 |
| PHC-1000 | 11.89 | 0.93 | 0.21 | 0.17 |
Reducing PHC iterations to match RoboGhost's latency (~6 s) costs 5–10 SR points — confirms the retargeting bottleneck is real.
| HumanML test set | Generalization (unseen MotionMillion subsets) | |
|---|---|---|
| MLP Policy: Succ / Empjpe / Empkpe | 0.96 / 0.17 / 0.11 | 0.54 / 0.48 / 0.45 |
| Diffusion Policy: | 0.97 / 0.12 / 0.09 | 0.68 / 0.42 / 0.39 |
Diffusion's generalization advantage on unseen subsets (fitness, perform, 100style, haa) is +14 SR — the imperfect latents from those subsets are absorbed better by a multi-modal distribution model.
| Backbone | IsaacGym Succ | Empjpe | Empkpe | HumanML3D R@3 | FID | Time (s) |
|---|---|---|---|---|---|---|
| DiT (4-layer) | 0.96 | 0.11 | 0.11 | 0.870 | 11.697 | 14.28 |
| MLP (16-layer) | 0.97 | 0.12 | 0.09 | 0.867 | 11.706 | 5.84 |
DiT marginally better on motion-generation metrics but no measurable tracking improvement and 2.4× higher latency. MLP chosen as default.
When observations are corrupted with Gaussian noise scale 0.2 in MuJoCo:
- MLP policy: maps noise to noisy actions → robot falls (failure).
- Diffusion policy: success, robust tracking.
Maximum noise scale tolerated: MLP 0.12 vs Diffusion 0.33. The diffusion policy's training-time noise injection translates directly to test-time robustness.
CAS increases sampling probability of kinematically challenging frames. Optimal (Fig 8a-b): λ = 0.8, p = 0.005.
Variation in expert count has measurable but limited impact; 5 experts optimal.
The paper supplements with GMT-style architecture trained on the same data — confirms the +SR gap from RoboGhost is architecture-driven, not just data-driven (specific numbers in Appendix 9.6).
The paper does not have an explicit "Limitations" section. Concrete gaps surfaced in the body and appendix:
- Two separate policies for humanml vs kungfu subsets — domain gap is too large for a single policy. Generalist single-policy across all motion families remains open.
- Real-world only on Unitree G1 — Unitree H1, Booster K1, Tien Kung etc. not validated.
- Flat-ground constraint — motions with non-planar terrains or infeasible contacts are filtered out before training.
- Stability filtering drops sequences with long unstable segments (CoM-CoP threshold). Some legitimate dynamic motions may be excluded.
- Reference State Initialization is critical (per the paper); without it, hard motions don't converge — fragility at training start.
- β scheduling — no adaptive scaling; the loss weighting is constant.
- No robot-on-robot collision handling in the rewards.
- Locomotion latency: 5.84 s end-to-end is much faster than the 17.85 s baseline, but still not real-time conversational. The bottleneck is now the motion generator (DDIM-accelerated MLP diffusion), not retargeting.
- Diversity vs success trade-off in MotionMillion: kungfu success is still only 72% on IsaacGym, 57% on MuJoCo.
RoboGhost removes the retargeting bottleneck that has long separated motion generation from humanoid control, and argues that motion latents — not joint trajectories — are the right shared interface between language and embodied control.
- vs LangWBC (Shao et al., 2025): scales poorly with no generalization guarantee to unseen instructions. RoboGhost adds diffusion + MoE teacher to handle open-vocabulary language.
- vs RLPF (Yue et al., 2025): risks catastrophic forgetting and limited diversity. RoboGhost's continuous AR + diffusion preserves diversity.
- vs UH-1 (Mao et al., 2025): transformer-based large model relying on motion retargeting and discrete action tokenization. RoboGhost is exactly the retargeting-free counterpoint.
- vs LeVERB (Xue et al., 2025): hierarchical CVAE-RL for vision-language WBC; lacks high-dynamic motion support. RoboGhost handles agile motions (Kungfu) on Unitree G1.
- vs OmniH2O (He et al., 2024) / HumanPlus (Fu et al., 2024): these prioritize specific robustness at the cost of generality or long-term accuracy. RoboGhost's MoE teacher + diffusion student enhances both.
- vs ExBody2 / GMT / Hover: these add masking, curriculum, MoE to enhance adaptability, but generalization is not guaranteed. RoboGhost's latent-driven design + diffusion student gives explicit generalization to unseen instructions (Table 5 right column).
- vs BFM-Zero: both move toward general-purpose, promptable humanoid policies. BFM-Zero focuses on a foundation locomotion controller; RoboGhost focuses on the language→control conditioning interface.
- vs HWC-Loco: orthogonal robustness axis (whole-body controller robustness vs language conditioning).
- vs WholeBodyVLA: WholeBodyVLA is end-to-end VLA for whole-body manipulation; RoboGhost is the locomotion-first variant.
The recipe — causal-AR + diffusion motion generator → MoE PPO teacher → diffusion DAgger student conditioned on latents, not retargeted motions — is the cleanest architectural answer to "how do we get from open-ended language to executable humanoid actions" without the multi-stage retargeting pipeline.
- OpenReview: https://openreview.net/forum?id=k3Cyx3Uets
- PDF: https://openreview.net/pdf?id=k3Cyx3Uets
← Back to ICLR-2026