ICLR 2026 Lang To Loco - Heungwoo/research GitHub Wiki

Lang-To-Loco (RoboGhost) — Retargeting-free Humanoid Control via Motion Latent Guidance

Venue: ICLR 2026 Category: Humanoid Control — Language-Conditioned Trend tag: Multimodal / motion latents Authors: Zhe Li, Yangyang Wei, Boan Zhu, Yibo Peng, Tao Huang, et al. (equal contribution); Cheng Chi (project leader); corresponding Shanghang Zhang, Chang Xu. BAAI + HIT + HKUST + SJTU + PKU + U Sydney.

Approach diagram

flowchart LR
  subgraph Stage1[Stage 1 · Continuous Autoregressive Motion Generator]
    Text[Task instruction T] --> LaMP[LaMP text transformer · Li 2024d]
    Mocap[MotionMillion humanml + kungfu<br/>50,378 sequences pretraining] --> ME[Causal motion encoder]
    ME --> CT[Causal transformer · masked AR<br/>cos-schedule mask ratio γ τ]
    LaMP --> CT
    CT --> AdaLN[AdaLN] --> Diff[Diffusion model · MLP 16-layer · 5.84s]
    Diff --> Lref[Motion latent l_ref · 64-d]
  end
  subgraph Stage2[Stage 2 · Hybrid Teacher-Student RL]
    Priv[Privileged sim info<br/>root vel · global joints<br/>friction · motor strength<br/>reference motion 69-d] --> MoE["5-expert MoE teacher policy<br/>Actor 512-256-128 · PPO"]
    Lref --> Stud[Student diffusion policy<br/>4-MLP 256·256·256<br/>AdaLN conditioning]
    Prop[Proprioception 75 × 10 history] --> Stud
    MoE -->|DAgger optimal a^t| Stud
    Stud --> Act[23-dim joint targets · PD @ 50 Hz]
    Act --> Robot[Unitree G1 humanoid · Jetson Orin NX<br/>500 Hz low-level control]
  end
  CAS[Causal Adaptive Sampling<br/>α u = γ^u · weight failure antecedents] -.- MoE
Loading

Problem

Standard language-to-humanoid pipelines decode human motion from language, retarget it to robot morphology, then track with a physics-based controller. The paper's three concrete failures of this pipeline:

  1. Cumulative errors across decoding → retargeting → tracking.
  2. High latency from sequential stages (paper measures 17.85 s end-to-end with PHC-1000).
  3. Loose coupling — each stage optimized in isolation, no end-to-end gradient flow from semantics to action.

Existing fixes (RLPF, Serifi et al., LangWBC) tweak local stages but leave the overall pipeline fragile. RoboGhost's fix: make the motion latent a first-class conditioning signal, skip decoding and retargeting altogether.

Detailed Method

Stage 1: Continuous Autoregressive Motion Generator

Architecture. Causal autoencoder + continuous masked autoregressive (MAR) transformer with causal attention masks (replacing bidirectional). Differs from prior MAR works (MARDM, OmniMotion) by avoiding token shuffling/batch-token prediction and using causal masking to mitigate low-rank approximation limitations.

Masking schedule. Cosine ratio γ(τ) = cos(πτ/2), τ ~ U(0, 1). γ(τ)·N tokens are randomly masked.

Text encoder. LaMP (Li et al., 2024d) text transformer extracts language features.

Diffusion head on top of transformer outputs. The transformer-predicted latent representations condition a diffusion model that produces refined latents lref. Default backbone: 16-layer MLP (chosen over 4-layer DiT for latency — see ablations). Velocity-prediction objective (after SiT/MARDM) instead of noise prediction; improves dynamics consistency.

Stage 2: Hybrid Teacher-Student RL

Teacher policy (PPO oracle in IsaacGym):

  • Inputs: privileged sim info (root vel, global joints, friction, motor strength) + reference motion + proprioception.
  • Outputs: 23-D joint targets for PD control.
  • Mixture-of-Experts with 5 experts; gating network produces expert distribution; final action a = Σ pi·ai. Ablation in Figure 9 confirms 5 experts is optimal.
  • Actor MLP [512, 256, 128]; AdamW; lr 1e-4; β=(0.9, 0.999); γ=0.99; clip 0.2; entropy coef 0.005; value loss coef 1; 5 update epochs; 4 minibatches; max grad norm 1; batch size 4096.

Student diffusion policy (DAgger with diffusion):

  • No privileged info; takes 75-dim proprioception × 10 timesteps history + 64-D motion latent lref as input.
  • No retargeted reference motion — this is the paper's main bet.
  • 4-layer MLP backbone with hidden size 256; AdaLN injects motion latent into denoiser.
  • x0-prediction strategy. Loss: ‖a − ât‖² where a is reconstructed clean action.
  • DDIM sampling at inference for low latency.

Causal Adaptive Sampling (CAS)

Failures in long-horizon motion typically have antecedent causes (s steps prior — a misstep, collision). The motion sequence is divided into K equal-length intervals. After a failed rollout terminates at interval kt:

Δpi = α(t − i)·p, α(u) = γu, γ ∈ (0, 1), for i ∈ [t−s, t]; else 0

p′i ← pi + Δpi, then renormalized. Restart frames sampled from Multinomial(p′1, ..., p′K). Best parameters: λ=0.8, p=0.005 (Fig 8).

Reward & curriculum

  • Task rewards (Table 11): root velocity (10), root velocity direction (6), root angular velocity (1), keypoint position (10), feet position (12), DoF position (6), DoF velocity (6).
  • Penalties: DoF position limits (-10), torque limits (-5), termination (-200).
  • Regularization: DoF acceleration (-3e-7), action rate (-0.5), feet air time (10), feet contact force (-0.003), stumble (-2), waist roll-pitch error (-0.5), ankle action (-0.3).
  • Termination curriculum (after ASAP): start at 1.5 m tracking-error tolerance and progressively tighten.
  • Reference State Initialization (RSI): uniform sampling of starting phase τ ∈ [0, 1] across the reference motion.

Domain randomization

Friction U(0.5, 2.2); P-gain U(0.75, 1.25)·default; control delay U(20, 40) ms; push robot every 8 s with vxy=0.5 m/s.

Sim-to-real

  • IsaacGym for RL training → MuJoCo zero-shot transfer → Unitree G1 humanoid (Jetson Orin NX).
  • Policy at 50 Hz, low-level controller at 500 Hz.
  • Communication via LCM with 18–30 ms latency.

Datasets

MotionMillion (Fan et al., 2025): pretrain on full humanml + kungfu subsets (50,378 sequences); after stability filtering (CoM-CoP distance threshold) and policy filter (tracking error ≤0.6), the curated set has 3,261 humanml sequences and 200 kungfu sequences. Train two separate policies for the two subsets (large domain gap). Index-0 sequences only (no mirrored variants); 8:2 train/test split.

Comprehensive Results

Motion generation quality (Table 1)

Model R@1 / R@2 / R@3 (HumanML3D) FID ↓ MM-Dist ↓ Diversity → 27.49
Ground Truth 0.702 / 0.864 / 0.914 0.002 15.151 27.492
MDM 0.523 / 0.692 / 0.764 23.454 17.423 26.325
MLD 0.546 / 0.730 / 0.792 18.236 16.638 26.352
T2M-GPT 0.606 / 0.774 / 0.838 12.475 16.812 27.275
MotionGPT 0.456 / 0.598 / 0.628 14.375 16.892 27.114
MoMask 0.621 / 0.784 / 0.846 12.232 16.138 27.127
AttT2M 0.592 / 0.765 / 0.834 15.428 15.726 26.674
MotionStreamer 0.631 / 0.802 / 0.859 11.790 16.081 27.284
Ours-DDPM 0.639 / 0.808 / 0.867 11.706 15.772 27.230
Ours-SiT 0.641 / 0.812 / 0.870 11.743 15.663 27.307

On the HumanML (MotionMillion) subset, Ours-SiT reaches 0.646 R@1, FID 11.716, MM-Dist 15.603, Diversity 26.471.

Motion tracking in physics simulators (Table 2)

Method IsaacGym Succ ↑ Empjpe ↓ Empkpe ↓ MuJoCo Succ ↑ Empjpe ↓ Empkpe ↓
HumanML (MotionMillion)
Baseline (Exbody2-style) 0.92 0.23 0.19 0.64 0.34 0.31
Ours-DDPM 0.97 0.12 0.09 0.74 0.24 0.20
Ours-SiT 0.98 0.14 0.08 0.72 0.26 0.23
Kungfu (MotionMillion)
Baseline 0.66 0.43 0.37 0.51 0.58 0.52
Ours-DDPM 0.72 0.34 0.31 0.57 0.54 0.50
Ours-SiT 0.71 0.36 0.32 0.55 0.53 0.48

+5–6% success on HumanML, +6 on Kungfu — agile motions still hardest.

Latent-driven vs explicit retargeted (Table 4)

Method (HumanML test set) IsaacGym Succ Empjpe Empkpe MuJoCo Succ Empjpe Empkpe Time Cost (s)
Ours-Explicit (PHC 1000 iter) 0.93 0.21 0.17 0.66 0.32 0.27 17.85
Ours-Implicit (latent-driven) 0.97 0.12 0.09 0.74 0.24 0.20 5.84

+4 IsaacGym SR, +8 MuJoCo SR, latency cut 67% by skipping retargeting.

PHC retargeting iteration trade-off (Table 3)

Method Time (s) Succ Empjpe Empkpe
PHC-100 1.63 0.81 0.45 0.40
PHC-500 6.09 0.88 0.31 0.25
PHC-800 9.87 0.91 0.25 0.21
PHC-1000 11.89 0.93 0.21 0.17

Reducing PHC iterations to match RoboGhost's latency (~6 s) costs 5–10 SR points — confirms the retargeting bottleneck is real.

Diffusion vs MLP student policy (Table 5)

HumanML test set Generalization (unseen MotionMillion subsets)
MLP Policy: Succ / Empjpe / Empkpe 0.96 / 0.17 / 0.11 0.54 / 0.48 / 0.45
Diffusion Policy: 0.97 / 0.12 / 0.09 0.68 / 0.42 / 0.39

Diffusion's generalization advantage on unseen subsets (fitness, perform, 100style, haa) is +14 SR — the imperfect latents from those subsets are absorbed better by a multi-modal distribution model.

Diffusion backbone choice (Table 6)

Backbone IsaacGym Succ Empjpe Empkpe HumanML3D R@3 FID Time (s)
DiT (4-layer) 0.96 0.11 0.11 0.870 11.697 14.28
MLP (16-layer) 0.97 0.12 0.09 0.867 11.706 5.84

DiT marginally better on motion-generation metrics but no measurable tracking improvement and 2.4× higher latency. MLP chosen as default.

Robustness to observation noise (Figure 4)

When observations are corrupted with Gaussian noise scale 0.2 in MuJoCo:

  • MLP policy: maps noise to noisy actions → robot falls (failure).
  • Diffusion policy: success, robust tracking.

Maximum noise scale tolerated: MLP 0.12 vs Diffusion 0.33. The diffusion policy's training-time noise injection translates directly to test-time robustness.

Ablation Studies

Causal Adaptive Sampling hyper-parameters (Table 12)

CAS increases sampling probability of kinematically challenging frames. Optimal (Fig 8a-b): λ = 0.8, p = 0.005.

Number of MoE experts (Figure 9)

Variation in expert count has measurable but limited impact; 5 experts optimal.

Tracking policy comparison (Table 13)

The paper supplements with GMT-style architecture trained on the same data — confirms the +SR gap from RoboGhost is architecture-driven, not just data-driven (specific numbers in Appendix 9.6).

Limitations

The paper does not have an explicit "Limitations" section. Concrete gaps surfaced in the body and appendix:

  • Two separate policies for humanml vs kungfu subsets — domain gap is too large for a single policy. Generalist single-policy across all motion families remains open.
  • Real-world only on Unitree G1 — Unitree H1, Booster K1, Tien Kung etc. not validated.
  • Flat-ground constraint — motions with non-planar terrains or infeasible contacts are filtered out before training.
  • Stability filtering drops sequences with long unstable segments (CoM-CoP threshold). Some legitimate dynamic motions may be excluded.
  • Reference State Initialization is critical (per the paper); without it, hard motions don't converge — fragility at training start.
  • β scheduling — no adaptive scaling; the loss weighting is constant.
  • No robot-on-robot collision handling in the rewards.
  • Locomotion latency: 5.84 s end-to-end is much faster than the 17.85 s baseline, but still not real-time conversational. The bottleneck is now the motion generator (DDIM-accelerated MLP diffusion), not retargeting.
  • Diversity vs success trade-off in MotionMillion: kungfu success is still only 72% on IsaacGym, 57% on MuJoCo.

Significance & Positioning

RoboGhost removes the retargeting bottleneck that has long separated motion generation from humanoid control, and argues that motion latents — not joint trajectories — are the right shared interface between language and embodied control.

  • vs LangWBC (Shao et al., 2025): scales poorly with no generalization guarantee to unseen instructions. RoboGhost adds diffusion + MoE teacher to handle open-vocabulary language.
  • vs RLPF (Yue et al., 2025): risks catastrophic forgetting and limited diversity. RoboGhost's continuous AR + diffusion preserves diversity.
  • vs UH-1 (Mao et al., 2025): transformer-based large model relying on motion retargeting and discrete action tokenization. RoboGhost is exactly the retargeting-free counterpoint.
  • vs LeVERB (Xue et al., 2025): hierarchical CVAE-RL for vision-language WBC; lacks high-dynamic motion support. RoboGhost handles agile motions (Kungfu) on Unitree G1.
  • vs OmniH2O (He et al., 2024) / HumanPlus (Fu et al., 2024): these prioritize specific robustness at the cost of generality or long-term accuracy. RoboGhost's MoE teacher + diffusion student enhances both.
  • vs ExBody2 / GMT / Hover: these add masking, curriculum, MoE to enhance adaptability, but generalization is not guaranteed. RoboGhost's latent-driven design + diffusion student gives explicit generalization to unseen instructions (Table 5 right column).
  • vs BFM-Zero: both move toward general-purpose, promptable humanoid policies. BFM-Zero focuses on a foundation locomotion controller; RoboGhost focuses on the language→control conditioning interface.
  • vs HWC-Loco: orthogonal robustness axis (whole-body controller robustness vs language conditioning).
  • vs WholeBodyVLA: WholeBodyVLA is end-to-end VLA for whole-body manipulation; RoboGhost is the locomotion-first variant.

The recipe — causal-AR + diffusion motion generator → MoE PPO teacher → diffusion DAgger student conditioned on latents, not retargeted motions — is the cleanest architectural answer to "how do we get from open-ended language to executable humanoid actions" without the multi-stage retargeting pipeline.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️