ICLR 2026 HWC Loco - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Authors: Sixu Lin, Guanren Qiao, Yunxin Tai, Ang Li, Kui Jia, Guiliang Liu (CUHK Shenzhen; Harbin Institute of Technology; DexForce Technology; Southeast University — corresponding: Guiliang Liu) Category: Humanoid Control — Locomotion Trend tag: Robust RL / safety-critical recovery
flowchart LR
Cmd[Velocity command v_x, v_y, w_z] --> Pi0[High-level π_0<br/>discrete switch policy<br/>Double-DQN, ε-greedy]
Obs[oH_t = ot..ot-H<br/>+ VAE-estimated P_t<br/>velocity + ZMP features] --> Pi0
Pi0 -->|goal-tracking| Pi1[π_1: PPO + Wasserstein imitation<br/>vs. CMU MoCap retargeted]
Pi0 -->|safety recovery| Pi2[π_2: robust optim. over P_α^L<br/>+ ZMP constraint]
Pi1 --> PD[PD controller 100 Hz<br/>target joint angles]
Pi2 --> PD
PD --> Robot[Unitree H1 / G1]
Reinforcement-learning humanoid locomotion policies trained in simulation suffer two failure modes: (1) Sim2Real transition mismatch causes falls under realistic disturbances (slips, pushes, terrain edges, payload changes); (2) standard robust-RL formulations (max-min over worst-case dynamics) produce overly conservative policies that fail to track velocity commands. HWC-Loco asks how to keep aggressive goal-tracking and guaranteed safety recovery in the same controller.
POMDP formulation. State s_t = [o^H_t, P_t] where o^H_t is a temporal stack of proprioception + velocity command and P_t is privileged information (base velocity, terrain height, external disturbance, ZMP features) inferred at deployment via a VAE estimator P(e_t, z_t | o^H_t) following DreamWaQ (Nahrendra et al. 2023). Action a_t = target joint angles fed to a 100 Hz PD controller.
Constrained-RL formulation (Eq. 2). Replace fixed penalty rewards with two explicit constraints:
- Distributional divergence
D_f(ρ_π || ρ_π^E) ≤ ε_fbetween learned and expert (CMU MoCap retargeted) occupancy measures, implemented as Wasserstein-1 under Kantorovich-Rubinstein duality with a discriminator trained via WGAN-GP-style gradient penalty. - Feasibility constraint
E[ϕ(τ)] ≤ ε_ϕwhere ϕ is a ZMP-based indicator.
Robust extension (Eq. 4). Worst-case feasibility under a mismatched-transition uncertainty set P_α^L = {αP_T^L + (1−α)P̄_T}; reward maximization stays on the learning dynamics P_T^L. This decouples "be safe under all dynamics" from "track well in expected dynamics".
Three-stage hierarchical training.
-
π_1 goal-tracking (Sec. 4.1): PPO + alternating discriminator updates, optimizing
r_T − λ f_d(s^d)with Lagrange-styleλ. -
π_2 safety recovery (Sec. 4.2): trained under the extreme-case uncertainty set — multi-scale external forces (up to 200 N / 200 N·m), high-intensity proprio + PD-gain noise, malicious velocity-command resampling, and aggressive domain randomization. ZMP feasibility:
ϕ(s,a) = ‖p_ZMP − p_ac‖_2withp_ZMP = p_CoM − (z_CoM / g) · p̈_CoM. Frequency encoding (NeRF-style) is applied to ϕ to expose subtle stability variations. -
π_0 high-level planner (Sec. 4.3): discrete
ā_t ∈ {0,1}^2, learned via Double-DQN with ε-greedy, rewardr_T(s_t, ā_t) − 1(ā_{t−1}≠ā_t) − α·1(s_t)where the switch penalty discourages chattering andαtrades off task vs. safety. Authors reportα ∈ {0, 20, 50}are stable;α = 200destabilizes training.
Effectiveness — locomotion across terrains (Isaac Gym; 1200 steps = 12 s; ±std over 3 seeds; Table 1):
| Method | Slopes Low SR | Slopes High SR | Stairs Low SR | Stairs High SR |
|---|---|---|---|---|
| DreamWaQ | 92.31 | 90.46 | 74.32 | 60.58 |
| AHL | 98.83 | 97.36 | 93.73 | 67.48 |
| Goal-tracking only | 99.90 | 98.51 | 96.60 | 72.60 |
| HWC-Loco-l (low α) | 100.00 | 99.95 | 99.80 | 78.92 |
| HWC-Loco | 100.00 | 100.00 | 99.98 | 84.34 |
The headline number: on high-speed stairs (the only setting where Goal-tracking-only fails substantially), adding the safety-recovery hierarchy lifts SR from 72.60% → 84.34%.
Robustness — disturbances (Table 2, success rate %; all entries are SR, higher is better. The paper's Table 2 reports success rate only — it has no ZMP-deviation column):
| Policy | Ext F Low-freq | Ext F Constant | Impulse Low | Impulse High | Payload Low (0–5 kg) | Payload High (0–10 kg) |
|---|---|---|---|---|---|---|
| DreamWaQ | 85.92 | 51.31 | 85.24 | 45.34 | 67.63 | 55.04 |
| AHL | 87.15 | 60.72 | 85.87 | 62.94 | 79.29 | 59.16 |
| Goal-tracking | 90.00 | 61.20 | 88.90 | 58.13 | 78.34 | 61.44 |
| HWC-Loco-l | 92.69 | 68.60 | 92.79 | 77.09 | 84.11 | 63.96 |
| HWC-Loco | 95.88 | 75.95 | 94.84 | 81.27 | 87.43 | 69.86 |
Scalability. Cross-embodiment Unitree G1 (Table 3) actually outperforms the H1 platform: 98.14% vs. 97.13% SR. Expressive motion tracking (Table 4) under impulse disturbance: Punching 94.01% / Dancing 86.44% / Expressive Walking 94.53% — beats a domain-randomized motion-tracking baseline by ~4 pts each.
Real-world deployment. Climbs 15 cm stairs and 20° slopes; handles pushes, pulls, kicks via automatic policy switching (Figs. 10–13 in appendix B.6). Tested outdoors on flat, grass, slopes.
Strong DR baseline (Large-DR-Hist, Table 19). Even with 4× randomization scale, a non-hierarchical history-aware policy reaches only 70.53% on Constant disturbances vs. HWC-Loco's 75.95%; gap is largest on high-impulse (71.36% vs. 81.27%). Confirms that the hierarchy adds value beyond DR scaling.
Switching strategies (Tables 22–25). Vs. fixed ZMP thresholds:
| Method | Low-Freq SR | Constant SR | Low-Imp SR | High-Imp SR | Switch count (Constant) |
|---|---|---|---|---|---|
| Fixed-0.2 ZMP | 94.83 | 70.71 | 95.04 | 80.31 | 459 |
| Fixed-0.4 ZMP | 91.54 | 64.86 | 91.81 | 75.84 | 206 |
| HWC-Loco-l | 92.69 | 68.60 | 92.79 | 77.09 | 45 |
| HWC-Loco (learned) | 95.88 | 75.95 | 94.84 | 81.27 | 195 |
Fixed thresholds over-trigger (~450 switches) and degrade tracking; the learned planner trades fewer switches for higher SR.
Robust-optimization ablation (Table 26). Removing the extreme-case uncertainty set (ZMP constraint only) drops high-impulse SR from 81.27% → 76.29%, confirming adversarial training is essential beyond the structural ZMP feasibility.
Non-robust constrained baselines (Table 27). CRL and RL-Penalty match success rate but degrade human-likeness (3.56 / 3.45 vs. 3.11) and tracking (1.06 / 1.11 vs. 1.12).
Hyperparameter sensitivity (Tables 17, 18). HWC-Loco is insensitive to mismatch scale α and imitation weight λ within ±2× of the nominal value: SR moves only 0.06–0.13 pts.
VAE estimation noise (Tables 20, 21). ZMP-feature MSE stays below 0.075 across disturbance regimes. Injecting Gaussian noise σ ≤ 1.0 on top of VAE estimates causes only mild SR degradation (78.65% → 76.99%) and modest flip-count increase (80 → 96); σ = 2.0 (well beyond observed VAE error) crashes performance to 21.81%.
- Policy switching is discrete and low-level controllers are frozen during high-level training. Joint optimization of the hierarchy could smooth transitions.
- The deployed humanoid has only 19 DOF, limiting whole-body coordination and recovery expressiveness.
- Recovery policy is trained in simulation only — extreme real-world disturbances may not be fully covered. Adversarial real-world data could improve coverage.
HWC-Loco makes safety recovery a structural component of the policy rather than a post-hoc safety filter or a single conservative max-min RL objective. The key empirical insight is that fixed-threshold ZMP heuristics over-switch — they treat any margin violation equally — while a learned switch uses temporal context to invoke recovery only when the current state is genuinely unrecoverable by π_1.
Compared to:
- BFM-Zero — a promptable foundation model; HWC-Loco is task-specific but adds a robust-optimization axis.
- Lang-To-Loco — language-conditioned humanoid control; orthogonal to HWC-Loco's safety focus.
- DreamWaQ / AHL — strong domain-randomization baselines that HWC-Loco beats by 14–22 pts on the hardest disturbance settings.
The framework is foundational for safety-critical loco-manipulation and is the natural pairing for promptable upper-body controllers; future work targets integrating HWC-Loco with manipulation skills.
- OpenReview: https://openreview.net/forum?id=3UE3Aatcjy
- BFM-Zero
- Lang-To-Loco
- WholeBodyVLA
- Sim2Real-VLA — sibling CUHK Shenzhen / DexForce work
- VLBiMan — sibling CUHK Shenzhen / DexForce work
← Back to ICLR-2026