ICLR 2026 HWC Loco - Heungwoo/research GitHub Wiki

HWC-Loco — Hierarchical Whole-Body Control for Robust Humanoid Locomotion

Venue: ICLR 2026 Authors: Sixu Lin, Guanren Qiao, Yunxin Tai, Ang Li, Kui Jia, Guiliang Liu (CUHK Shenzhen; Harbin Institute of Technology; DexForce Technology; Southeast University — corresponding: Guiliang Liu) Category: Humanoid Control — Locomotion Trend tag: Robust RL / safety-critical recovery

Approach diagram

flowchart LR
  Cmd[Velocity command v_x, v_y, w_z] --> Pi0[High-level π_0<br/>discrete switch policy<br/>Double-DQN, ε-greedy]
  Obs[oH_t = ot..ot-H<br/>+ VAE-estimated P_t<br/>velocity + ZMP features] --> Pi0
  Pi0 -->|goal-tracking| Pi1[π_1: PPO + Wasserstein imitation<br/>vs. CMU MoCap retargeted]
  Pi0 -->|safety recovery| Pi2[π_2: robust optim. over P_α^L<br/>+ ZMP constraint]
  Pi1 --> PD[PD controller 100 Hz<br/>target joint angles]
  Pi2 --> PD
  PD --> Robot[Unitree H1 / G1]
Loading

Problem

Reinforcement-learning humanoid locomotion policies trained in simulation suffer two failure modes: (1) Sim2Real transition mismatch causes falls under realistic disturbances (slips, pushes, terrain edges, payload changes); (2) standard robust-RL formulations (max-min over worst-case dynamics) produce overly conservative policies that fail to track velocity commands. HWC-Loco asks how to keep aggressive goal-tracking and guaranteed safety recovery in the same controller.

Detailed Method

POMDP formulation. State s_t = [o^H_t, P_t] where o^H_t is a temporal stack of proprioception + velocity command and P_t is privileged information (base velocity, terrain height, external disturbance, ZMP features) inferred at deployment via a VAE estimator P(e_t, z_t | o^H_t) following DreamWaQ (Nahrendra et al. 2023). Action a_t = target joint angles fed to a 100 Hz PD controller.

Constrained-RL formulation (Eq. 2). Replace fixed penalty rewards with two explicit constraints:

  1. Distributional divergence D_f(ρ_π || ρ_π^E) ≤ ε_f between learned and expert (CMU MoCap retargeted) occupancy measures, implemented as Wasserstein-1 under Kantorovich-Rubinstein duality with a discriminator trained via WGAN-GP-style gradient penalty.
  2. Feasibility constraint E[ϕ(τ)] ≤ ε_ϕ where ϕ is a ZMP-based indicator.

Robust extension (Eq. 4). Worst-case feasibility under a mismatched-transition uncertainty set P_α^L = {αP_T^L + (1−α)P̄_T}; reward maximization stays on the learning dynamics P_T^L. This decouples "be safe under all dynamics" from "track well in expected dynamics".

Three-stage hierarchical training.

  • π_1 goal-tracking (Sec. 4.1): PPO + alternating discriminator updates, optimizing r_T − λ f_d(s^d) with Lagrange-style λ.
  • π_2 safety recovery (Sec. 4.2): trained under the extreme-case uncertainty set — multi-scale external forces (up to 200 N / 200 N·m), high-intensity proprio + PD-gain noise, malicious velocity-command resampling, and aggressive domain randomization. ZMP feasibility: ϕ(s,a) = ‖p_ZMP − p_ac‖_2 with p_ZMP = p_CoM − (z_CoM / g) · p̈_CoM. Frequency encoding (NeRF-style) is applied to ϕ to expose subtle stability variations.
  • π_0 high-level planner (Sec. 4.3): discrete ā_t ∈ {0,1}^2, learned via Double-DQN with ε-greedy, reward r_T(s_t, ā_t) − 1(ā_{t−1}≠ā_t) − α·1(s_t) where the switch penalty discourages chattering and α trades off task vs. safety. Authors report α ∈ {0, 20, 50} are stable; α = 200 destabilizes training.

Comprehensive Results

Effectiveness — locomotion across terrains (Isaac Gym; 1200 steps = 12 s; ±std over 3 seeds; Table 1):

Method Slopes Low SR Slopes High SR Stairs Low SR Stairs High SR
DreamWaQ 92.31 90.46 74.32 60.58
AHL 98.83 97.36 93.73 67.48
Goal-tracking only 99.90 98.51 96.60 72.60
HWC-Loco-l (low α) 100.00 99.95 99.80 78.92
HWC-Loco 100.00 100.00 99.98 84.34

The headline number: on high-speed stairs (the only setting where Goal-tracking-only fails substantially), adding the safety-recovery hierarchy lifts SR from 72.60% → 84.34%.

Robustness — disturbances (Table 2, success rate %; all entries are SR, higher is better. The paper's Table 2 reports success rate only — it has no ZMP-deviation column):

Policy Ext F Low-freq Ext F Constant Impulse Low Impulse High Payload Low (0–5 kg) Payload High (0–10 kg)
DreamWaQ 85.92 51.31 85.24 45.34 67.63 55.04
AHL 87.15 60.72 85.87 62.94 79.29 59.16
Goal-tracking 90.00 61.20 88.90 58.13 78.34 61.44
HWC-Loco-l 92.69 68.60 92.79 77.09 84.11 63.96
HWC-Loco 95.88 75.95 94.84 81.27 87.43 69.86

Scalability. Cross-embodiment Unitree G1 (Table 3) actually outperforms the H1 platform: 98.14% vs. 97.13% SR. Expressive motion tracking (Table 4) under impulse disturbance: Punching 94.01% / Dancing 86.44% / Expressive Walking 94.53% — beats a domain-randomized motion-tracking baseline by ~4 pts each.

Real-world deployment. Climbs 15 cm stairs and 20° slopes; handles pushes, pulls, kicks via automatic policy switching (Figs. 10–13 in appendix B.6). Tested outdoors on flat, grass, slopes.

Ablation Studies

Strong DR baseline (Large-DR-Hist, Table 19). Even with 4× randomization scale, a non-hierarchical history-aware policy reaches only 70.53% on Constant disturbances vs. HWC-Loco's 75.95%; gap is largest on high-impulse (71.36% vs. 81.27%). Confirms that the hierarchy adds value beyond DR scaling.

Switching strategies (Tables 22–25). Vs. fixed ZMP thresholds:

Method Low-Freq SR Constant SR Low-Imp SR High-Imp SR Switch count (Constant)
Fixed-0.2 ZMP 94.83 70.71 95.04 80.31 459
Fixed-0.4 ZMP 91.54 64.86 91.81 75.84 206
HWC-Loco-l 92.69 68.60 92.79 77.09 45
HWC-Loco (learned) 95.88 75.95 94.84 81.27 195

Fixed thresholds over-trigger (~450 switches) and degrade tracking; the learned planner trades fewer switches for higher SR.

Robust-optimization ablation (Table 26). Removing the extreme-case uncertainty set (ZMP constraint only) drops high-impulse SR from 81.27% → 76.29%, confirming adversarial training is essential beyond the structural ZMP feasibility.

Non-robust constrained baselines (Table 27). CRL and RL-Penalty match success rate but degrade human-likeness (3.56 / 3.45 vs. 3.11) and tracking (1.06 / 1.11 vs. 1.12).

Hyperparameter sensitivity (Tables 17, 18). HWC-Loco is insensitive to mismatch scale α and imitation weight λ within ±2× of the nominal value: SR moves only 0.06–0.13 pts.

VAE estimation noise (Tables 20, 21). ZMP-feature MSE stays below 0.075 across disturbance regimes. Injecting Gaussian noise σ ≤ 1.0 on top of VAE estimates causes only mild SR degradation (78.65% → 76.99%) and modest flip-count increase (80 → 96); σ = 2.0 (well beyond observed VAE error) crashes performance to 21.81%.

Limitations stated by authors

  1. Policy switching is discrete and low-level controllers are frozen during high-level training. Joint optimization of the hierarchy could smooth transitions.
  2. The deployed humanoid has only 19 DOF, limiting whole-body coordination and recovery expressiveness.
  3. Recovery policy is trained in simulation only — extreme real-world disturbances may not be fully covered. Adversarial real-world data could improve coverage.

Significance & Positioning

HWC-Loco makes safety recovery a structural component of the policy rather than a post-hoc safety filter or a single conservative max-min RL objective. The key empirical insight is that fixed-threshold ZMP heuristics over-switch — they treat any margin violation equally — while a learned switch uses temporal context to invoke recovery only when the current state is genuinely unrecoverable by π_1.

Compared to:

  • BFM-Zero — a promptable foundation model; HWC-Loco is task-specific but adds a robust-optimization axis.
  • Lang-To-Loco — language-conditioned humanoid control; orthogonal to HWC-Loco's safety focus.
  • DreamWaQ / AHL — strong domain-randomization baselines that HWC-Loco beats by 14–22 pts on the hardest disturbance settings.

The framework is foundational for safety-critical loco-manipulation and is the natural pairing for promptable upper-body controllers; future work targets integrating HWC-Loco with manipulation skills.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️