ICLR 2026 BFM Zero - Heungwoo/research GitHub Wiki

BFM-Zero — Promptable Behavioral Foundation Model for Humanoids

Venue: ICLR 2026 Authors: Carnegie Mellon University + Meta (Yitang Li, Zhengyi Luo, Tonghe Zhang, Cunxi Dai, Andrea Tirinzoni, Anssi Kanervisto, Haoyang Weng, Kris Kitani, Mateusz Guzek, Ahmed Touati, Alessandro Lazaric, Matteo Pirotta, Guanya Shi). Two authors (Yitang Li, Haoyang Weng) interned at CMU and are now with Tsinghua University. Category: Humanoid Control — Behavioral Foundation Model Trend tag: Unsupervised RL · Forward-Backward representations · sim-to-real

Approach diagram

flowchart LR
  Mocap[LAFAN1 mocap dataset<br/>40 motions, retargeted to G1] --> Disc[GAN-style discriminator D<br/>style regularization]
  Sim[IsaacLab @ 200 Hz<br/>1024 parallel envs<br/>DR + perturbations] --> Train[Off-policy actor-critic]
  Disc --> Train
  Train --> FB[Forward-Backward representation<br/>F: S×A×Z→Z<br/>B: S→Z<br/>z ∈ R^256]
  Train --> Pol["History-conditioned policy<br/>π(o_t,H, z), H=4"]
  FB --> Z["Latent task space Z<br/>r_z = φ(s)^T z"]
  Z --> Track["Zero-shot motion tracking<br/>z_t = Σ B(s_t prime)"]
  Z --> Goal["Goal reaching<br/>z_g = B(s_g)"]
  Z --> Rew["Reward inference<br/>z_r = E(B(s)·r(s))"]
  Z --> Few[Few-shot CEM / DIAL-MPC<br/>in latent space]
  Pol --> G1[Real Unitree G1 @ 50 Hz<br/>29 DoF]
Loading

Problem

Humanoid whole-body control is dominated by per-task on-policy PPO with hand-tuned tracking rewards (HumanPlus, ExBody, ASAP, HOVER, OmniH2O…). These are task-specific, non-adaptive, and lack a unified prompt interface for goal/motion/reward specification. Behavioral Foundation Models (BFMs) trained with unsupervised RL — proven on virtual characters (Tessler '23, Tirinzoni '25 FB-CPR) — have never been deployed on a real humanoid. Two unknowns: (i) whether off-policy unsupervised RL can scale to real humanoids, and (ii) whether the resulting latent space survives sim-to-real.

Detailed Method

Backbone: FB-CPR + sim-to-real extensions

BFM-Zero builds on FB-CPR (Tirinzoni et al., 2025), which combines Forward-Backward (FB) representations (Touati & Ollivier 2021) with online training and motion-data regularization.

The FB framework learns:

  • Forward map F: S × A × R^d → R^d
  • Backward map B: S → R^d
  • such that the discounted state-visitation measure decomposes as M^{π_z}(ds' | s, a) ≃ F(s, a, z)^T B(s') ρ(ds').
  • This implies F(s, a, z)^T z is the Q-function of π_z under reward r_z(s) = φ(s)^T z, where φ = (E[BB^T])^{-1} B.

The latent z ∈ R^256 simultaneously parameterizes (i) reward functions, (ii) goal embeddings (z_g = B(s_g)), and (iii) imitation embeddings z_τ = (1/|τ|) Σ_{(o,s)∈τ} B(o, s).

Sim-to-real design choices (the paper's contribution beyond FB-CPR)

A) Asymmetric history training. Policy sees history o_{t,H} = {o_{t-H}, a_{t-H}, …, o_t} ∈ R^{93·H+64}; critics see (o_{t,H}, s_t) including privileged info s ∈ R^463 (root height, body pose/rotation, lin/ang velocities). Closes proprio-vs-privileged gap. History length H = 4.

B) Massively parallel off-policy. N_env = 1024 parallel envs (FastTD3-style scaling), batch size 1024, UTD = 16 grad steps per env step, 3M total grad updates ≈ 192M env steps, replay buffer ~5M transitions. Episode length T = 500.

C) Domain randomization. COM offset U(±0.02 m), link mass U(0.95, 1.05), friction U(-0.5, 1.25), default joint pos U(±0.02 m), random push U(0, 0.5 m/s). Observation noise on joint pos/vel/gravity/angvel.

D) Reward regularization. 6 auxiliary penalties: DoF-limit (-10), action rate (-0.1), self-contact (-1), feet orientation (-0.4), ankle roll (-4), feet slip (-2). Optimized through a separate auxiliary critic Q_R.

Loss structure

Three coupled objectives:

  • FB loss (Bellman residual on successor measures): L(F, B) trains F, B jointly via temporal-difference on (F(o,s,a,z)^T B(o',s'))² minus a cross term.
  • Discriminator loss (GAN-style on motion data M): L(D) = -E_{τ∼M}[log D] - E_D[log(1-D)], yielding an imitation reward r_d = D / (1-D).
  • Auxiliary critic Q_R (Bellman residual with the 6 penalty rewards above).

Combined actor loss: L(π) = -E[F(o,s,a,z)^T z + λ_D Q_D(o,s,a,z) + λ_R Q_R(o,s,a,z)] with λ_D (αD) = 0.05, λ_R (αR) = 0.02.

Network architecture (Table 2, ~440M total params)

  • Critics F, Q_D, Q_R: 4 embedding residual blocks + 6 residual blocks (2048 hidden, Mish), feed-forward 2048×1, 2 parallel networks. F has output dim d=256, Q_D/Q_R output 1.
  • Actor π: 4+6 residual blocks, 2048 hidden, output 29 (PD targets).
  • Discriminator D: 2-layer FF, 1024 hidden, ReLU.
  • Backward map B: 1-layer FF, 256 hidden, ReLU, output normalized to unit ball.
  • Param counts: F = 135.8M, Q_D = Q_R = 134.8M, π = 31.9M, D = 2.9M, B = 0.2M → 440.5M total.

Inference modes (zero-shot)

  • Reward: z_r = (1/N) Σ r(s_i) B(s_i), 400k samples from replay buffer.
  • Goal: z_g = B(s_g).
  • Tracking: rolling z_t = Σ_{t'=t}^{t+H} B(s_{t'}), look-ahead H = 3 (real) or = sequence length (sim).

Few-shot adaptation in latent space

  • Single-pose: CEM optimization starting from z_init = B(s_g, o_g) with sparse reward 1{h_right_foot > 0.15 ∧ no-contact}.
  • Trajectory: dual-loop annealing trajectory optimization (DIAL-MPC style, Xue '25), particles N=2048, β₁=0.85, β₂=0.9, M=6 iterations.

Training setup

  • Simulator: IsaacLab @ 200 Hz, control rate 50 Hz.
  • Robot: Unitree G1, 29 DoF, action ∈ R^29 (PD targets).
  • Mocap data: LAFAN1 (40 motions, retargeted to G1, split into 10s chunks) with EMD-based prioritized sampling: p(m) ∝ 2^{max(0.5, min(EMD(m),2)·4)}.
  • Initial-state distribution: 30% falling-pose / 70% mocap-frame initialization.
  • Discount γ = 0.98; sequence length 8; orthonormality loss coefficient 100; gradient penalty 10. LRs: F = π = Q_D = Q_R = 3e-4, B = D = 1e-5.
  • Generality: also demonstrated on a Booster T1 humanoid in appendix.

Comprehensive Results

Sim performance (Fig. 3) — tracking error E_mpjpe (↓), reward (↑), pose error (↓)

Model Test env Test data Track ↓ Reward ↑ Pose ↓
BFM-Zero-priv (no DR, privileged) Isaac no-DR LAFAN1 1.0749 299.3 1.0291
BFM-Zero Isaac DR LAFAN1 1.1015 221.9 1.1387
BFM-Zero Mujoco DR LAFAN1 1.0789 207.3 1.1041
BFM-Zero Mujoco DR AMASS (OOD) 1.0342 207.3 1.4735

DR drops vs the privileged idealization: tracking -2.47%, reward -25.86%, pose-reaching -10.65%. Sim-to-sim (Isaac→Mujoco) gap < 7%. Reward inference suffers most under DR — paper attributes this to subsampled-replay-buffer brittleness on sparse rewards (e.g. move-ego-0.0).

Real-world Unitree G1 (qualitative)

All from a single trained model:

  • Tracking of walking, dancing, fighting, sports motions; even retargeted monocular YouTube clips with occlusions/artifacts.
  • Goal reaching: discontinuous goal-pose sequences executed with smooth natural transitions (incl. infeasible "in-the-air" target → robot converges to nearest feasible pose).
  • Reward optimization: locomotion rewards, arm-movement (wrist height), pelvis-height (sit/crouch). Composites work via simple linear reward combination (e.g. raisearm + move-backward).
  • Disturbance rejection: kicks, pushes, drags-to-ground all recovered with running-style adaptive recovery — without any explicit recovery training.

Few-shot adaptation (real)

  • Single-pose with 4 kg attached payload: zero-shot prompt z_init collides within 5 s; CEM-optimized z* keeps single-leg balance > 15 s.
  • Trajectory (leaping under altered ground friction): dual-annealing in latent space reduces tracking error by ~29.1%.

Ablation Studies (Appendix D)

The paper's quantitative ablations cover data size and model size (Fig. 13, Table 3) and reward-inference data source (Fig. 14). There is no latent-dimension or auxiliary-reward on/off ablation in the paper; d_z = 256 and the six regularization rewards are fixed design choices (Table 1 / Fig. 11).

Data size & model size (Fig. 13, Table 3)

Trained on LAFAN1 alone, plus LAFAN1+AMASS (CMU & BMLHandball subsets) mixes at X% ∈ {12.5, 25, 50, 75, 100}, across ResNet (3/6/9 blocks, 1024/2048 dim) and MLP (2/4 layers, 1024/2048 dim). Tracking improves with total model capacity for almost all training mocaps (LAFAN1 saturates early, being close to the held-out CMU/BMLHandball test motions). Residual architectures outperform MLPs and scale to larger sizes; MLPs become unstable when scaled up. Total params for the configurations span 40.1M → 679.9M; the chosen ResNet⋆ (6-block, 2048 dim) is 440.5M.

Per-component params for the ablated configurations (Table 3; F / Q_R columns):

Model π Q_R F Total
MLP 2-layer 1024d 4.4M 10.7M 11.2M 40.1M
MLP 2-layer 2048d 15.1M 34.0M 35.0M 121.2M
MLP 4-layer 1024d 6.5M 14.9M 15.4M 54.8M
MLP 4-layer 2048d 23.5M 50.8M 51.8M 179.9M
ResNet 3-block 2048d 19.3M 59.2M 60.3M 201.1M
ResNet⋆ 6-block 2048d (chosen) 31.9M 134.8M 135.9M 440.5M
ResNet 9-block 2048d 44.5M 210.4M 211.5M 679.9M

Reward inference data source (Fig. 14)

Reward inference using the online replay buffer vs the training motion set: both lower average reward-inference performance vs the default (400k replay-buffer subsample), with the larger drop when using the full training buffer — attributed to buffer samples being collected under domain randomization while the motion set is not. Some sparse rewards (e.g. move-ego-0.0) show occasional collapse/repetition under subsampling.

Limitations (as stated by authors)

  1. Scope tied to motion dataset. Behaviors are bounded by the diversity of LAFAN1; no scaling-law study yet linking dataset size / sim data / architecture / model performance.
  2. Sim-to-real gap not closed. Despite DR + history + asymmetric learning, complex movements still need better online adaptation algorithms.
  3. Test-time adaptation underexplored. Few-shot CEM / DIAL-MPC works but a thorough fast-adaptation / fine-tuning study is missing.

Significance & Positioning

Vs on-policy humanoid RL (HumanPlus, ExBody, ASAP, HOVER, OmniH2O, GMT, Lipschitz-locomotion). All train per-task with PPO + hand-engineered tracking rewards, then add specialized recovery / extension modules. BFM-Zero replaces all of that with one off-policy unsupervised pretrain — a single 440M-param model handles tracking + goal reaching + reward optimization + disturbance recovery + few-shot adaptation with no per-task training.

Vs FB-CPR / behavior FMs on virtual characters (Tirinzoni '25, Tessler '23 CALM, Peng '22 ASE). Those works showed BFMs only on simulated MuJoCo characters. BFM-Zero is the first deployment of an unsupervised-RL BFM on a real humanoid, with the sim-to-real recipe (asymmetric history, DR, aux-reward shaping) as the core contribution.

Vs VLA-style humanoid foundation models (GR00T N1, Helix). GR00T uses behavior cloning on teleoperation; BFM-Zero argues humanoid whole-body control lacks the actuator-level demonstration data manipulation has, making BC-style VLA a fundamental mismatch. The forward-backward latent z is to humanoids what action tokens are to manipulation VLAs — but unsupervised and prompt-able by reward, goal, or motion.

Complement to instruction-conditioned humanoid policies like Lang-To-Loco: BFM-Zero gives the substrate (smooth promptable latent space), language conditioning could sit on top.

A counterweight to the "scale teleop data" hypothesis in humanoid VLAs: shows that purely-physical pretraining + unsupervised RL gives you most of the prompt-ability for free. SLERP interpolation between latents produces meaningful intermediate skills with zero retraining.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️