ICLR 2026 Pretrain Finetune Humanoid - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 · Authors: Weidong Huang, Zhehan Li, Hangxin Liu, Biao Hou, Yao Su, Jingwen Zhang (BIGAI State Key Lab of General AI; Xidian University) · arXiv:2601.21363 · Category: Humanoid Control · Trend tag: Pretrain-then-finetune for locomotion
flowchart LR
subgraph Pretrain
SAC[Off-policy SAC<br/>large batch, high UTD] --> Policy[Pretrained<br/>locomotion policy]
end
Policy --> Zero[Zero-shot real<br/>deployment]
Policy --> WM[Physics-informed<br/>world model pretrain]
subgraph Finetune
Det[Deterministic policy<br/>collects real data] --> WM2[Stochastic exploration<br/>inside world model]
WM2 --> FT[Efficient policy +<br/>world model finetune]
end
WM --> Finetune
Humanoid control faces a tension between scale and adaptation. On-policy RL (e.g. PPO) is the workhorse for large-scale pretraining but is sample-inefficient, which makes safe adaptation to a new environment expensive and risky — random exploration on real hardware is dangerous. The goal is to bridge the gap: pretrain at scale, then finetune efficiently and safely.
LIFT (Large-scale pretraIning and efficient FineTuning) is a three-stage framework:
- Large-scale policy pretraining. The paper shows that off-policy Soft Actor-Critic (SAC) — with large-batch updates and a high Update-To-Data (UTD) ratio — reliably supports large-scale humanoid locomotion pretraining, contrary to the usual reliance on on-policy methods, and yields policies deployable zero-shot on real robots.
- Physics-informed world model pretraining. A world model is pretrained alongside the policy, grounded in physics priors.
- Efficient finetuning. In a new environment, data collection executes a deterministic policy (safe), while stochastic exploration is confined to the physics-informed world model rather than the real robot. This brings model-based sample efficiency to adaptation while mitigating the risk of random real-world exploration.
- SAC-based pretraining with large-batch / high-UTD updates achieves zero-shot real-robot deployment of humanoid locomotion policies.
- Model-based finetuning adapts pretrained policies to new environments with improved sample efficiency and reduced exploration risk versus on-policy adaptation.
(Specific sample-efficiency and success-rate numbers are reported in the full paper; code at bigai-ai/LIFT-humanoid.)
Challenges the default that humanoid locomotion must be pretrained with on-policy RL: a well-tuned off-policy SAC scales and transfers zero-shot. By confining risky stochastic exploration to a physics-informed world model during finetuning, LIFT offers a safe, sample-efficient adaptation recipe — a practical pretrain-then-finetune pipeline for real humanoids.
← Back to ICLR-2026