ICLR 2026 Pretrain Finetune Humanoid - Heungwoo/research GitHub Wiki

LIFT — Bridging large-scale pretraining and efficient finetuning for humanoid control

Venue: ICLR 2026 · Authors: Weidong Huang, Zhehan Li, Hangxin Liu, Biao Hou, Yao Su, Jingwen Zhang (BIGAI State Key Lab of General AI; Xidian University) · arXiv:2601.21363 · Category: Humanoid Control · Trend tag: Pretrain-then-finetune for locomotion

Approach diagram

flowchart LR
  subgraph Pretrain
    SAC[Off-policy SAC<br/>large batch, high UTD] --> Policy[Pretrained<br/>locomotion policy]
  end
  Policy --> Zero[Zero-shot real<br/>deployment]
  Policy --> WM[Physics-informed<br/>world model pretrain]
  subgraph Finetune
    Det[Deterministic policy<br/>collects real data] --> WM2[Stochastic exploration<br/>inside world model]
    WM2 --> FT[Efficient policy +<br/>world model finetune]
  end
  WM --> Finetune
Loading

Problem

Humanoid control faces a tension between scale and adaptation. On-policy RL (e.g. PPO) is the workhorse for large-scale pretraining but is sample-inefficient, which makes safe adaptation to a new environment expensive and risky — random exploration on real hardware is dangerous. The goal is to bridge the gap: pretrain at scale, then finetune efficiently and safely.

Method

LIFT (Large-scale pretraIning and efficient FineTuning) is a three-stage framework:

  1. Large-scale policy pretraining. The paper shows that off-policy Soft Actor-Critic (SAC) — with large-batch updates and a high Update-To-Data (UTD) ratio — reliably supports large-scale humanoid locomotion pretraining, contrary to the usual reliance on on-policy methods, and yields policies deployable zero-shot on real robots.
  2. Physics-informed world model pretraining. A world model is pretrained alongside the policy, grounded in physics priors.
  3. Efficient finetuning. In a new environment, data collection executes a deterministic policy (safe), while stochastic exploration is confined to the physics-informed world model rather than the real robot. This brings model-based sample efficiency to adaptation while mitigating the risk of random real-world exploration.

Results

  • SAC-based pretraining with large-batch / high-UTD updates achieves zero-shot real-robot deployment of humanoid locomotion policies.
  • Model-based finetuning adapts pretrained policies to new environments with improved sample efficiency and reduced exploration risk versus on-policy adaptation.

(Specific sample-efficiency and success-rate numbers are reported in the full paper; code at bigai-ai/LIFT-humanoid.)

Significance

Challenges the default that humanoid locomotion must be pretrained with on-policy RL: a well-tuned off-policy SAC scales and transfers zero-shot. By confining risky stochastic exploration to a physics-informed world model during finetuning, LIFT offers a safe, sample-efficient adaptation recipe — a practical pretrain-then-finetune pipeline for real humanoids.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️