CVPR 2026 Humanoid GPT - Heungwoo/research GitHub Wiki
Venue: CVPR 2026 (Poster #39399) Category: Humanoid Motion / Pretraining Trend tag: Trend 2 Affiliations: Tsinghua University · Galbot · Beihang University · Shanghai Jiao Tong University · Peking University · Shanghai Qi Zhi Institute
flowchart LR
MOCAP["2B-frame retargeted mocap corpus<br/>AMASS · LAFAN1 · Motion-X++ · PHUMA · MotionMillion + in-house"] --> HME["HME diversity clustering"]
HME --> EXP["per-cluster RL motion experts"]
EXP --> DAGGER["DAgger distillation"]
DAGGER --> GPT["single causal-attention GPT<br/>generalist tracker"]
TARGET["target trajectory"] --> GPT
GPT --> ACT["zero-shot humanoid action sequence"]
Humanoid motion-tracking policies have always been task-specialist: train per task, retrain for each new motion. The agility-vs-generalization trade-off is fundamental in conventional designs.
Humanoid-GPT is not a single GPT trained directly on raw mocap. The pipeline is expert-distillation:
- Corpus + diversity metric. A 2 B-frame retargeted corpus unifies all major mocap datasets (AMASS, LAFAN1, Motion-X++, PHUMA, MotionMillion) with large-scale in-house recordings — ~4–5× larger log-volume than AMASS under the authors' Harmonic Motion Embedding (HME), a novel metric that quantifies and categorizes motion diversity directly from motion data.
- Per-cluster RL experts. Motion experts are trained via reinforcement learning on HME-clustered data.
- DAgger distillation. A GPT-style Transformer with causal temporal attention is trained via DAgger to consolidate all expert controllers into one generalist tracker.
At deployment, conditioning on a target trajectory yields a zero-shot motion tracker that runs on a Unitree G1.
- Zero-shot generalization to unseen motions and control tasks (e.g., unseen dance sequences) while still tracking highly dynamic in-domain motions — directly attacking the agility-vs-generalization trade-off.
- Clear scaling laws: enlarging both corpus and model capacity yields consistent gains in tracking accuracy and stability.
- Latency: end-to-end inference under 1.5 ms on a single NVIDIA RTX 4090 (ONNX + TensorRT), ~5× faster than TWIST.
Humanoid-GPT shows that a single causal Transformer, distilled from many per-cluster RL experts over a billion-scale, diversity-balanced corpus, can beat per-task specialist trackers and generalize zero-shot. The contribution is two-fold: (i) the HME-driven data pipeline that makes "scale" meaningful by measuring and balancing motion diversity, and (ii) the DAgger distillation that turns a fleet of RL experts into one generalist. Its lineage is the whole-body motion-tracking line (it benchmarks against TWIST), not VLA manipulation.
- Project page: https://qizekun.github.io/humanoid-gpt/
- Code: https://github.com/qizekun/Humanoid-GPT
- CVPR Poster: #39399 (ExHall F, Sat Jun 6 2026)
- Authors: Zekun Qi, Xuchuan Chen, Jilong Wang, Chenghuai Lin, Yunrui Lian, Zhikai Zhang, Yu Guan, Wenyao Zhang, Xinqiang Yu, He Wang, Li Yi
← Back to CVPR-2026