Review Psi0 - Heungwoo/research GitHub Wiki
In-Depth Review β Ξ¨β (Psi-Zero): An Open Foundation Model Towards Universal Humanoid Loco-Manipulation
Paper: Ξ¨β: An Open Foundation Model Towards Universal Humanoid Loco-Manipulation Β· RSS 2026 (Sydney, oral β Humanoids session) Authors: Songlin Wei*, Hongyi Jing*, Boqian Li*, Zhenyu Zhao*, Jiageng Mao, β¦ Marco Pavone, Di Huang, Yue Wangβ β USC Physical Superintelligence (PSI) Lab Γ NVIDIA Γ WorldEngine arXiv: 2603.12263 (v1 Mar 12, 2026) Β· Project: psi-lab.ai/Psi0 Β· Code: github.com/physical-superintelligence-lab/Psi0 Openness: full ecosystem release stated β teleoperation stack, data/training pipeline, model weights (~2.5B), real-time inference engine
Companion reviews: RSS 2026 survey Β· Humanoid VLA Β· GR00T series Β· System 0/1/2 Β· Cross-Embodiment Β· EgoDex.
- The anti-scaling thesis for humanoid VLAs: don't co-train one monolithic policy on mixed human + humanoid data β the two action distributions are kinematically incompatible, so a single model wastes capacity reconciling them. Instead, decouple: pre-train the VLM on human video for semantics and visual-action representations, then post-train a separate action expert on humanoid data alone for joint-space control.
- The data-efficiency headline: with only ~800 h of egocentric human video (EgoDex) + ~30 h of real humanoid data (Humanoid Everyday), Ξ¨β beats baselines pre-trained on >10Γ the data β including GR00T N1.6 β by >40 percentage points in overall success rate on eight real-world long-horizon loco-manipulation tasks on a Unitree G1.
- Triple-system architecture mapping cleanly onto the System 0/1/2 framing: System-2 = Qwen3-VL-2B-Instruct backbone, System-1 = 500M flow-matching MM-DiT action expert (SD3-style dual modulation + joint attention β ablated to beat a naive DiT head), System-0 = the off-the-shelf AMO RL controller mapping 8-DoF base/torso commands to 15-DoF lower-body joints. 36-DoF policy action β 43-DoF whole-body control.
- A pre-training trick worth stealing: the VLM is trained to predict only a single next-step action (48-DoF task-space, FAST-tokenized to ~20 tokens with a custom-fit tokenizer) instead of action chunks β enough to learn task semantics and aligned visual representations at a fraction of the compute, since chunked control is delegated to the post-trained expert anyway. The ablation validates it: frozen stock Qwen3-VL scores 0β2/10; EgoDex next-action pre-training lifts it to 6/10 before any real-robot fine-tuning beyond the expert.
- Training-time Real-Time Chunking (RTC) rather than test-time guidance: inference-delay simulation during training (first d ~ U(0, d_max) action tokens un-noised and loss-masked) because test-time RTC guidance was unstable on their model β a practical negative result for the RTC line.
- It attacks the dominant humanoid-VLA recipe head-on. GR00T's data pyramid, EgoVLA, H-RDT, and Being-H0 all co-train across human and robot action distributions in one model. Ξ¨β argues (and shows on hardware) that this is the wrong use of human data: human video should shape the representation, not the action head. This is the humanoid-domain echo of VLM4VLA's finding that what matters is aligning the backbone's representations with control β and of LBM-style stage-decoupled recipes.
- "Scale the right data in the right way." The 800 h + 30 h result is the strongest data-efficiency claim in the humanoid space to date β a direct counter to the volume-first posture, and a challenge to GR00T N1.7's ~20k-hour EgoScale bet. Notably the quality bar is specific: high-quality egocentric manipulation video (EgoDex's Vision-Pro-tracked 829 h), not in-the-wild internet clips.
- Openness at the humanoid tier. GR00T is the only comparable open humanoid foundation model; Ξ¨β ships smaller (2.5B vs 3B), with the full teleop + training + deployment stack, from an academic lab. If the release holds, it becomes the default academic humanoid-VLA baseline.
- One more Qwen-backbone VLA built outside Alibaba β Qwen3-VL-2B joins GR00T (Qwen3-VL-2B) as external adopters while Qwen's own robot models stay closed.

Figure 1 of the Ξ¨β paper: a composite of Unitree G1 rollouts in the pantry evaluation environment β taking a cup from the coffee machine, pushing a cart, wiping the table, grasping a bottle into the sink, pushing the fridge door. The single scene captures the paper's task spectrum: dexterous-hand grasping, whole-body pushes, and locomotion between stations, all >2,000-step episodes at 30 Hz.
flowchart LR
subgraph S2["System-2 β Qwen3-VL-2B-Instruct (pre-trained on EgoDex)"]
OBS["Head camera (D435i) + instruction<br/>(+ 28-DoF arm/hand state at post-train)"]
end
S2 -- "hidden features z_t" --> S1
subgraph S1["System-1 β MM-DiT action expert (500M, flow matching)"]
MOD["Dual FiLM modulation by flow timestep Ο<br/>(VL branch and Action branch separately)"]
JA["Joint global attention over VL + action tokens"]
end
S1 --> ACT["36-DoF action chunk a_{t:t+H}:<br/>14 hand + 14 arm + torso rpy(3)<br/>+ base height + v_x, v_y, v_yaw + p_yaw"]
ACT -- "8-DoF base/torso commands" --> S0["System-0 β AMO RL tracking policy<br/>β 15-DoF legs + waist"]
ACT -- "28-DoF upper body direct" --> G1["Unitree G1 (29 DoF) + 2Γ Dex3-1 hands (7 DoF each)<br/>43-DoF whole-body control"]
S0 --> G1
classDef vlm fill:#bbdefb,stroke:#1565c0,color:#000
classDef act fill:#c8e6c9,stroke:#2e7d32,color:#000
classDef ctl fill:#fff9c4,stroke:#f57f17,color:#000
class S2 vlm
class S1,ACT act
class S0,G1 ctl
- MM-DiT vs naive DiT: the expert adapts Stable Diffusion 3's MM-DiT β the flow timestep modulates the vision-language and action branches separately (dual FiLM), and both token sets share joint global attention per block, versus a naive DiT's cross-attention into VL features. Ablated: MM-DiT consistently wins; the authors attribute naive DiT's weakness to its text-to-image conditioning heritage.
- Single head camera only; whole-body proprioception (upper joints, torso rpy, base height) as state. ~160 ms per forward pass at ~2.5B total parameters.
| Stage | Data | Trains | Objective | Compute |
|---|---|---|---|---|
| 1. VLM pre-training | EgoDex (829 h / ~900M frames, actions re-expressed in head-camera frame, 3Γ frame-rate upsampling, q01/q99 normalization, no state input) + Humanoid Everyday (31 h, 260 tasks, as visual-gap mitigation) | VLM (full) | Autoregressive single next-step action prediction in a unified 48-DoF task space ({9-DoF wrist pose + 5Γ3D fingertips} Γ 2 hands), FAST-tokenized (custom tokenizer: L1 0.005 vs 0.01 stock; 48 dims β ~20 tokens) | 64ΓA100, 10 days, LR 1e-4, batch 1024 |
| 2. Expert post-training | Humanoid Everyday (~3M frames real teleop) | MM-DiT from scratch, VLM frozen | Flow matching on 36-DoF joint-space chunks | 32ΓA100, ~30 h, batch 2048 |
| 3. Task fine-tuning | 80 teleop trajectories per task | Action expert only | Flow matching, 40k steps, cosine LR | (small) |
Plus training-time RTC: each prediction conditions on the previously committed chunk; the first d ~ U(0, d_max) tokens are exposed un-noised and masked from the loss, simulating inference delay. Test-time RTC guidance (gradient steering) was tried and found unstable on this model.
Teleoperation system (the data engine): PICO headset + wrist trackers β multi-target IK for arms/torso; MANUS gloves β full dexterous-hand DoF; waist/foot trackers β high-level locomotion commands to the AMO policy. Locomotion is deliberately decoupled from whole-body retargeting β trading expressiveness for lower-body stability and single-operator practicality (contrast TWIST2-style end-to-end whole-body teleop).
Setup: 8 real-world long-horizon tasks (pantry/household; 3β5 sub-tasks each; most >2,000 steps at 30 Hz; pick-place, pushing carts and fridge doors, wiping, faucet turning, chip-tray pulling), 10 rollouts per model per task; evaluators may intervene after a failed sub-task so later sub-tasks remain measurable; overall success = all sub-tasks completed.
- Headline: Ξ¨β's average overall success rate is β₯40 pp above the second-best baseline, GR00T N1.6 (3B, fine-tuned on the same 80-traj/task data with default recipe). Other baselines β Ο0.5 (action space padded 30β36), InternVLA-M1 (jittery chunks), H-RDT (2B DiT, imprecise), EgoVLA (weak lower-body priors), Diffusion Policy and ACT (fail broadly) β trail further. Several baselines score 0 on multiple tasks (task-wise chart bottoms out at 0 for most non-Ξ¨β policies on the harder tasks).
- Skill-level radar: strong across grasping, placing, pouring, rotating, squatting, walking, carrying, pushing, pulling β i.e., the wins are not concentrated in one skill family.
- Ablation (dual-arm 3-step task: right-arm P&P β left-arm P&P β dual-arm box lift):
| EgoDex pre-train | HE post-train | RTC | Head | Overall |
|---|---|---|---|---|
| β | β | β | naive DiT | 0/10 |
| β | β | β | MM-DiT | 2/10 |
| β | β | β | MM-DiT | 6/10 |
| β | β | β | MM-DiT | 8/10 |
| β | β | β | MM-DiT | 9/10 |
The decomposition is clean: human-video pre-training is worth +4/10, cross-task humanoid post-training +2, RTC +1 (with collision reduction and higher rollout throughput noted qualitatively). The frozen-stock-Qwen3-VL row (0β2/10) is the direct evidence that generic VLM features do not suffice β next-action pre-training is what aligns the representation, echoing VLM4VLA's action-supervision finding at the humanoid scale.
Genuinely novel / load-bearing: (a) the decoupled human-videoβVLM / robot-dataβexpert recipe with hardware-validated superiority over co-training; (b) single next-step action pre-training as a compute-cheap alignment objective; (c) MM-DiT as an action head (first humanoid VLA to ablate SD3-style dual-stream conditioning against naive DiT); (d) the 800 h + 30 h data-efficiency result itself.
Caveats a careful reader should hold:
- Task-specific fine-tuning (80 trajs/task) is inside the loop β this is not a zero-shot generalist result; every policy including baselines is per-suite fine-tuned. The claim is about recipe efficiency, not open-world generalization (no unseen-object/scene OOD axis is reported, in contrast to RobotManip-style OOD protocols).
- Baseline adaptation asymmetry. Ο0.5 required action-space padding and hyperparameter surgery; GR00T ran without RTC (not public); InternVLA-M1's backbone was frozen. The authors are transparent about the reproduction effort, but β₯40 pp should be read with the usual adapted-baseline salt.
- Ablation is one task, 10 trials β the recipe decomposition is directional, not statistically tight (cf. RSS 2026's own Beyond Binary Success evaluation-rigor thread).
- The "10Γ less data" comparison conflates axes β baselines' larger pre-training corpora are also differently distributed; the paper's own framing ("the right data in the right way") is the honest version of the claim.
- Payload limits of the G1 constrain task difficulty (authors' stated limitation); dexterity is Dex3-1-level (7-DoF hands), below Shadow/Allegro-class in-hand work.
| Axis | Ξ¨β | GR00T N1.6/1.7 | Ο0.5 (adapted) | EgoVLA / H-RDT / Being-H0 line |
|---|---|---|---|---|
| Human-data role | Representation pre-training only (VLM) | Data-pyramid co-training (incl. ~20k h ego at N1.7) | None (robot fleet) | Unified human-robot action space, co-trained |
| Action expert | 500M MM-DiT flow, joint space | Cross-attention DiT | Same-stack flow expert | DiT / VLM heads, task-space + IK |
| Lower body | Delegated to AMO RL (System-0) | Whole-body in-policy (varies) | N/A (wheeled) | Mostly upper-body |
| Data | 829 h ego + 31 h robot | Orders of magnitude more | PI fleet | Large mixed corpora |
| Open | Full stack (stated) | Weights (Apache 2.0) | Weights (Ο0.5 ckpt) | Partial |
Ξ¨β is best read as the humanoid instantiation of a broader 2026 convergence: representation from cheap heterogeneous data, control from small high-quality in-domain data β the same shape as LBM's staged specialization and the inverse of the everything-in-one-pot co-training the field defaulted to in 2025.
- arXiv: https://arxiv.org/abs/2603.12263 Β· Project: https://psi-lab.ai/Psi0 Β· Code: https://github.com/physical-superintelligence-lab/Psi0
- RSS 2026 β VLA & Manipulation Survey β conference context (humanoid loco-manipulation thread)
- Humanoid VLA Β· System 0/1/2 Β· GR00T series
- EgoDex β the 829 h pre-training corpus
- VLM4VLA β the representation-alignment finding this recipe operationalizes
- Real-Time Chunking β the test-time method whose training-time variant Ξ¨β adopts
- HoMMI Β· RSS-2026 survey humanoid thread β sibling RSS 2026 whole-body work
β Back to Home