Review Psi0 - Heungwoo/research GitHub Wiki

In-Depth Review β€” Ξ¨β‚€ (Psi-Zero): An Open Foundation Model Towards Universal Humanoid Loco-Manipulation

Paper: Ξ¨β‚€: An Open Foundation Model Towards Universal Humanoid Loco-Manipulation Β· RSS 2026 (Sydney, oral β€” Humanoids session) Authors: Songlin Wei*, Hongyi Jing*, Boqian Li*, Zhenyu Zhao*, Jiageng Mao, … Marco Pavone, Di Huang, Yue Wang† β€” USC Physical Superintelligence (PSI) Lab Γ— NVIDIA Γ— WorldEngine arXiv: 2603.12263 (v1 Mar 12, 2026) Β· Project: psi-lab.ai/Psi0 Β· Code: github.com/physical-superintelligence-lab/Psi0 Openness: full ecosystem release stated β€” teleoperation stack, data/training pipeline, model weights (~2.5B), real-time inference engine

Companion reviews: RSS 2026 survey Β· Humanoid VLA Β· GR00T series Β· System 0/1/2 Β· Cross-Embodiment Β· EgoDex.


1. TL;DR

  1. The anti-scaling thesis for humanoid VLAs: don't co-train one monolithic policy on mixed human + humanoid data β€” the two action distributions are kinematically incompatible, so a single model wastes capacity reconciling them. Instead, decouple: pre-train the VLM on human video for semantics and visual-action representations, then post-train a separate action expert on humanoid data alone for joint-space control.
  2. The data-efficiency headline: with only ~800 h of egocentric human video (EgoDex) + ~30 h of real humanoid data (Humanoid Everyday), Ξ¨β‚€ beats baselines pre-trained on >10Γ— the data β€” including GR00T N1.6 β€” by >40 percentage points in overall success rate on eight real-world long-horizon loco-manipulation tasks on a Unitree G1.
  3. Triple-system architecture mapping cleanly onto the System 0/1/2 framing: System-2 = Qwen3-VL-2B-Instruct backbone, System-1 = 500M flow-matching MM-DiT action expert (SD3-style dual modulation + joint attention β€” ablated to beat a naive DiT head), System-0 = the off-the-shelf AMO RL controller mapping 8-DoF base/torso commands to 15-DoF lower-body joints. 36-DoF policy action β†’ 43-DoF whole-body control.
  4. A pre-training trick worth stealing: the VLM is trained to predict only a single next-step action (48-DoF task-space, FAST-tokenized to ~20 tokens with a custom-fit tokenizer) instead of action chunks β€” enough to learn task semantics and aligned visual representations at a fraction of the compute, since chunked control is delegated to the post-trained expert anyway. The ablation validates it: frozen stock Qwen3-VL scores 0–2/10; EgoDex next-action pre-training lifts it to 6/10 before any real-robot fine-tuning beyond the expert.
  5. Training-time Real-Time Chunking (RTC) rather than test-time guidance: inference-delay simulation during training (first d ~ U(0, d_max) action tokens un-noised and loss-masked) because test-time RTC guidance was unstable on their model β€” a practical negative result for the RTC line.

2. Why this paper matters

  • It attacks the dominant humanoid-VLA recipe head-on. GR00T's data pyramid, EgoVLA, H-RDT, and Being-H0 all co-train across human and robot action distributions in one model. Ξ¨β‚€ argues (and shows on hardware) that this is the wrong use of human data: human video should shape the representation, not the action head. This is the humanoid-domain echo of VLM4VLA's finding that what matters is aligning the backbone's representations with control β€” and of LBM-style stage-decoupled recipes.
  • "Scale the right data in the right way." The 800 h + 30 h result is the strongest data-efficiency claim in the humanoid space to date β€” a direct counter to the volume-first posture, and a challenge to GR00T N1.7's ~20k-hour EgoScale bet. Notably the quality bar is specific: high-quality egocentric manipulation video (EgoDex's Vision-Pro-tracked 829 h), not in-the-wild internet clips.
  • Openness at the humanoid tier. GR00T is the only comparable open humanoid foundation model; Ξ¨β‚€ ships smaller (2.5B vs 3B), with the full teleop + training + deployment stack, from an academic lab. If the release holds, it becomes the default academic humanoid-VLA baseline.
  • One more Qwen-backbone VLA built outside Alibaba β€” Qwen3-VL-2B joins GR00T (Qwen3-VL-2B) as external adopters while Qwen's own robot models stay closed.

3. Architecture

Psi-0 real-world loco-manipulation (Figure 1 of arXiv 2603.12263, Β© the authors)

Figure 1 of the Ξ¨β‚€ paper: a composite of Unitree G1 rollouts in the pantry evaluation environment β€” taking a cup from the coffee machine, pushing a cart, wiping the table, grasping a bottle into the sink, pushing the fridge door. The single scene captures the paper's task spectrum: dexterous-hand grasping, whole-body pushes, and locomotion between stations, all >2,000-step episodes at 30 Hz.

flowchart LR
  subgraph S2["System-2 β€” Qwen3-VL-2B-Instruct (pre-trained on EgoDex)"]
    OBS["Head camera (D435i) + instruction<br/>(+ 28-DoF arm/hand state at post-train)"]
  end
  S2 -- "hidden features z_t" --> S1
  subgraph S1["System-1 β€” MM-DiT action expert (500M, flow matching)"]
    MOD["Dual FiLM modulation by flow timestep Ο„<br/>(VL branch and Action branch separately)"]
    JA["Joint global attention over VL + action tokens"]
  end
  S1 --> ACT["36-DoF action chunk a_{t:t+H}:<br/>14 hand + 14 arm + torso rpy(3)<br/>+ base height + v_x, v_y, v_yaw + p_yaw"]
  ACT -- "8-DoF base/torso commands" --> S0["System-0 β€” AMO RL tracking policy<br/>β†’ 15-DoF legs + waist"]
  ACT -- "28-DoF upper body direct" --> G1["Unitree G1 (29 DoF) + 2Γ— Dex3-1 hands (7 DoF each)<br/>43-DoF whole-body control"]
  S0 --> G1

  classDef vlm fill:#bbdefb,stroke:#1565c0,color:#000
  classDef act fill:#c8e6c9,stroke:#2e7d32,color:#000
  classDef ctl fill:#fff9c4,stroke:#f57f17,color:#000
  class S2 vlm
  class S1,ACT act
  class S0,G1 ctl
Loading
  • MM-DiT vs naive DiT: the expert adapts Stable Diffusion 3's MM-DiT β€” the flow timestep modulates the vision-language and action branches separately (dual FiLM), and both token sets share joint global attention per block, versus a naive DiT's cross-attention into VL features. Ablated: MM-DiT consistently wins; the authors attribute naive DiT's weakness to its text-to-image conditioning heritage.
  • Single head camera only; whole-body proprioception (upper joints, torso rpy, base height) as state. ~160 ms per forward pass at ~2.5B total parameters.

4. Training recipe β€” three stages, three objectives

Stage Data Trains Objective Compute
1. VLM pre-training EgoDex (829 h / ~900M frames, actions re-expressed in head-camera frame, 3Γ— frame-rate upsampling, q01/q99 normalization, no state input) + Humanoid Everyday (31 h, 260 tasks, as visual-gap mitigation) VLM (full) Autoregressive single next-step action prediction in a unified 48-DoF task space ({9-DoF wrist pose + 5Γ—3D fingertips} Γ— 2 hands), FAST-tokenized (custom tokenizer: L1 0.005 vs 0.01 stock; 48 dims β†’ ~20 tokens) 64Γ—A100, 10 days, LR 1e-4, batch 1024
2. Expert post-training Humanoid Everyday (~3M frames real teleop) MM-DiT from scratch, VLM frozen Flow matching on 36-DoF joint-space chunks 32Γ—A100, ~30 h, batch 2048
3. Task fine-tuning 80 teleop trajectories per task Action expert only Flow matching, 40k steps, cosine LR (small)

Plus training-time RTC: each prediction conditions on the previously committed chunk; the first d ~ U(0, d_max) tokens are exposed un-noised and masked from the loss, simulating inference delay. Test-time RTC guidance (gradient steering) was tried and found unstable on this model.

Teleoperation system (the data engine): PICO headset + wrist trackers β†’ multi-target IK for arms/torso; MANUS gloves β†’ full dexterous-hand DoF; waist/foot trackers β†’ high-level locomotion commands to the AMO policy. Locomotion is deliberately decoupled from whole-body retargeting β€” trading expressiveness for lower-body stability and single-operator practicality (contrast TWIST2-style end-to-end whole-body teleop).

5. Results

Setup: 8 real-world long-horizon tasks (pantry/household; 3–5 sub-tasks each; most >2,000 steps at 30 Hz; pick-place, pushing carts and fridge doors, wiping, faucet turning, chip-tray pulling), 10 rollouts per model per task; evaluators may intervene after a failed sub-task so later sub-tasks remain measurable; overall success = all sub-tasks completed.

  • Headline: Ξ¨β‚€'s average overall success rate is β‰₯40 pp above the second-best baseline, GR00T N1.6 (3B, fine-tuned on the same 80-traj/task data with default recipe). Other baselines β€” Ο€0.5 (action space padded 30β†’36), InternVLA-M1 (jittery chunks), H-RDT (2B DiT, imprecise), EgoVLA (weak lower-body priors), Diffusion Policy and ACT (fail broadly) β€” trail further. Several baselines score 0 on multiple tasks (task-wise chart bottoms out at 0 for most non-Ξ¨β‚€ policies on the harder tasks).
  • Skill-level radar: strong across grasping, placing, pouring, rotating, squatting, walking, carrying, pushing, pulling β€” i.e., the wins are not concentrated in one skill family.
  • Ablation (dual-arm 3-step task: right-arm P&P β†’ left-arm P&P β†’ dual-arm box lift):
EgoDex pre-train HE post-train RTC Head Overall
βœ— βœ— βœ— naive DiT 0/10
βœ— βœ— βœ— MM-DiT 2/10
βœ“ βœ— βœ— MM-DiT 6/10
βœ“ βœ“ βœ— MM-DiT 8/10
βœ“ βœ“ βœ“ MM-DiT 9/10

The decomposition is clean: human-video pre-training is worth +4/10, cross-task humanoid post-training +2, RTC +1 (with collision reduction and higher rollout throughput noted qualitatively). The frozen-stock-Qwen3-VL row (0–2/10) is the direct evidence that generic VLM features do not suffice β€” next-action pre-training is what aligns the representation, echoing VLM4VLA's action-supervision finding at the humanoid scale.

6. Critical assessment

Genuinely novel / load-bearing: (a) the decoupled human-video→VLM / robot-data→expert recipe with hardware-validated superiority over co-training; (b) single next-step action pre-training as a compute-cheap alignment objective; (c) MM-DiT as an action head (first humanoid VLA to ablate SD3-style dual-stream conditioning against naive DiT); (d) the 800 h + 30 h data-efficiency result itself.

Caveats a careful reader should hold:

  1. Task-specific fine-tuning (80 trajs/task) is inside the loop β€” this is not a zero-shot generalist result; every policy including baselines is per-suite fine-tuned. The claim is about recipe efficiency, not open-world generalization (no unseen-object/scene OOD axis is reported, in contrast to RobotManip-style OOD protocols).
  2. Baseline adaptation asymmetry. Ο€0.5 required action-space padding and hyperparameter surgery; GR00T ran without RTC (not public); InternVLA-M1's backbone was frozen. The authors are transparent about the reproduction effort, but β‰₯40 pp should be read with the usual adapted-baseline salt.
  3. Ablation is one task, 10 trials β€” the recipe decomposition is directional, not statistically tight (cf. RSS 2026's own Beyond Binary Success evaluation-rigor thread).
  4. The "10Γ— less data" comparison conflates axes β€” baselines' larger pre-training corpora are also differently distributed; the paper's own framing ("the right data in the right way") is the honest version of the claim.
  5. Payload limits of the G1 constrain task difficulty (authors' stated limitation); dexterity is Dex3-1-level (7-DoF hands), below Shadow/Allegro-class in-hand work.

7. Positioning

Axis Ξ¨β‚€ GR00T N1.6/1.7 Ο€0.5 (adapted) EgoVLA / H-RDT / Being-H0 line
Human-data role Representation pre-training only (VLM) Data-pyramid co-training (incl. ~20k h ego at N1.7) None (robot fleet) Unified human-robot action space, co-trained
Action expert 500M MM-DiT flow, joint space Cross-attention DiT Same-stack flow expert DiT / VLM heads, task-space + IK
Lower body Delegated to AMO RL (System-0) Whole-body in-policy (varies) N/A (wheeled) Mostly upper-body
Data 829 h ego + 31 h robot Orders of magnitude more PI fleet Large mixed corpora
Open Full stack (stated) Weights (Apache 2.0) Weights (Ο€0.5 ckpt) Partial

Ξ¨β‚€ is best read as the humanoid instantiation of a broader 2026 convergence: representation from cheap heterogeneous data, control from small high-quality in-domain data β€” the same shape as LBM's staged specialization and the inverse of the everything-in-one-pot co-training the field defaulted to in 2025.

8. Links & related pages

← Back to Home

⚠️ **GitHub.com Fallback** ⚠️