pi series evolution - Heungwoo/research GitHub Wiki

Ο€ Series Evolution β€” Ο€0 β†’ Ο€0.5 β†’ Ο€0.6 β†’ Ο€*0.6 β†’ Ο€0.7

A side-by-side comparison of Physical Intelligence's VLA releases through model, data, and training lenses. Each version is an additive delta on the previous one β€” the table makes the deltas legible.

Timeline

Model Date arXiv / Ref Headline claim
Ο€0 Oct 2024 (RSS 2025) 2410.24164 First cross-embodiment flow-matching VLA
Ο€0-FAST Jan 2025 FAST tokenizer (RSS 2025) 5Γ— faster autoregressive variant via action tokenization
Ο€0.5 Apr 2025 (CoRL 2025) 2504.16054 Open-world generalization in unseen homes via co-training
Ο€0.6 Nov 2025 Model card Out-of-the-box dexterity β€” no post-training needed
Ο€*0.6 + RECAP Nov 2025 2511.14759 ("Ο€0.6: a VLA That Learns From Experience"*) RL-from-experience for flow-matching VLAs (RECAP = RL with Experience & Corrections via Advantage-conditioned Policies)
Ο€0.7 Apr 2026 pi.website/pi07 Steerable generalist with compositional generalization

Model architecture deltas

Ο€0 Ο€0.5 Ο€0.6 Ο€0.7
VLM backbone PaliGemma-3B Standard VLM (web-pretrained) Gemma3-4B (switch) Gemma3-4B (same)
Action expert Flow matching, ~300M Flow matching (smaller than backbone) 860M flow matching 860M flow matching (same)
Total params ~3B ~3–4B ~5B ~5B + 14B world model (BAGEL)
Hierarchy Flat High-level subtask prediction added Same hierarchical (preserved) + world-model subgoal images in prompt
Memory / history Single-frame Single-frame Single-frame MEM video history encoder (6 frames, temporal+spatial compression)
Proprioception Discretized tokens Discretized tokens Discretized tokens Linear projection to backbone dim
Input cameras 3 up to 3 up to 4 (base + 2 wrist + rear) @ 448Β² 4 cameras + up to 3 subgoal images
Action chunk / denoise 50 / 10 50 / fewer 50 / 5 (63ms on H100) 50 / 5 (+ RTC training for 0–240ms latency)
Attention mask Causal Global bidirectional (image + text λͺ¨λ‘ bidir, per Ο€0.7 Appendix B) (paper λͺ…μ‹œ μ—†μŒ β€” Ο€0.7 Appendix B "no image goals β†’ Ο€0.5 동일"μ—μ„œ μ—­μΆ”λ‘ ν•˜λ©΄ Ο€0.5와 같은 global bidirectional일 κ°€λŠ₯μ„± λ†’μŒ) Conditional block-causal (Appendix B): subgoal images μžˆμ„ λ•Œ β†’ obs & subgoal tokens bidir within themselves, subgoal이 obs attend, 후속 text tokens (subtask/metadata) causal, 50 action tokens bidir + VLM activations attend; 이미지 goals 없을 λ•Œ β†’ Ο€0.5-style global bidirectional둜 fallback

Training recipe deltas

Ο€0 Ο€0.5 Ο€0.6 Ο€0.7
Main loss Flow-matching (continuous) Two-stage: FAST-token pretrain β†’ flow-matching post-train Knowledge Insulation (KI): VLM learns FAST + web, action expert learns continuous; gradients don't flow back KI (inherited)
Prompt Ct Task string only Task + semantic subtask β„“Μ‚ Task + subtask + optional metadata Task + subtask + subgoal images + metadata + control-mode, each with independent dropout
Post-training Required per task Required for many tasks Out-of-the-box for many tasks Out-of-the-box matches RL specialists
RL β€” β€” β€” Ο€*0.6 RECAP is a separate branch
CFG β€” β€” β€” Classifier-free guidance on metadata (Ξ² ∈ {1.3, 1.7, 2.2})
Inference-latency handling β€” β€” β€” Real-Time Chunking (RTC) trained with simulated 0–240ms delays

Data deltas

Ο€0 Ο€0.5 Ο€0.6 Ο€0.7
Demonstrations Multi-robot (single-arm, bimanual, mobile) ~400 hrs mobile manipulator Γ— ~100 homes + cross-embodiment lab + OXE Ο€0.5 mix + expanded in-home Ο€0.5/0.6 mix further expanded
Cross-embodiment Yes Yes (97.6% pretraining = non-MM) Yes Yes + UR5e zero-shot transfer demonstrated
Web / multimodal data Inherited via VLM Added (detection, captioning, VQA, semantic prediction co-training) Same + bbox/keypoint prediction Same + video captioning of robot data and web videos
Human video β€” β€” β€” Egocentric human video as first-class data source
Suboptimal / failure data Excluded Excluded Excluded Heavily used: failures, RL rollouts from Ο€*0.6, autonomous eval data, interventions
Open-source robot data OXE OXE + community sets Same + DROID (Franka)

Where each release sits conceptually

flowchart TB
  pi0["Ο€0 (2024)<br/>Cross-embodiment<br/>flow-matching VLA"]
  pi05["Ο€0.5 (2025)<br/>+ co-training<br/>+ subtask hierarchy<br/>+ FAST-token pretrain<br/>β†’ open-world generalization"]
  pi06["Ο€0.6 (Nov 2025)<br/>+ Gemma3-4B + KI<br/>+ optional metadata<br/>β†’ out-of-the-box dexterity"]
  pistar["Ο€*0.6 (Nov 2025)<br/>+ RECAP<br/>(advantage-conditioned RL)<br/>β†’ specialist-grade policies"]
  pi07["Ο€0.7 (Apr 2026)<br/>+ MEM history + subgoal<br/>  images + metadata<br/>+ suboptimal + RL data<br/>β†’ compositional generalization"]
  pi0 --> pi05 --> pi06
  pi06 --> pistar
  pi06 --> pi07
  pistar -. rollouts distilled .-> pi07
Loading

Axis-by-axis summary of novelties

Model axis

  • Ο€0 β†’ Ο€0.5: Hierarchical split (high-level subtask head + low-level action expert). Discrete FAST-token stage for pretraining.
  • Ο€0.5 β†’ Ο€0.6: Backbone upgraded to Gemma3-4B, action expert grown to 860M, Knowledge Insulation stabilizes joint training.
  • Ο€0.6 β†’ Ο€0.7: Keeps the same backbone/expert; adds MEM video-history encoder, subgoal-image conditioning, linear-projected proprioception, RTC, and CFG on metadata. Worldmodel (BAGEL-14B) is a separate module that feeds the prompt, not an architectural change to the policy itself.

Data axis

  • Ο€0 β†’ Ο€0.5: Large-scale home data (~400 hrs MM Γ— 100 homes), web multimodal co-training.
  • Ο€0.5 β†’ Ο€0.6: More of the same, richer metadata annotation.
  • Ο€0.6 β†’ Ο€0.7: Doubles down on heterogeneity β€” failures, autonomous rollouts, RL-specialist traces, egocentric human video, DROID. The novelty is not the sources but that metadata conditioning lets them be used productively instead of averaged.

Training axis

  • Ο€0 β†’ Ο€0.5: Two-stage FASTβ†’flow-matching; co-training on VQA/detection/subtask text.
  • Ο€0.5 β†’ Ο€0.6: Knowledge Insulation (one-stage joint training with gradient-insulated VLM).
  • Ο€0.6 β†’ Ο€*0.6: Advantage-conditioned RL (RECAP) β€” no log-prob RL needed (flow-matching policies give no action log-probs, so PPO/SAC are inapplicable). A multi-task distributional value function $V^{\pi_{ref}}(o_t,\ell)$ is trained on all data (demos + autonomous rollouts + teleoperated interventions); the policy then conditions on a binarized advantage indicator as a text token ("Advantage: positive/negative"), and at inference is run on the positive branch. Trained end-to-end with Knowledge Insulation. On the hardest tasks (laundry folding, box assembly, espresso) RECAP more than doubles task throughput and roughly halves the failure rate vs the base policy.
  • Ο€0.6 β†’ Ο€0.7: Rich prompt conditioning + per-component dropout + CFG, distilling Ο€*0.6 RL behavior into a generalist without RL at post-train time.

Key takeaway

The Ο€ series is evolving along a consistent thesis: push more capability into pre-training so post-training becomes optional. Ο€0 needed per-task finetuning; Ο€0.5 reduced it; Ο€0.6 removed it for many tasks; Ο€0.7 removes it even for tasks previously only solvable by RL specialists β€” and shows the first signs of compositional (not just in-distribution) generalization in the series (the Ο€0.7 blog phrases it as "the first signs of compositional generalization," recombining skills across tasks).

Links

← Back to Home

⚠️ **GitHub.com Fallback** ⚠️