PI pi06 - Heungwoo/research GitHub Wiki
Venue: Physical Intelligence (technical report) ยท Date: Nov 2025 Category: Baseline (VLA Architecture) Trend tag: Baseline for Trend 1
flowchart LR
V[Camera frames<br/>3 views] --> B[Gemma3-4B<br/>VLM backbone]
L[Language instruction] --> B
B --> HL[High-level subtask predictor]
B --> AE[~860M Flow-Matching<br/>Action Expert]
HL --> AE
noise[Noise prior] --> AE
AE -- 5 Euler steps --> A[Continuous action chunk]
A -- 63 ms on H100 --> R[Real robot]
See the ฯ0.6 model card (link below) for the authors' own architecture diagram.
A generalist robot policy needs strong visual-linguistic grounding and precise low-level control and sub-100 ms inference. Earlier generalist models either decoded actions autoregressively (too slow at billion-scale) or used ad-hoc MLP heads (no generative modeling of action distributions).
Hierarchical VLA (preserving the ฯ0.5 design): a high-level component predicts language-level subtasks; a low-level component generates continuous action chunks. The backbone is initialized from Gemma3-4B (with a SigLIP-400M vision encoder). The ~860M-parameter "action expert" is a parallel transformer expert with the same number of layers as the backbone (a mixture-of-experts arrangement, as in ฯ0/ฯ0.5 โ not an external MLP head), and is trained with flow matching โ it learns a continuous vector field that transports a noise distribution to action chunks, conditioned on backbone features; gradients from the action expert do not flow back into the VLM backbone (Knowledge Insulation). The backbone additionally predicts FAST discrete action tokens and co-training web data alongside the flow-matched continuous actions. At inference, 5 denoising (Euler) steps suffice.
On a single H100 GPU with 3 cameras: ~63 ms per action chunk. Generalizes across multiple embodiments and real-world manipulation tasks. Deployed on Physical Intelligence's commercial and academic platforms.
The de facto production baseline for 2026 VLA research. Almost every ICLR 2026 VLA paper compares against ฯ0, ฯ0.5, or ฯ0.6. Architectural legacy: established "VLM backbone + separate flow-matching action expert" as the dominant pattern โ the pattern that ICLR 2026's discrete-diffusion VLAs directly challenge (Discrete Diffusion VLA, Unified Diffusion VLA).
Now superseded by ฯ0.7 (Apr 2026), which keeps the same backbone+expert and adds MEM history, subgoal-image world-model conditioning, expanded episode-metadata prompting (ฯ0.6 already supports optional metadata conditioning), and suboptimal-data distillation.
- ฯ0.6 model card (Physical Intelligence): https://website.pi-asset.com/pi06star/PI06_model_card.pdf
- ฯ0 paper (predecessor): https://arxiv.org/html/2410.24164v1
- OpenPi (open implementation): https://github.com/Physical-Intelligence/openpi
- ฯ0 / ฯ0-FAST blog: https://huggingface.co/blog/pi0
For a long-form review with model-card figures, accuracy tables, Knowledge Insulation training details, limitations, and the ฯ0.6 โ ฯ*0.6 โ ฯ0.7 delta: In-Depth Review of ฯ0.6.
Compare the ฯ-series' flow-matching action expert to other VLA architecture families: VLA Architectures Review.
- ฯ0.7 (the Apr 2026 successor)
- ฯ series evolution (ฯ0 โ ฯ0.7 side-by-side)
- ฯ*0.6 + RECAP (the experience-learning extension)
- Discrete Diffusion VLA (the structural alternative)
- Survey: VLA & Manipulation
โ Back to ICLR-2026