PI pi07 - Heungwoo/research GitHub Wiki
Venue: Physical Intelligence (technical report) ยท Date: April 16, 2026 Category: Baseline (VLA Architecture + Data) Trend tag: Baseline for Trend 1 (supersedes ฯ0.6)
flowchart LR
subgraph Prompt[Multimodal Prompt Ct]
L[Task instruction<br/>e.g. 'clean the kitchen']
SL[Subtask โฬ<br/>e.g. 'open fridge']
SG[Subgoal images gt<br/>multi-view]
M[Metadata<br/>speed / quality 1-5 / mistake]
C[Control mode<br/>joint or EE]
end
V[4 cameras ร up to<br/>6 history frames<br/>448ร448] --> MEM[MEM-style video<br/>history encoder]
MEM --> B[Gemma3-4B<br/>VLM backbone]
Prompt --> B
PR[Proprioception qt<br/>linear projection] --> B
B --> AE[860M flow-matching<br/>Action Expert<br/>50-step chunk]
noise[Noise] --> AE
AE -- 5 Euler steps --> A[Continuous action chunk]
A -- RTC, 240ms tolerance --> R[Real robot]
HL[High-level policy<br/>same arch] -.generates.-> SL
WM[World model gฯ<br/>BAGEL-14B MoT] -.generates.-> SG
Prior VLAs (ฯ0, ฯ0.5, ฯ0.6) were trained on high-quality demos with minimal prompt context โ just a task string. This produces generalists that need task-specific RL/SFT post-training (ฯ*0.6) to match specialist performance, and that fail to compositionally generalize to unseen combinations (new appliances, new embodiments). Naively mixing failed/suboptimal/autonomous data into training just averages modes and degrades performance.
Core idea: expand the training context Ct from a plain instruction to a rich multimodal prompt so the model can absorb heterogeneous data without mode collapse:
- Subtask text โฬ โ intermediate semantic step (inherited from ฯ0.5).
- Subgoal images โ multi-view near-future images, generated at runtime by a BAGEL-14B world model that is web-pretrained on image editing/video.
- Episode metadata โ discretized speed, quality score 1โ5, mistake flag. At test time, always prompt "quality = 5, mistake = false" to steer toward high-quality modes.
- Control mode โ joint vs. end-effector label. Each component is randomly dropped out at train time (subgoal kept in 25%, subtask in 70% of those, metadata dropped 15%), so any subset can be used at inference. Classifier-free guidance on metadata further amplifies high-quality behavior.
Architecture: ~5B total = Gemma3-4B VLM backbone + MEM video-history encoder (temporal+spatial compression, fixed token count regardless of #frames) + 860M flow-matching action expert (50-token chunk, 5 Euler steps). Inherits Knowledge Insulation from ฯ0.6 (VLM trained on FAST tokens + web; action-expert gradients don't propagate back). Training-time Real-Time Chunking (RTC) simulates 0โ240ms inference delay for smooth deployment.
Data: the prompt-conditioning is what unlocks the real novelty โ heavy use of suboptimal data: lower-quality demos, failure episodes, autonomous rollouts from ฯ*0.6 RL training, human egocentric video, open-source robot data (DROID etc.), and web multimodal (VQA, detection, captioning). Metadata labels disambiguate each episode's quality so the model can distill RL-specialist behavior rather than average it away.
On the four ฯ*0.6 RL-specialist tasks (laundry, diverse laundry, box building, espresso), a single general-purpose ฯ0.7 matches or exceeds specialist success rate and throughput โ no task-specific post-training. Also matches SFT specialists on shirt-inside-out, PB sandwich, drive-through-door, slice zucchini, peel, trash. Zero-shot cross-embodiment: ฯ0.7 folds t-shirts on a UR5e (never trained for laundry on UR5e) at the level of expert teleoperators โ 85.6% task progress / 80% success vs. 10 experienced teleoperators' 90.9% / 80.6% โ discovering embodiment-adapted strategies (e.g. single-arm or vertical grasps) rather than copying source behaviors. Compositional generalization: operates a never-seen air fryer (loading a sweet potato, unloading, toasting a bagel) with zero air-fryer training episodes in the robot data โ only similar appliances seen in human and external datasets โ by being coached through with step-by-step language. Ablations: removing metadata or eval-data both hurt โ diverse data and detailed context are synergistic.
Shifts the frontier of the ฯ series from "post-train per task" to "one prompt-steerable generalist." Three structural moves:
- Prompt as the integration surface โ instead of changing architecture for new capabilities, package them as extra prompt modalities (ร la LLM prompt expansion).
- Suboptimal-data-as-a-feature โ RL rollouts and failures become training signal via metadata conditioning, sidestepping the ฯ*0.6 + RECAP detour for many tasks.
- World-model subgoal images bring visual grounding into the prompt, complementing subtask text when language is ambiguous. Positions ฯ0.7 as the new production baseline against which ICLR 2026's discrete-diffusion and memory VLAs will be compared, and gives the first ฯ-series paper with strong claims of compositional (not just in-distribution) generalization.
- Blog / project: https://www.pi.website/blog/pi07
- Paper PDF: https://www.pi.website/download/pi07.pdf
- arXiv (v1 Apr 16 2026, v2 Apr 24 2026): https://arxiv.org/abs/2604.15483
- Project page: https://pi.website/pi07
- ฯ0.6 model card: https://website.pi-asset.com/pi06star/PI06_model_card.pdf
- ฯ*0.6 (RECAP) arXiv: https://arxiv.org/abs/2511.14759
- ฯ0.5 arXiv: https://arxiv.org/abs/2504.16054
- ฯ0 arXiv: https://arxiv.org/abs/2410.24164
- OpenPi: https://github.com/Physical-Intelligence/openpi
- TechCrunch coverage: https://techcrunch.com/2026/04/16/physical-intelligence-a-hot-robotics-startup-says-its-new-robot-brain-can-figure-out-tasks-it-was-never-taught/
For a long-form review with full results tables, all ablations, and limitations: In-Depth Review of ฯ0.7.
- ฯ series evolution โ side-by-side model/data/training comparison across ฯ0 โ ฯ0.7
- ฯ0.6 (direct predecessor)
- ฯ*0.6 + RECAP (the RL sibling whose specialists ฯ0.7 now distills)
- MemoryVLA (memory axis โ ฯ0.7 uses MEM-style history encoder)
- Discrete Diffusion VLA (the architectural alternative)
- Survey: VLA & Manipulation
โ Back to ICLR-2026