Review pi06 - Heungwoo/research GitHub Wiki
Paper: ฯ0.6 Model Card ยท Physical Intelligence ยท November 17, 2025 Related summary: ฯ0.6
This page is the long-form companion to the ฯ0.6 summary. Everything is sourced from the ฯ0.6 model card at website.pi-asset.com/pi06star/PI06_model_card.pdf.
| Release | Date | Page | Role |
|---|---|---|---|
| ฯ0 | Oct 2024 (RSS 2025) | โ | Flow-matching VLA foundation |
| ฯ0.5 | Apr 2025 (CoRL 2025 Oral) | ฯ0.5 | Hierarchical + co-training |
| ฯ0.5-KI | Dec 2025 | Knowledge Insulation | The training recipe ฯ0.6 adopts |
| ฯ0.6 | Nov 2025 | this review / summary | Gemma3-4B backbone + KI + metadata prompting |
| ฯ*0.6 + RECAP | Nov 2025 | ฯ*0.6 + RECAP | RL-from-experience on ฯ0.6 base |
| ฯ0.7 | Apr 2026 | ฯ0.7 / long-form | MEM + world-model subgoals + metadata CFG |
| (meta) | โ | ฯ series evolution | Side-by-side of all releases |
ฯ0.6 is the production baseline for 2026. It's the model every ICLR 2026 VLA paper compares against; it's the base for ฯ*0.6's RL specialists; and it's the architectural template ฯ0.7 inherits verbatim (same Gemma3-4B + 860M action expert, same Knowledge Insulation).
ฯ0.6 is Physical Intelligence's November 2025 VLA that upgrades ฯ0.5 along three axes โ backbone (PaliGemma โ Gemma3-4B), training (joint โ Knowledge-Insulation two-path), and prompt (task-only โ task + optional metadata) โ while keeping the hierarchical design. Result: first ฯ-series release that achieves strong out-of-the-box performance without task-specific post-training. Laundry (T-shirts + shorts) goes from ~0% (ฯ0.5 needed fine-tuning) to "folds reliably" out-of-the-box (the card gives no laundry %); box assembly goes from 0% to 20% full assembly out-of-the-box (the one hard static-task number in the card); shirt folding, table bussing, and mobile tasks see consistent throughput/speed gains (exact multipliers are bar-chart estimates, not quoted). 63 ms per action chunk on a single H100 with 3 cameras. ฯ0.6 is the base model for ฯ*0.6 + RECAP (RL-from-experience specialists) and the architectural template for ฯ0.7 (which inherits the backbone + action expert verbatim).
ฯ0.5 (April 2025, CoRL 2025 Oral) demonstrated open-world generalization in unseen homes via co-training on heterogeneous data + a hierarchical subtask head. But:
- Post-training was still necessary. For dexterous tasks (laundry folding, box assembly), ฯ0.5 needed task-specific fine-tuning with curated high-quality data to reach non-zero success rates. This is expensive and brittle.
- Action-expert gradients corrupted the VLM. Joint training mixed language-grounded representations with flow-matching gradients โ a known stability problem that limited how hard you could push the action expert.
- No handle to steer behavior at deployment. A trained policy was what it was โ no metadata or prompt knob to bias toward high-quality modes at test time.
- Backbone was showing its age. PaliGemma (ฯ0 / ฯ0.5) was fine but Gemma3's multimodal improvements were worth capturing.
ฯ0.6 addresses all four.

Figure 1 of the ฯ0.6 model card (Physical Intelligence, Nov 17 2025). The VLA consists of a pre-trained VLM (SigLIP 400M + Gemma3 4B) that consumes up to four 448ร448 images, a language prompt, tokenized proprioceptive state, and optional episode metadata. The backbone produces both discretized actions (FAST tokens) โ which supply the VLM's training signal via Knowledge Insulation โ and features feeding a separate 860M action expert that generates continuous actions via flow matching. Included for scholarly review.

Figure 2 of the model card. Out-of-the-box (no task-specific fine-tuning) comparison of ฯ0.5 vs. ฯ0.6 on four static-robot tasks. Top row: success rate. Bottom row: throughput (successes per hour). ฯ0.6's biggest wins are on laundry folding (now folds reliably out-of-the-box vs. ~0% before) and box assembly (0 โ 20% full assembly, the one number quoted in the card text) โ previously both required fine-tuning with curated data.
flowchart TB
subgraph Prompt[Prompt Ct]
L[Task: 'clean the bedroom']
SL[Subtask: 'pick up the pillow']
M[Optional metadata<br/>speed / quality / etc.]
end
V[Up to 4 cameras<br/>448ร448<br/>base + 2 wrist + rear] --> B[ฯ0.6 VLA<br/>SigLIP 400M + Gemma3-4B]
PR[Proprioception<br/>tokenized] --> B
Prompt --> B
B -- FAST-tokenized actions --> CE[Discrete CE loss<br/>trains VLM representation]
B -- Web co-training<br/>subtask prediction --> CE
B -- conditioning features --> AE[860M Action Expert<br/>flow matching]
noise[Noise] --> AE
AE -- 5 denoising steps --> A[Continuous action chunk<br/>63 ms / chunk on H100]
AE -. ๐ซ NO gradient back to VLM<br/>(Knowledge Insulation) .-> B
ฯ0.5 used PaliGemma (SigLIP vision + Gemma language). ฯ0.6 swaps in Gemma3-4B (with a 400M SigLIP vision encoder) โ a later-generation VLM with stronger multimodal pretraining. This is a ~1-generation upgrade and is kept unchanged in ฯ0.7.
The action expert:
- Has the same number of layers as the VLM backbone.
- Consists of ~860M parameters.
- Uses flow matching to produce continuous action chunks from noise.
- 5 denoising (Euler) steps at inference.
- 63 ms per action chunk on a single H100 with 3 camera inputs.
- Bidirectional attention among action tokens.
Identical to ฯ0.7's action expert โ ฯ0.7 keeps this subsystem unchanged and adds new prompt modalities + MEM history around it.
- Up to 4 images at 448ร448: base camera, up to two wrist cameras, optional backward camera for mobile manipulators.
- Image tokens + tokenized language prompt + tokenized proprioception concatenated.
- Bidirectional attention among image tokens (inherited from ฯ0.5); causal attention among text tokens.
The VLM and action expert are trained simultaneously but with no gradient flow from the action expert into the VLM:
- VLM is supervised by discretized FAST action tokens (cross-entropy loss โ the stable objective it was built for) + co-training tasks including multi-modal web data, subtask prediction, bounding-box / keypoint prediction.
- Action expert learns continuous actions via flow matching, attending to VLM activations.
- Gradients from the action-expert flow-matching loss are blocked from propagating into the VLM's parameters.
This is the Knowledge Insulation recipe (NeurIPS 2025 Spotlight โ Driess et al.) โ formalized at NeurIPS in December 2025 but already deployed in ฯ0.6 (November 2025). The recipe then propagates unchanged into ฯ0.7.
ฯ0.6 optionally accepts conditioning metadata alongside the language command โ a small prompt knob that biases how the task is performed. The space of metadata is not fully documented in the model card but is expanded considerably in ฯ0.7 (speed / quality / mistake / control mode, each with independent dropout and CFG).
Largely inherits ฯ0.5's mix:
- Cross-embodiment data from PI robots (static + mobile + bimanual + single-arm).
- External data sources (OXE, community datasets).
- Diverse mobile and non-mobile home data collected in-house.
- High-level subtask prediction examples.
- Multi-modal web datasets including bounding-box and keypoint prediction co-training tasks.
Notable non-inclusion (vs. ฯ0.7): no autonomous rollouts / failures / RL traces are explicitly part of ฯ0.6's training data โ that's a ฯ0.7 addition.
All results are out-of-the-box (no task-specific fine-tuning for either model). ฯ0.6 is compared against the Knowledge-Insulation-trained ฯ0.5 (which is itself an improvement over CoRL 2025 ฯ0.5 โ open-sourced at openpi).
Caveat (verified against the model card): The model card's text reports only two hard quantitative statements for the static tasks โ that ฯ0.6 can "fully assemble the box 20% of the time" out-of-the-box, and that it "can out-of-the-box fold laundry reliably/consistently" (no percentage given). All other figures below are approximate values read off the Figure 2 bar charts โ they are not stated numerically anywhere in the card text, nor in the RECAP report. The reads were cross-checked against the actual Figure 2 image (each task panel has its own y-axis scale, with standard-error bars), and the throughput bars match the table values within reading precision: Shirt โ21โโ50, Laundry 0โโ19, Box 0โโ5 (very large error bar), Table bussing โ28โโ45 succ/hr. Treat all such values as approximate chart reads, not quoted numbers.
| Task | ฯ0.5 success | ฯ0.6 success | ฯ0.5 throughput (succ/hr) | ฯ0.6 throughput |
|---|---|---|---|---|
| Shirt folding (flat start) | ~85%* | ~87%* | ~20* | ~50* |
| Laundry folding (T-shirts + shorts from basket) | ~0% (needed fine-tuning) | "folds reliably" (~65%*) | 0 | ~20* |
| Box assembly | ~0% (needed fine-tuning) | 20% full assembly (card text) | 0 | ~5* |
| Table bussing | ~80%* | ~100%* | ~28* | ~45* |
* = approximate value read off the Figure 2 bar chart (cross-checked against the figure image), not a number stated in the card text or RECAP report.
Card text says: "Across tasks, ฯ0.6 shows significant improvement over ฯ0.5 in speed and often success rates. The biggest differences lie in laundry folding and box assembly โ previously these two tasks require fine-tuning with high-quality data to achieve non-zero success rates." So: success on already-solved tasks is roughly flat; the consistently improved axis is throughput/speed, and the headline new capability is non-zero performance on laundry folding and box assembly without fine-tuning.
Picking up laundry โ basket ยท Tidying bed ยท Putting dishes in sink ยท Putting items in drawer.
Results: "ฯ0.6 improves throughput over ฯ0.5 when the average task progress is saturated, or otherwise improves both performance and throughput." No single headline number โ the pattern is uniform improvement across all four.
5.3 Generalization tasks (Figure 4) โ 4 task suites ร 3 difficulty levels ร 12โ18 instructions
Mixed static + mobile tasks requiring either:
- Language-following generalization โ "pick up the third fruit from the left", "move to where the fresh milk is kept"
- Novel-skill generalization โ "wipe the spill with the bread", "hang the shorts into oven handle"
Most language instructions and objects are not seen in training. ฯ0.6 shows "healthy improvements over ฯ0.5 across all settings." Mobile settings are generally harder (longer horizon, more distractors).
63 ms per action chunk on a single H100 GPU with 3 camera inputs, using 5 denoising steps. This is the production-deployment number the ฯ-series carries forward โ ฯ0.7 keeps the same 5-step inference budget.
The model card attributes the gain to three changes, not one:
- Gemma3-4B backbone โ better multimodal pretraining transfer.
- Knowledge Insulation โ lets the VLM learn stable representations while the action expert is free to be aggressive on flow matching.
- Diverse training data + optional metadata prompting โ the metadata gives a knob for test-time steering without additional training; the diverse data mix provides the coverage that prior models had to fine-tune in.
The model card does not separately ablate these three โ a key scientific limitation. ฯ0.7 later ablates the metadata contribution heavily and shows it's load-bearing for scaling on suboptimal data; Knowledge Insulation's paper provides the independent ablation for the KI training change.
The ฯ0.6 model card is a production release document, not a research paper โ it does not enumerate limitations directly. Implicit gaps from the text:
- No separate ablation of backbone upgrade / KI / metadata.
- Benchmarks are PI-internal; no LIBERO / RLBench / SimplerEnv numbers for external comparison.
- Generalization results are qualitative ("healthy improvements") rather than fully tabulated per task.
- Throughput comparisons use normalized axes in several plots; raw successes-per-hour for the hardest tasks are not always disclosed.
- No head-to-head vs. discrete-diffusion VLAs. The ICLR 2026 alternative family (Discrete Diffusion VLA, Unified Diffusion VLA, dVLA) directly challenges the ฯ-series' separate-action-expert design โ ฯ0.6 doesn't engage with this.
- No open weights. Like every PI tech report, ฯ0.6 is not an independently reproducible artifact; the reference open implementation is ฯ0.5-KI (openpi), not ฯ0.6.
- Dexterous tasks still fail ~35% of the time. Laundry at 65% is a huge jump from 0%, but it's not production-ready โ hence the immediate follow-up with ฯ*0.6 + RECAP RL specialists.
- Box assembly at 20% full assembly is more a "proof of feasibility" than a deployable capability.
- The metadata space is under-documented. What metadata fields exist? What's their distribution at train time? How are they prompted at test time? ฯ0.7 fills most of this in retroactively; ฯ0.6 leaves it opaque.
ฯ0.6 is the single most-compared-against VLA of 2026. Concretely:
- ~All ICLR 2026 VLA papers benchmark against ฯ0.6 or ฯ0.5-KI.
- ฯ*0.6 + RECAP uses ฯ0.6 as the base model for RL-from-experience โ specialist policies for laundry, box, espresso that are then distilled back into ฯ0.7 via metadata conditioning.
- ฯ0.7 inherits the Gemma3-4B + 860M action expert + Knowledge Insulation stack unchanged. The ฯ0.7 delta is entirely new prompt modalities + MEM history.
- Architectural legacy: ฯ0.6 cements the "VLM backbone + separate flow-matching action expert" pattern that every flow-matching VLA at CoRL 2025 and ICLR 2026 follows. It's the pattern that Discrete Diffusion VLA / Unified Diffusion VLA explicitly challenge by unifying action generation back into the VLM transformer.
Category placement (see Review-VLA-Architecture ยง5.B): ฯ0.6 is the canonical exemplar of Category B (Flow-matching action expert) โ the one other VLA architectures are compared against.
A one-screen summary for readers jumping through the lineage:
| Axis | ฯ0.6 (Nov 2025) | ฯ*0.6 + RECAP (Nov 2025) | ฯ0.7 (Apr 2026) |
|---|---|---|---|
| Backbone | Gemma3-4B | same | same |
| Action expert | 860M flow-matching | same | same |
| Training recipe | Knowledge Insulation | same + advantage-conditioned RL | same + suboptimal data distillation |
| Prompt | Task + optional metadata | same + positive/negative conditioning | Task + subtask + subgoal images + rich metadata + control mode (each dropout) |
| History | Single frame | same | MEM video encoder โ 4 cameras ร 6 history frames |
| World model | None | None | BAGEL-14B generating subgoal images |
| Data | ฯ0.5 mix | ฯ0.6 mix + real-robot RL | ฯ0.6 mix + failures + autonomous + RL rollouts + egocentric human + DROID |
| Post-training needed | Sometimes | Per-task RL | Zero-shot matches specialists |
Core insight: the architectural chassis from ฯ0.6 is stable across the entire Nov-2025-to-Apr-2026 window. Improvements come from data, prompt, and history โ not from the VLM + action-expert core.
- Model card (primary source): https://website.pi-asset.com/pi06star/PI06_model_card.pdf
- ฯ0.6 summary page (this wiki): PI-pi06
- Knowledge Insulation (NeurIPS 2025 Spotlight): NeurIPS-2025-Knowledge-Insulation ยท https://arxiv.org/abs/2505.23705
- Predecessor (ฯ0.5): CoRL-2025-pi05 ยท https://arxiv.org/abs/2504.16054
- RL successor (ฯ*0.6 + RECAP): PI-RECAP ยท https://arxiv.org/abs/2511.14759
- Architectural successor (ฯ0.7): PI-pi07 / Review-pi07 ยท https://www.pi.website/blog/pi07
- ฯ series side-by-side: pi-series-evolution
- OpenPi (ฯ0.5-KI reference implementation): https://github.com/Physical-Intelligence/openpi
- Gemma3 technical report: https://arxiv.org/abs/2503.19786
- FAST action tokenizer: https://arxiv.org/abs/2501.09747