Review pi06 - Heungwoo/research GitHub Wiki

In-Depth Review โ€” ฯ€0.6: A VLA Model for General Robot Control

Paper: ฯ€0.6 Model Card ยท Physical Intelligence ยท November 17, 2025 Related summary: ฯ€0.6

This page is the long-form companion to the ฯ€0.6 summary. Everything is sourced from the ฯ€0.6 model card at website.pi-asset.com/pi06star/PI06_model_card.pdf.

๐Ÿ“Ž ฯ€ series context

Release Date Page Role
ฯ€0 Oct 2024 (RSS 2025) โ€” Flow-matching VLA foundation
ฯ€0.5 Apr 2025 (CoRL 2025 Oral) ฯ€0.5 Hierarchical + co-training
ฯ€0.5-KI Dec 2025 Knowledge Insulation The training recipe ฯ€0.6 adopts
ฯ€0.6 Nov 2025 this review / summary Gemma3-4B backbone + KI + metadata prompting
ฯ€*0.6 + RECAP Nov 2025 ฯ€*0.6 + RECAP RL-from-experience on ฯ€0.6 base
ฯ€0.7 Apr 2026 ฯ€0.7 / long-form MEM + world-model subgoals + metadata CFG
(meta) โ€” ฯ€ series evolution Side-by-side of all releases

ฯ€0.6 is the production baseline for 2026. It's the model every ICLR 2026 VLA paper compares against; it's the base for ฯ€*0.6's RL specialists; and it's the architectural template ฯ€0.7 inherits verbatim (same Gemma3-4B + 860M action expert, same Knowledge Insulation).


1. TL;DR

ฯ€0.6 is Physical Intelligence's November 2025 VLA that upgrades ฯ€0.5 along three axes โ€” backbone (PaliGemma โ†’ Gemma3-4B), training (joint โ†’ Knowledge-Insulation two-path), and prompt (task-only โ†’ task + optional metadata) โ€” while keeping the hierarchical design. Result: first ฯ€-series release that achieves strong out-of-the-box performance without task-specific post-training. Laundry (T-shirts + shorts) goes from ~0% (ฯ€0.5 needed fine-tuning) to "folds reliably" out-of-the-box (the card gives no laundry %); box assembly goes from 0% to 20% full assembly out-of-the-box (the one hard static-task number in the card); shirt folding, table bussing, and mobile tasks see consistent throughput/speed gains (exact multipliers are bar-chart estimates, not quoted). 63 ms per action chunk on a single H100 with 3 cameras. ฯ€0.6 is the base model for ฯ€*0.6 + RECAP (RL-from-experience specialists) and the architectural template for ฯ€0.7 (which inherits the backbone + action expert verbatim).

2. Motivation โ€” what ฯ€0.5 left on the table

ฯ€0.5 (April 2025, CoRL 2025 Oral) demonstrated open-world generalization in unseen homes via co-training on heterogeneous data + a hierarchical subtask head. But:

  • Post-training was still necessary. For dexterous tasks (laundry folding, box assembly), ฯ€0.5 needed task-specific fine-tuning with curated high-quality data to reach non-zero success rates. This is expensive and brittle.
  • Action-expert gradients corrupted the VLM. Joint training mixed language-grounded representations with flow-matching gradients โ€” a known stability problem that limited how hard you could push the action expert.
  • No handle to steer behavior at deployment. A trained policy was what it was โ€” no metadata or prompt knob to bias toward high-quality modes at test time.
  • Backbone was showing its age. PaliGemma (ฯ€0 / ฯ€0.5) was fine but Gemma3's multimodal improvements were worth capturing.

ฯ€0.6 addresses all four.


3. Representative diagrams

Figure 1 from the model card โ€” ฯ€0.6 architecture

ฯ€0.6 architecture (Figure 1 from Physical Intelligence, Nov 2025)

Figure 1 of the ฯ€0.6 model card (Physical Intelligence, Nov 17 2025). The VLA consists of a pre-trained VLM (SigLIP 400M + Gemma3 4B) that consumes up to four 448ร—448 images, a language prompt, tokenized proprioceptive state, and optional episode metadata. The backbone produces both discretized actions (FAST tokens) โ€” which supply the VLM's training signal via Knowledge Insulation โ€” and features feeding a separate 860M action expert that generates continuous actions via flow matching. Included for scholarly review.

Figure 2 from the model card โ€” static-task results

ฯ€0.6 static-task results (Figure 2 from Physical Intelligence, Nov 2025)

Figure 2 of the model card. Out-of-the-box (no task-specific fine-tuning) comparison of ฯ€0.5 vs. ฯ€0.6 on four static-robot tasks. Top row: success rate. Bottom row: throughput (successes per hour). ฯ€0.6's biggest wins are on laundry folding (now folds reliably out-of-the-box vs. ~0% before) and box assembly (0 โ†’ 20% full assembly, the one number quoted in the card text) โ€” previously both required fine-tuning with curated data.

Our reconstruction as mermaid

flowchart TB
  subgraph Prompt[Prompt Ct]
    L[Task: 'clean the bedroom']
    SL[Subtask: 'pick up the pillow']
    M[Optional metadata<br/>speed / quality / etc.]
  end

  V[Up to 4 cameras<br/>448ร—448<br/>base + 2 wrist + rear] --> B[ฯ€0.6 VLA<br/>SigLIP 400M + Gemma3-4B]
  PR[Proprioception<br/>tokenized] --> B
  Prompt --> B

  B -- FAST-tokenized actions --> CE[Discrete CE loss<br/>trains VLM representation]
  B -- Web co-training<br/>subtask prediction --> CE
  B -- conditioning features --> AE[860M Action Expert<br/>flow matching]
  noise[Noise] --> AE
  AE -- 5 denoising steps --> A[Continuous action chunk<br/>63 ms / chunk on H100]

  AE -. ๐Ÿšซ NO gradient back to VLM<br/>(Knowledge Insulation) .-> B
Loading

4. Method โ€” what changed vs. ฯ€0.5

4.1 Backbone: PaliGemma โ†’ Gemma3-4B

ฯ€0.5 used PaliGemma (SigLIP vision + Gemma language). ฯ€0.6 swaps in Gemma3-4B (with a 400M SigLIP vision encoder) โ€” a later-generation VLM with stronger multimodal pretraining. This is a ~1-generation upgrade and is kept unchanged in ฯ€0.7.

4.2 Action expert: ~860M parameters, same layer count as backbone

The action expert:

  • Has the same number of layers as the VLM backbone.
  • Consists of ~860M parameters.
  • Uses flow matching to produce continuous action chunks from noise.
  • 5 denoising (Euler) steps at inference.
  • 63 ms per action chunk on a single H100 with 3 camera inputs.
  • Bidirectional attention among action tokens.

Identical to ฯ€0.7's action expert โ€” ฯ€0.7 keeps this subsystem unchanged and adds new prompt modalities + MEM history around it.

4.3 Input format

  • Up to 4 images at 448ร—448: base camera, up to two wrist cameras, optional backward camera for mobile manipulators.
  • Image tokens + tokenized language prompt + tokenized proprioception concatenated.
  • Bidirectional attention among image tokens (inherited from ฯ€0.5); causal attention among text tokens.

4.4 Knowledge Insulation โ€” the key training change

The VLM and action expert are trained simultaneously but with no gradient flow from the action expert into the VLM:

  • VLM is supervised by discretized FAST action tokens (cross-entropy loss โ€” the stable objective it was built for) + co-training tasks including multi-modal web data, subtask prediction, bounding-box / keypoint prediction.
  • Action expert learns continuous actions via flow matching, attending to VLM activations.
  • Gradients from the action-expert flow-matching loss are blocked from propagating into the VLM's parameters.

This is the Knowledge Insulation recipe (NeurIPS 2025 Spotlight โ€” Driess et al.) โ€” formalized at NeurIPS in December 2025 but already deployed in ฯ€0.6 (November 2025). The recipe then propagates unchanged into ฯ€0.7.

4.5 Optional metadata in the prompt

ฯ€0.6 optionally accepts conditioning metadata alongside the language command โ€” a small prompt knob that biases how the task is performed. The space of metadata is not fully documented in the model card but is expanded considerably in ฯ€0.7 (speed / quality / mistake / control mode, each with independent dropout and CFG).

4.6 Data

Largely inherits ฯ€0.5's mix:

  • Cross-embodiment data from PI robots (static + mobile + bimanual + single-arm).
  • External data sources (OXE, community datasets).
  • Diverse mobile and non-mobile home data collected in-house.
  • High-level subtask prediction examples.
  • Multi-modal web datasets including bounding-box and keypoint prediction co-training tasks.

Notable non-inclusion (vs. ฯ€0.7): no autonomous rollouts / failures / RL traces are explicitly part of ฯ€0.6's training data โ€” that's a ฯ€0.7 addition.


5. Results โ€” every number the model card reports

All results are out-of-the-box (no task-specific fine-tuning for either model). ฯ€0.6 is compared against the Knowledge-Insulation-trained ฯ€0.5 (which is itself an improvement over CoRL 2025 ฯ€0.5 โ€” open-sourced at openpi).

5.1 Static tasks (Figure 2)

Caveat (verified against the model card): The model card's text reports only two hard quantitative statements for the static tasks โ€” that ฯ€0.6 can "fully assemble the box 20% of the time" out-of-the-box, and that it "can out-of-the-box fold laundry reliably/consistently" (no percentage given). All other figures below are approximate values read off the Figure 2 bar charts โ€” they are not stated numerically anywhere in the card text, nor in the RECAP report. The reads were cross-checked against the actual Figure 2 image (each task panel has its own y-axis scale, with standard-error bars), and the throughput bars match the table values within reading precision: Shirt โ‰ˆ21โ†’โ‰ˆ50, Laundry 0โ†’โ‰ˆ19, Box 0โ†’โ‰ˆ5 (very large error bar), Table bussing โ‰ˆ28โ†’โ‰ˆ45 succ/hr. Treat all such values as approximate chart reads, not quoted numbers.

Task ฯ€0.5 success ฯ€0.6 success ฯ€0.5 throughput (succ/hr) ฯ€0.6 throughput
Shirt folding (flat start) ~85%* ~87%* ~20* ~50*
Laundry folding (T-shirts + shorts from basket) ~0% (needed fine-tuning) "folds reliably" (~65%*) 0 ~20*
Box assembly ~0% (needed fine-tuning) 20% full assembly (card text) 0 ~5*
Table bussing ~80%* ~100%* ~28* ~45*

* = approximate value read off the Figure 2 bar chart (cross-checked against the figure image), not a number stated in the card text or RECAP report.

Card text says: "Across tasks, ฯ€0.6 shows significant improvement over ฯ€0.5 in speed and often success rates. The biggest differences lie in laundry folding and box assembly โ€” previously these two tasks require fine-tuning with high-quality data to achieve non-zero success rates." So: success on already-solved tasks is roughly flat; the consistently improved axis is throughput/speed, and the headline new capability is non-zero performance on laundry folding and box assembly without fine-tuning.

5.2 Mobile tasks (Figure 3) โ€” four tasks from the ฯ€0.5 paper

Picking up laundry โ†’ basket ยท Tidying bed ยท Putting dishes in sink ยท Putting items in drawer.

Results: "ฯ€0.6 improves throughput over ฯ€0.5 when the average task progress is saturated, or otherwise improves both performance and throughput." No single headline number โ€” the pattern is uniform improvement across all four.

5.3 Generalization tasks (Figure 4) โ€” 4 task suites ร— 3 difficulty levels ร— 12โ€“18 instructions

Mixed static + mobile tasks requiring either:

  • Language-following generalization โ€” "pick up the third fruit from the left", "move to where the fresh milk is kept"
  • Novel-skill generalization โ€” "wipe the spill with the bread", "hang the shorts into oven handle"

Most language instructions and objects are not seen in training. ฯ€0.6 shows "healthy improvements over ฯ€0.5 across all settings." Mobile settings are generally harder (longer horizon, more distractors).

5.4 Latency headline

63 ms per action chunk on a single H100 GPU with 3 camera inputs, using 5 denoising steps. This is the production-deployment number the ฯ€-series carries forward โ€” ฯ€0.7 keeps the same 5-step inference budget.


6. Why ฯ€0.6 works โ€” three contributing factors

The model card attributes the gain to three changes, not one:

  1. Gemma3-4B backbone โ€” better multimodal pretraining transfer.
  2. Knowledge Insulation โ€” lets the VLM learn stable representations while the action expert is free to be aggressive on flow matching.
  3. Diverse training data + optional metadata prompting โ€” the metadata gives a knob for test-time steering without additional training; the diverse data mix provides the coverage that prior models had to fine-tune in.

The model card does not separately ablate these three โ€” a key scientific limitation. ฯ€0.7 later ablates the metadata contribution heavily and shows it's load-bearing for scaling on suboptimal data; Knowledge Insulation's paper provides the independent ablation for the KI training change.


7. Limitations

7.1 Stated limitations (model card)

The ฯ€0.6 model card is a production release document, not a research paper โ€” it does not enumerate limitations directly. Implicit gaps from the text:

  • No separate ablation of backbone upgrade / KI / metadata.
  • Benchmarks are PI-internal; no LIBERO / RLBench / SimplerEnv numbers for external comparison.
  • Generalization results are qualitative ("healthy improvements") rather than fully tabulated per task.
  • Throughput comparisons use normalized axes in several plots; raw successes-per-hour for the hardest tasks are not always disclosed.

7.2 Reviewer's concerns

  • No head-to-head vs. discrete-diffusion VLAs. The ICLR 2026 alternative family (Discrete Diffusion VLA, Unified Diffusion VLA, dVLA) directly challenges the ฯ€-series' separate-action-expert design โ€” ฯ€0.6 doesn't engage with this.
  • No open weights. Like every PI tech report, ฯ€0.6 is not an independently reproducible artifact; the reference open implementation is ฯ€0.5-KI (openpi), not ฯ€0.6.
  • Dexterous tasks still fail ~35% of the time. Laundry at 65% is a huge jump from 0%, but it's not production-ready โ€” hence the immediate follow-up with ฯ€*0.6 + RECAP RL specialists.
  • Box assembly at 20% full assembly is more a "proof of feasibility" than a deployable capability.
  • The metadata space is under-documented. What metadata fields exist? What's their distribution at train time? How are they prompted at test time? ฯ€0.7 fills most of this in retroactively; ฯ€0.6 leaves it opaque.

8. Significance โ€” the production baseline of 2026

ฯ€0.6 is the single most-compared-against VLA of 2026. Concretely:

  • ~All ICLR 2026 VLA papers benchmark against ฯ€0.6 or ฯ€0.5-KI.
  • ฯ€*0.6 + RECAP uses ฯ€0.6 as the base model for RL-from-experience โ€” specialist policies for laundry, box, espresso that are then distilled back into ฯ€0.7 via metadata conditioning.
  • ฯ€0.7 inherits the Gemma3-4B + 860M action expert + Knowledge Insulation stack unchanged. The ฯ€0.7 delta is entirely new prompt modalities + MEM history.
  • Architectural legacy: ฯ€0.6 cements the "VLM backbone + separate flow-matching action expert" pattern that every flow-matching VLA at CoRL 2025 and ICLR 2026 follows. It's the pattern that Discrete Diffusion VLA / Unified Diffusion VLA explicitly challenge by unifying action generation back into the VLM transformer.

Category placement (see Review-VLA-Architecture ยง5.B): ฯ€0.6 is the canonical exemplar of Category B (Flow-matching action expert) โ€” the one other VLA architectures are compared against.


9. What changed from ฯ€0.6 โ†’ ฯ€*0.6 โ†’ ฯ€0.7

A one-screen summary for readers jumping through the lineage:

Axis ฯ€0.6 (Nov 2025) ฯ€*0.6 + RECAP (Nov 2025) ฯ€0.7 (Apr 2026)
Backbone Gemma3-4B same same
Action expert 860M flow-matching same same
Training recipe Knowledge Insulation same + advantage-conditioned RL same + suboptimal data distillation
Prompt Task + optional metadata same + positive/negative conditioning Task + subtask + subgoal images + rich metadata + control mode (each dropout)
History Single frame same MEM video encoder โ€” 4 cameras ร— 6 history frames
World model None None BAGEL-14B generating subgoal images
Data ฯ€0.5 mix ฯ€0.6 mix + real-robot RL ฯ€0.6 mix + failures + autonomous + RL rollouts + egocentric human + DROID
Post-training needed Sometimes Per-task RL Zero-shot matches specialists

Core insight: the architectural chassis from ฯ€0.6 is stable across the entire Nov-2025-to-Apr-2026 window. Improvements come from data, prompt, and history โ€” not from the VLM + action-expert core.


10. Links

โ† Back to PI-pi06 ยท ICLR-2026 ยท Home

โš ๏ธ **GitHub.com Fallback** โš ๏ธ