PI pi06 - Heungwoo/research GitHub Wiki

ฯ€0.6 โ€” A Vision-Language-Action Flow Model for General Robot Control

Venue: Physical Intelligence (technical report) ยท Date: Nov 2025 Category: Baseline (VLA Architecture) Trend tag: Baseline for Trend 1

Approach diagram

flowchart LR
  V[Camera frames<br/>3 views] --> B[Gemma3-4B<br/>VLM backbone]
  L[Language instruction] --> B
  B --> HL[High-level subtask predictor]
  B --> AE[~860M Flow-Matching<br/>Action Expert]
  HL --> AE
  noise[Noise prior] --> AE
  AE -- 5 Euler steps --> A[Continuous action chunk]
  A -- 63 ms on H100 --> R[Real robot]
Loading

See the ฯ€0.6 model card (link below) for the authors' own architecture diagram.

Problem

A generalist robot policy needs strong visual-linguistic grounding and precise low-level control and sub-100 ms inference. Earlier generalist models either decoded actions autoregressively (too slow at billion-scale) or used ad-hoc MLP heads (no generative modeling of action distributions).

Method

Hierarchical VLA (preserving the ฯ€0.5 design): a high-level component predicts language-level subtasks; a low-level component generates continuous action chunks. The backbone is initialized from Gemma3-4B (with a SigLIP-400M vision encoder). The ~860M-parameter "action expert" is a parallel transformer expert with the same number of layers as the backbone (a mixture-of-experts arrangement, as in ฯ€0/ฯ€0.5 โ€” not an external MLP head), and is trained with flow matching โ€” it learns a continuous vector field that transports a noise distribution to action chunks, conditioned on backbone features; gradients from the action expert do not flow back into the VLM backbone (Knowledge Insulation). The backbone additionally predicts FAST discrete action tokens and co-training web data alongside the flow-matched continuous actions. At inference, 5 denoising (Euler) steps suffice.

Results

On a single H100 GPU with 3 cameras: ~63 ms per action chunk. Generalizes across multiple embodiments and real-world manipulation tasks. Deployed on Physical Intelligence's commercial and academic platforms.

Significance

The de facto production baseline for 2026 VLA research. Almost every ICLR 2026 VLA paper compares against ฯ€0, ฯ€0.5, or ฯ€0.6. Architectural legacy: established "VLM backbone + separate flow-matching action expert" as the dominant pattern โ€” the pattern that ICLR 2026's discrete-diffusion VLAs directly challenge (Discrete Diffusion VLA, Unified Diffusion VLA).

Now superseded by ฯ€0.7 (Apr 2026), which keeps the same backbone+expert and adds MEM history, subgoal-image world-model conditioning, expanded episode-metadata prompting (ฯ€0.6 already supports optional metadata conditioning), and suboptimal-data distillation.

Links

๐Ÿ“– In-depth review

For a long-form review with model-card figures, accuracy tables, Knowledge Insulation training details, limitations, and the ฯ€0.6 โ†’ ฯ€*0.6 โ†’ ฯ€0.7 delta: In-Depth Review of ฯ€0.6.

๐Ÿ“– In-depth architecture review

Compare the ฯ€-series' flow-matching action expert to other VLA architecture families: VLA Architectures Review.

Related pages

โ† Back to ICLR-2026

โš ๏ธ **GitHub.com Fallback** โš ๏ธ