PI pi07 - Heungwoo/research GitHub Wiki

ฯ€0.7 โ€” A Steerable Generalist Robotic Foundation Model with Emergent Capabilities

Venue: Physical Intelligence (technical report) ยท Date: April 16, 2026 Category: Baseline (VLA Architecture + Data) Trend tag: Baseline for Trend 1 (supersedes ฯ€0.6)

Approach diagram

flowchart LR
  subgraph Prompt[Multimodal Prompt Ct]
    L[Task instruction<br/>e.g. 'clean the kitchen']
    SL[Subtask โ„“ฬ‚<br/>e.g. 'open fridge']
    SG[Subgoal images gt<br/>multi-view]
    M[Metadata<br/>speed / quality 1-5 / mistake]
    C[Control mode<br/>joint or EE]
  end
  V[4 cameras ร— up to<br/>6 history frames<br/>448ร—448] --> MEM[MEM-style video<br/>history encoder]
  MEM --> B[Gemma3-4B<br/>VLM backbone]
  Prompt --> B
  PR[Proprioception qt<br/>linear projection] --> B
  B --> AE[860M flow-matching<br/>Action Expert<br/>50-step chunk]
  noise[Noise] --> AE
  AE -- 5 Euler steps --> A[Continuous action chunk]
  A -- RTC, 240ms tolerance --> R[Real robot]

  HL[High-level policy<br/>same arch] -.generates.-> SL
  WM[World model gฯˆ<br/>BAGEL-14B MoT] -.generates.-> SG
Loading

Problem

Prior VLAs (ฯ€0, ฯ€0.5, ฯ€0.6) were trained on high-quality demos with minimal prompt context โ€” just a task string. This produces generalists that need task-specific RL/SFT post-training (ฯ€*0.6) to match specialist performance, and that fail to compositionally generalize to unseen combinations (new appliances, new embodiments). Naively mixing failed/suboptimal/autonomous data into training just averages modes and degrades performance.

Method

Core idea: expand the training context Ct from a plain instruction to a rich multimodal prompt so the model can absorb heterogeneous data without mode collapse:

  • Subtask text โ„“ฬ‚ โ€” intermediate semantic step (inherited from ฯ€0.5).
  • Subgoal images โ€” multi-view near-future images, generated at runtime by a BAGEL-14B world model that is web-pretrained on image editing/video.
  • Episode metadata โ€” discretized speed, quality score 1โ€“5, mistake flag. At test time, always prompt "quality = 5, mistake = false" to steer toward high-quality modes.
  • Control mode โ€” joint vs. end-effector label. Each component is randomly dropped out at train time (subgoal kept in 25%, subtask in 70% of those, metadata dropped 15%), so any subset can be used at inference. Classifier-free guidance on metadata further amplifies high-quality behavior.

Architecture: ~5B total = Gemma3-4B VLM backbone + MEM video-history encoder (temporal+spatial compression, fixed token count regardless of #frames) + 860M flow-matching action expert (50-token chunk, 5 Euler steps). Inherits Knowledge Insulation from ฯ€0.6 (VLM trained on FAST tokens + web; action-expert gradients don't propagate back). Training-time Real-Time Chunking (RTC) simulates 0โ€“240ms inference delay for smooth deployment.

Data: the prompt-conditioning is what unlocks the real novelty โ€” heavy use of suboptimal data: lower-quality demos, failure episodes, autonomous rollouts from ฯ€*0.6 RL training, human egocentric video, open-source robot data (DROID etc.), and web multimodal (VQA, detection, captioning). Metadata labels disambiguate each episode's quality so the model can distill RL-specialist behavior rather than average it away.

Results

On the four ฯ€*0.6 RL-specialist tasks (laundry, diverse laundry, box building, espresso), a single general-purpose ฯ€0.7 matches or exceeds specialist success rate and throughput โ€” no task-specific post-training. Also matches SFT specialists on shirt-inside-out, PB sandwich, drive-through-door, slice zucchini, peel, trash. Zero-shot cross-embodiment: ฯ€0.7 folds t-shirts on a UR5e (never trained for laundry on UR5e) at the level of expert teleoperators โ€” 85.6% task progress / 80% success vs. 10 experienced teleoperators' 90.9% / 80.6% โ€” discovering embodiment-adapted strategies (e.g. single-arm or vertical grasps) rather than copying source behaviors. Compositional generalization: operates a never-seen air fryer (loading a sweet potato, unloading, toasting a bagel) with zero air-fryer training episodes in the robot data โ€” only similar appliances seen in human and external datasets โ€” by being coached through with step-by-step language. Ablations: removing metadata or eval-data both hurt โ€” diverse data and detailed context are synergistic.

Significance

Shifts the frontier of the ฯ€ series from "post-train per task" to "one prompt-steerable generalist." Three structural moves:

  1. Prompt as the integration surface โ€” instead of changing architecture for new capabilities, package them as extra prompt modalities (ร  la LLM prompt expansion).
  2. Suboptimal-data-as-a-feature โ€” RL rollouts and failures become training signal via metadata conditioning, sidestepping the ฯ€*0.6 + RECAP detour for many tasks.
  3. World-model subgoal images bring visual grounding into the prompt, complementing subtask text when language is ambiguous. Positions ฯ€0.7 as the new production baseline against which ICLR 2026's discrete-diffusion and memory VLAs will be compared, and gives the first ฯ€-series paper with strong claims of compositional (not just in-distribution) generalization.

Links

๐Ÿ“– In-depth review

For a long-form review with full results tables, all ablations, and limitations: In-Depth Review of ฯ€0.7.

Related pages

โ† Back to ICLR-2026

โš ๏ธ **GitHub.com Fallback** โš ๏ธ