Review pi07 - Heungwoo/research GitHub Wiki

In-Depth Review โ€” ฯ€0.7: A Steerable Generalist Robotic Foundation Model

Paper: ฯ€0.7: A Steerable Generalist Robotic Foundation Model with Emergent Capabilities ยท Physical Intelligence ยท April 16, 2026 Related summary: ฯ€0.7

This page is the long-form, numbers-heavy companion to the ฯ€0.7 summary. Everything here is sourced from the paper at https://www.pi.website/download/pi07.pdf (and the blog at https://www.pi.website/blog/pi07).

๐Ÿ“Ž ฯ€ series context โ€” read alongside this review

ฯ€0.7 is the latest step in a 2024โ€“2026 lineage. These pages give the context this review assumes:

Release Date Page Role in the lineage
ฯ€0 Oct 2024 (RSS 2025) โ€” First cross-embodiment flow-matching VLA; the architectural ancestor
ฯ€0.5 Apr 2025 (CoRL 2025 Oral) ฯ€0.5 Hierarchical subtask head + co-training โ€” the recipe ฯ€0.7 still inherits
ฯ€0.6 Nov 2025 ฯ€0.6 Gemma3-4B backbone + Knowledge Insulation + optional metadata
ฯ€*0.6 + RECAP Nov 2025 ฯ€*0.6 + RECAP RL-from-experience specialists; ฯ€0.7 distills these back into a generalist
ฯ€0.7 Apr 2026 this review / summary MEM history + subgoal-image world model + metadata prompting
(meta) โ€” ฯ€ series evolution ๐Ÿ“Š Side-by-side comparison table of model / data / training across ฯ€0 โ†’ ฯ€0.7

If you're reading this without the series context, start with ฯ€ series evolution first โ€” it takes 5 minutes and makes every section below more useful.


1. TL;DR

ฯ€0.7 is a steerable robotic foundation model that achieves out-of-the-box performance matching (and sometimes exceeding) task-specific RL-finetuned specialists โ€” without any task-specific post-training. The central idea: expand the prompt/context passed to the VLA from just a task string to a rich multimodal bundle (subtask text + subgoal images from a world model + episode metadata + control mode), each randomly dropped during training. This "prompt expansion" is what lets the model ingest heterogeneous/suboptimal/autonomous data (including ฯ€*0.6 RL rollouts, failures, human egocentric video) without averaging away the good modes. The result: the first ฯ€-series model with strong evidence of compositional generalization โ€” e.g., operating a never-before-seen air fryer from two tangential training episodes plus web pretraining priors.

2. Motivation โ€” what was broken in ฯ€0.6?

  • Forgetting a model's own specialists. ฯ€*0.6 + RECAP produced strong task-specific RL specialists (laundry, boxes, espresso). Those specialists sat in silos โ€” you couldn't collapse them into one generalist without losing performance.
  • Post-training hassle. For a new task, ฯ€0.6 still typically needed fine-tuning or at least metadata tuning. The field needed a generalist that worked zero-shot across a wider surface.
  • No convincing compositional generalization. Every prior ฯ€ release could do in-distribution tasks and some novel instruction-following, but could not remix skills to handle genuinely new tasks (e.g., use an unfamiliar kitchen appliance).
  • Suboptimal data was a liability. Mixing failures, RL rollouts, human video naively averages modes and degrades performance โ€” so most pipelines excluded them, throwing away signal.

ฯ€0.7 targets all four.

3. Representative diagrams

Figure 2 from the paper โ€” architecture overview

ฯ€0.7 architecture (Figure 2 from Physical Intelligence, 2026)

Figure 2 of the ฯ€0.7 paper (Physical Intelligence, Apr 2026). The 5B-parameter VLA = Gemma3-4B VLM backbone (including a 400M vision encoder) + MEM-style video-history encoder + 860M flow-matching action expert. At runtime, the subtask instruction is produced by a high-level semantic policy based on the same architecture; the subgoal images are produced by a lightweight BAGEL-14B world model. Included for scholarly review.

Figure 3 from the paper โ€” prompt composition

ฯ€0.7 prompt overview (Figure 3 from Physical Intelligence, 2026)

Figure 3 of the ฯ€0.7 paper. ฯ€0.7 uses diverse modalities of context in the prompt: subtask instructions, subgoal images, and episode metadata (speed / quality / mistake). Each component is trained with dropout so any subset can be provided at test time โ€” e.g., for UR5e bimanual shirt folding the authors use subgoal image + metadata prompting (bottom row).

Our reconstruction as mermaid

flowchart TB
  subgraph Prompt[Multimodal prompt Ct โ€” each component dropped independently]
    L[Task โ„“<br/>'peel vegetables']
    SL[Subtask โ„“ฬ‚<br/>'pick up the peeler']
    SG[Subgoal images gt<br/>multi-view, web-pretrained<br/>world model]
    M[Metadata<br/>speed / quality 1-5 / mistake]
    C[Control mode<br/>joint or end-effector]
  end

  V[4 cameras ร— up to 6 history frames<br/>448ร—448] --> MEM[MEM video-history encoder<br/>temporal+spatial compression]
  MEM --> B[Gemma3-4B VLM backbone]
  Prompt --> B
  PR[Proprioception qt<br/>linear projection] --> B

  B --> AE[860M flow-matching<br/>Action Expert<br/>50-token chunk]
  noise[Noise z] --> AE
  AE -- 5 Euler steps + RTC --> A[Action chunk]
  A --> R[Real robot @ 50 Hz / 20 Hz]

  HL[High-level policy<br/>same architecture] -.generates.-> SL
  WM[World model gฯˆ<br/>BAGEL-14B MoT,<br/>web pretrained] -.generates.-> SG

  classDef new fill:#ffe8c2,stroke:#b47820,color:#000
  class MEM,SG,M,WM,AE new
Loading

Orange boxes = changes relative to ฯ€0.6. The VLM backbone + flow-matching action expert are unchanged; all new capability is routed through the prompt and the history/subgoal channels.

4. Method โ€” full recipe

4.1 Architecture

  • VLM backbone. Gemma3-4B (same as ฯ€0.6), including a 400M SigLIP vision encoder.
  • History encoder (NEW). MEM-style video encoder. Takes up to 4 cameras (base + 2 wrist + optional rear) ร— up to 6 history frames at 448ร—448, plus up to 3 subgoal images (rear view omitted) processed through the same encoder; compresses spatially and temporally to a single frame's worth of tokens regardless of # frames. History stride = 1 s; full history dropped with p = 0.3, rear view dropped with p = 0.3.
  • Proprioception (NEW). Unlike ฯ€0.6 which used discretized text tokens, ฯ€0.7 embeds the robot state via a linear projection to the backbone dimension; each history-state gets its own token.
  • Action expert. 860M flow-matching transformer, 50-token chunk. Uses adaptive RMSNorm for timestep injection. Action tokens attend bidirectionally among themselves and can attend all VLM activations.
  • Attention (conditional, per Appendix B). ฯ€0.7 uses different attention patterns depending on what's in the prompt:
    • No image goals in prompt โ†’ falls back to the same global bidirectional attention as ฯ€0.5 (image and text embeddings all attend each other bidirectionally). This is the important correction over earlier wiki notes that claimed ฯ€0.6 already used causal-on-text โ€” Appendix B explicitly says "in absence of image goals we use the same attention patterns as in ฯ€0.5, with global bidirectional attention between embeddings for all".
    • With image goals (subgoal images) โ†’ switches to block-causal: observation tokens and subgoal-image tokens are bidirectional within themselves; subgoal-image tokens can additionally attend observations; the following text tokens (subtask, metadata) use causal attention. This is the change motivated by the new prompt modalities (subgoal images, metadata) โ€” bidirectional everywhere would let metadata leak across components and clash with the per-component dropout + CFG-on-metadata recipe.
    • Action tokens (50 tokens, in both modes) attend bidirectionally among themselves and can attend all VLM activations.
  • Total parameters. ~5B (ฯ€0.7 policy) + 14B (BAGEL world model) + small high-level policy.

4.2 Knowledge Insulation (inherited from ฯ€0.6)

The VLM backbone is supervised with FAST tokens and web-style co-training; the action expert attends all VLM activations but gradients from the action expert do not flow back into the VLM. This keeps the VLM training signal stable (discrete cross-entropy) while the action expert learns continuous targets via flow matching.

4.3 Training objective (Equation 1)

$$\max_\theta \mathbb{E}{\mathcal D}, \log \pi\theta(a_{t:t+H} \mid o_{t-T:t}, C_t)$$ with $C_t$ = (task โ„“, subtask โ„“ฬ‚, subgoal images g, metadata m, control mode c). Flow matching optimizes an approximate lower bound rather than exact log-likelihood.

4.4 The key trick โ€” multimodal context + dropout

Each component of $C_t$ is trained with independent dropout so the model handles any subset at test time:

  • Subgoal images included in 25% of training examples; of those, 30% also drop the subtask text (subgoal alone substitutes for it).
  • Metadata dropped entirely 15% of the time; additionally each field (speed/quality/mistake) dropped with 5% probability.
  • Subtask text dropped per data source's annotation coverage.
  • Control mode never dropped (always set).

At inference, a default prompt is Quality: 5, Mistake: false, Speed: 15th-percentile-per-task, plus subtask from the high-level policy and (optionally) generated subgoal images.

4.5 Classifier-Free Guidance (CFG) on metadata

Apply CFG at each denoising step with unconditional prompt = metadata dropped: $$\nabla_a \log \pi_\theta(a|o,C) + \beta\left(\nabla_a \log \pi_\theta(a|o,C) - \nabla_a \log \pi_\theta(a|o,C^{\text{uncond}})\right)$$ with ฮฒ โˆˆ {1.3, 1.7, 2.2}. Amplifies the effect of "quality = 5" conditioning, steering the model toward higher-quality modes.

4.6 Real-Time Chunking (RTC)

During training, simulate inference delays of 0โ€“12 timesteps (up to 240 ms on a 50 Hz robot). At runtime, inference runs async; new chunks overlap the previous chunk's tail to handle latency and produce smooth trajectories.

4.7 World model for subgoal images

A separate model gฯˆ โ€” initialized from BAGEL-14B, a mixture-of-transformers image editing/generation model with web-scale pretraining โ€” is trained with flow-matching to predict future subgoal images given the current observation and subtask. Because BAGEL is web-pretrained, the world model can generate plausible subgoals for unseen appliances (air fryer, bagel toaster) and feed them back into ฯ€0.7's prompt. Trained on the subset of data with high-quality subtask labels; combined with sampling from real future frames (25% end-of-segment, 75% uniform 0โ€“4 s ahead) to mitigate train/test mismatch.

4.8 Training data mix

  • Robot demonstrations: multiple robot platforms (static / mobile / bimanual / UR5e)
  • Autonomous rollouts โ€” policy evaluation episodes, including ฯ€*0.6 RL-specialist rollouts (this is new for the ฯ€ series).
  • Failures and lower-quality demos.
  • Human interventions during policy rollouts.
  • Open-source: OXE + DROID (Franka arms).
  • Egocentric human video.
  • Web multimodal: VQA, detection, captioning, text-only.

5. Experimental setup

5.1 Platforms

Platform Type Rate
Static bimanual (standard ฯ€ platform) 6-DoF ร— 2 50 Hz
BiPi (lightweight static bimanual) 6-DoF ร— 2 50 Hz
Mobile bimanual manipulator 6-DoF ร— 2 + base 50 Hz
Bimanual UR5e (cross-embodiment) 6-DoF ร— 2, longer/heavier arms, Robotiq grippers 20 Hz

5.2 Baselines

  • ฯ€0.5, ฯ€0.6 โ€” prior ฯ€ series.
  • ฯ€*0.6 RL specialists โ€” per-task RL-finetuned policies from Physical Intelligence's November 2025 release.
  • ฯ€0.6 SFT specialists โ€” task-specific supervised fine-tunes.
  • ฯ€0.7 ablations โ€” no metadata, no eval data (excludes autonomous rollouts from training), w/o most diverse 20%, w/o random 20%.

5.3 Inference settings

5 Euler denoising steps; execute ฤค โˆˆ {15, 25} of 50 action tokens per chunk; async subgoal refresh every 4 s or when subtask changes.


6. Results โ€” every number the paper reports

6.1 Out-of-the-box dexterity vs. RL / SFT specialists

On the four tasks ฯ€*0.6 released with, a single ฯ€0.7 matches or exceeds per-task specialists with no post-training:

Task Specialist baseline ฯ€0.7 outcome
Laundry (T-shirts + shorts) ฯ€*0.6 RL specialist Matches success rate; higher normalized throughput
Laundry (diverse, hardest items) ฯ€*0.6 RL specialist Matches success rate; higher throughput
Box building ฯ€*0.6 RL specialist Matches success rate; higher throughput
Make espresso ฯ€*0.6 RL specialist Matches success rate; matches throughput

Additional dexterous tasks vs. ฯ€0.6 SFT specialists (ฯ€0.7 roughly matches each):

  • Shirt inside-out ยท Make PB sandwich ยท Drive-through door
  • Slice zucchini ยท Peel fruits & vegetables ยท Take out trash

6.2 Memory-requiring tasks (vs. ฯ€0.6-MEM SFT specialists)

Out-of-the-box, no fine-tuning:

  • Swap 3 mugs ยท Find object ยท Scoop coffee ยท Window cleaning

ฯ€0.7 matches or exceeds the MEM-paper specialists on all four.

6.3 Instruction following in unseen environments

14 instruction-following scenarios across 4 unseen kitchens + 2 unseen bedrooms; each scenario = 3โ€“6 open-ended instructions. ฯ€0.7 significantly outperforms ฯ€0.5 and ฯ€0.6 on overall per-instruction success rate. Authors report "high absolute success rates" but do not publish a single headline percentage.

On the Office Desk Object Rearrangement suite split into standard and complex instructions:

  • Standard instructions ("pick up the spoon"): all models succeed.
  • Complex/referential instructions ("pick up the object I would use to eat soup", "pick up the fruit on the largest plate"): ฯ€0.7 significantly outperforms ฯ€0.5/ฯ€0.6; adding subgoal images (ฯ€0.7 GC) further boosts performance.

6.4 Breaking dataset bias

Two adversarial tasks where the correct instruction is the opposite of the dataset bias:

  • Reverse Bussing: put trash in the dish bin, dishes in the trash. ฯ€0.7 โ‰ซ ฯ€0.5/ฯ€0.6.
  • Reverse Fridge โ†’ Microwave: take food out of the microwave and put it in the fridge (reverse of training data). ฯ€0.7 (GC) with world-model subgoals is critical for success โ€” language alone doesn't fully break the bias.

6.5 Cross-embodiment transfer (zero-shot)

Rearrangement suite (bigger embodiment gap โ†’ ฯ€0.5 collapses, ฯ€0.6 holds, ฯ€0.7 leads):

Task Source robot Target robot Gap Winner
Table Setting Many robots (mobile/static/single-arm) Static bimanual Small All methods strong
Bag In Backpack UR5e bimanual (larger, heavier) Static bimanual (smaller) Medium ฯ€0.5 collapses, ฯ€0.6 holds, ฯ€0.7 strongest
Organize Tupperware UR5e bimanual Static bimanual Medium Same โ€” ฯ€0.7 โ‰ฅ ฯ€0.6 โ‰ซ ฯ€0.5
Shirt Bagging Static bimanual (smaller) Single-arm UR5e (bigger, heavier) Large (one arm!) ฯ€0.7 significantly outperforms prior models

Laundry folding โ€” the headline cross-embodiment result: Training data for folding was almost entirely from the lightweight static bimanual. Tested on the bimanual UR5e (never seen folding training data). ฯ€0.7 adapts its strategy โ€” uses vertical grasps instead of the source robot's tilted grasps โ€” because UR5e kinematics make tilts bad.

Human subject study for context: 10 expert teleoperators, mean 375 hours of teleop experience across robots (top 2% by experience). They attempted shirt folding on UR5e for the first time (same zero-shot setting as ฯ€0.7):

Agent Task progress Success rate
Expert human teleoperators (first attempt on UR5e) 90.9% 80.6%
ฯ€0.7 (zero-shot on UR5e) 85.6% 80.0%

ฯ€0.7 matches expert operators on their first attempt โ€” striking evidence of cross-embodiment transfer.

6.6 Compositional generalization

New short-horizon tasks, no training data collected for any of them, out-of-the-box:

  • Pressing a French press plunger
  • Scooping rice into a rice cooker
  • Wiping office objects (ruler, headphones) with a cloth
  • Spinning articulated items (gear set, desk fan)

ฯ€0.7 performs at roughly equal levels when prompted by language or by generated subgoal images.

New long-horizon tasks via language coaching:

  • Loading an air fryer (cook a sweet potato)
  • Unloading an air fryer
  • Toasting a bagel

Authors walk the robot through step-by-step ("grasp the air fryer handle", "pick up the sweet potato", "open the air fryer", โ€ฆ). Prior models fail these because they don't follow coaching; ฯ€0.7 succeeds. The paper states the robot data contained no training episodes for these tasks, "although similar appliances were seen in different contexts in human data and external datasets" โ€” i.e., the capability comes from remixing related skills plus web-pretraining priors rather than task-specific demonstrations.

Coaching โ†’ autonomous policy: Use the coaching episodes as data to train the high-level subtask policy; ฯ€0.7 then performs the full task autonomously. Tested on 5 tasks:

  • Scoop rice into rice cooker
  • Reverse fridge โ†’ microwave
  • Press French-press plunger
  • Wiping office supplies
  • Spinning articulated objects

Autonomous version โ‰ˆ coaching version on all 5 โ€” new autonomous capability acquired without any teleop data collected for the task.


7. Ablations โ€” proving the recipe matters

7.1 Prompt components (Fig. 7)

On ฯ€*0.6-release tasks:

  • ฯ€0.7 (no metadata): drops episode metadata from context โ†’ significant throughput and success drops across the board.
  • ฯ€0.7 (no eval data): excludes autonomous rollouts from training โ†’ significant drops.

Interpretation: both metadata and heterogeneous data matter โ€” neither is sufficient alone.

7.2 Scaling with mixed-quality data (Fig. 18, left)

Laundry (T-shirts+shorts) split into 4 buckets by annotated quality & speed: top 30% / top 50% / top 80% / all data. For each bucket, train ฯ€0.7 with and without metadata (8 models):

  • ฯ€0.7 without metadata: performance gets worse as more lower-quality data is added.
  • ฯ€0.7 with metadata: performance continues to improve even as average data quality drops.

This is the single strongest result in the paper โ€” it validates the whole thesis that metadata is what makes scaling on suboptimal data work.

7.3 Scaling with diverse data (Fig. 18, right)

On short-horizon unseen tasks from ยง6.6:

  • ฯ€0.7 (w/o most diverse 20%): remove the 20% of data with highest task diversity.
  • ฯ€0.7 (w/o random 20%): random 20% removed (data-controlled comparison).

ฯ€0.7 โ‰ˆ ฯ€0.7 (w/o random 20%) โ‰ซ ฯ€0.7 (w/o most diverse 20%). Diverse data specifically, not just more data, drives compositional generalization.

7.4 End-effector vs. joint control (Appendix E)

End-effector control does not give noticeable gains over joint control with prior models; authors use joint-space control throughout.


8. Limitations (authors' own + critical reading)

8.1 Authors' stated limitations

  1. Zero-shot success rates still 60โ€“80% on unseen tasks, vs. >90% on in-distribution. Not deployment-ready for new tasks without at least coaching.
  2. "Seen vs. unseen" is ill-defined on large diverse datasets. With millions of episodes, "never seen" is nearly unfalsifiable โ€” the model may be achieving generalization primarily by remixing seen skills. The authors concede this but argue remixing is compositional generalization.
  3. Coaching and subgoal images still require a human to steer in the long-horizon cases โ€” fully autonomous long-horizon execution requires the follow-up high-level policy training step.
  4. Laundry comparisons are per-task throughput-normalized, not raw throughput โ€” raw numbers not released.

8.2 Reviewer's concerns worth flagging (not in the paper)

  • No head-to-head against ICLR 2026 discrete-diffusion VLAs. The structural alternative (Discrete Diffusion VLA, Unified Diffusion VLA) is not evaluated here.
  • BAGEL-14B is a large module in the inference path; the paper doesn't break down latency impact. Subgoal refresh is asynchronous (every 4 s), so it's not on the critical path per chunk, but VRAM cost is 3ร— the policy.
  • Metadata quality annotation is human-labeled and coarse. Scaling the metadata-conditioning thesis further would require either annotator models or self-labeling.
  • Data-leak risk with distillation of ฯ€*0.6 rollouts. ฯ€0.7's "match RL specialists" claim could partly reflect direct distillation of those specialists' trajectories, not emergent generalization. The paper acknowledges this as "distillation of experience" but doesn't cleanly decompose the effect.
  • Reproducibility. No open weights, no training code release. Like prior ฯ€ papers, this is a technical report more than an independently reproducible artifact.
  • Compositional generalization claim is demonstrational, not statistical. Specific anecdotes (air fryer) are compelling, but the paper does not report success rates on a large held-out compositional-task suite.

9. Takeaways โ€” what ฯ€0.7 changes about the field

  1. Prompt is the new integration surface. Instead of new architectures, package new capability as new prompt modalities with independent dropout. ฯ€0.7's architectural delta over ฯ€0.6 is small; the recipe delta is huge.
  2. Suboptimal data becomes signal. With metadata conditioning, failures and RL rollouts become useful training examples โ€” inverting the decade-old data-curation default.
  3. World models move inside the prompt. Subgoal-image generation from BAGEL is a cheaper form of "world model as teacher" than DreamGen-style in-imagination RL, but shares the thesis that web-pretrained generative priors have direct use in robot policies.
  4. First strong compositional generalization signal in the ฯ€ series. Up until ฯ€0.7, the series gave in-distribution dexterity and some language following; ฯ€0.7 is the first release with qualitatively new zero-shot capability (operating unfamiliar kitchen appliances).
  5. The RL-vs-BC debate softens. ฯ€*0.6 RL specialists are now being distilled back into a BC generalist via the autonomous-data pathway; RL becomes an intermediate teacher rather than a deployment mechanism.

10. Open questions

  • How does ฯ€0.7 compare directly to the ICLR 2026 unified discrete-diffusion line?
  • Does prompt-expansion scale further โ€” can you add, say, force-torque traces, acoustic cues, per-step rewards as new prompt modalities without destabilizing training?
  • Can the metadata annotator be self-trained (model labels its own episodes) to remove human-annotation as a bottleneck?
  • What's the minimum BAGEL-scale world model that retains the subgoal-image benefit? Is 14B overkill?
  • Is there a principled way to decompose "distillation of RL specialists" from "true compositional generalization" in the evaluation metrics?

11. Links & references

โ† Back to PI-pi07 ยท ICLR-2026 ยท Home

โš ๏ธ **GitHub.com Fallback** โš ๏ธ