Review pi07 - Heungwoo/research GitHub Wiki
Paper: ฯ0.7: A Steerable Generalist Robotic Foundation Model with Emergent Capabilities ยท Physical Intelligence ยท April 16, 2026 Related summary: ฯ0.7
This page is the long-form, numbers-heavy companion to the ฯ0.7 summary. Everything here is sourced from the paper at https://www.pi.website/download/pi07.pdf (and the blog at https://www.pi.website/blog/pi07).
ฯ0.7 is the latest step in a 2024โ2026 lineage. These pages give the context this review assumes:
| Release | Date | Page | Role in the lineage |
|---|---|---|---|
| ฯ0 | Oct 2024 (RSS 2025) | โ | First cross-embodiment flow-matching VLA; the architectural ancestor |
| ฯ0.5 | Apr 2025 (CoRL 2025 Oral) | ฯ0.5 | Hierarchical subtask head + co-training โ the recipe ฯ0.7 still inherits |
| ฯ0.6 | Nov 2025 | ฯ0.6 | Gemma3-4B backbone + Knowledge Insulation + optional metadata |
| ฯ*0.6 + RECAP | Nov 2025 | ฯ*0.6 + RECAP | RL-from-experience specialists; ฯ0.7 distills these back into a generalist |
| ฯ0.7 | Apr 2026 | this review / summary | MEM history + subgoal-image world model + metadata prompting |
| (meta) | โ | ฯ series evolution | ๐ Side-by-side comparison table of model / data / training across ฯ0 โ ฯ0.7 |
If you're reading this without the series context, start with ฯ series evolution first โ it takes 5 minutes and makes every section below more useful.
ฯ0.7 is a steerable robotic foundation model that achieves out-of-the-box performance matching (and sometimes exceeding) task-specific RL-finetuned specialists โ without any task-specific post-training. The central idea: expand the prompt/context passed to the VLA from just a task string to a rich multimodal bundle (subtask text + subgoal images from a world model + episode metadata + control mode), each randomly dropped during training. This "prompt expansion" is what lets the model ingest heterogeneous/suboptimal/autonomous data (including ฯ*0.6 RL rollouts, failures, human egocentric video) without averaging away the good modes. The result: the first ฯ-series model with strong evidence of compositional generalization โ e.g., operating a never-before-seen air fryer from two tangential training episodes plus web pretraining priors.
- Forgetting a model's own specialists. ฯ*0.6 + RECAP produced strong task-specific RL specialists (laundry, boxes, espresso). Those specialists sat in silos โ you couldn't collapse them into one generalist without losing performance.
- Post-training hassle. For a new task, ฯ0.6 still typically needed fine-tuning or at least metadata tuning. The field needed a generalist that worked zero-shot across a wider surface.
- No convincing compositional generalization. Every prior ฯ release could do in-distribution tasks and some novel instruction-following, but could not remix skills to handle genuinely new tasks (e.g., use an unfamiliar kitchen appliance).
- Suboptimal data was a liability. Mixing failures, RL rollouts, human video naively averages modes and degrades performance โ so most pipelines excluded them, throwing away signal.
ฯ0.7 targets all four.

Figure 2 of the ฯ0.7 paper (Physical Intelligence, Apr 2026). The 5B-parameter VLA = Gemma3-4B VLM backbone (including a 400M vision encoder) + MEM-style video-history encoder + 860M flow-matching action expert. At runtime, the subtask instruction is produced by a high-level semantic policy based on the same architecture; the subgoal images are produced by a lightweight BAGEL-14B world model. Included for scholarly review.

Figure 3 of the ฯ0.7 paper. ฯ0.7 uses diverse modalities of context in the prompt: subtask instructions, subgoal images, and episode metadata (speed / quality / mistake). Each component is trained with dropout so any subset can be provided at test time โ e.g., for UR5e bimanual shirt folding the authors use subgoal image + metadata prompting (bottom row).
flowchart TB
subgraph Prompt[Multimodal prompt Ct โ each component dropped independently]
L[Task โ<br/>'peel vegetables']
SL[Subtask โฬ<br/>'pick up the peeler']
SG[Subgoal images gt<br/>multi-view, web-pretrained<br/>world model]
M[Metadata<br/>speed / quality 1-5 / mistake]
C[Control mode<br/>joint or end-effector]
end
V[4 cameras ร up to 6 history frames<br/>448ร448] --> MEM[MEM video-history encoder<br/>temporal+spatial compression]
MEM --> B[Gemma3-4B VLM backbone]
Prompt --> B
PR[Proprioception qt<br/>linear projection] --> B
B --> AE[860M flow-matching<br/>Action Expert<br/>50-token chunk]
noise[Noise z] --> AE
AE -- 5 Euler steps + RTC --> A[Action chunk]
A --> R[Real robot @ 50 Hz / 20 Hz]
HL[High-level policy<br/>same architecture] -.generates.-> SL
WM[World model gฯ<br/>BAGEL-14B MoT,<br/>web pretrained] -.generates.-> SG
classDef new fill:#ffe8c2,stroke:#b47820,color:#000
class MEM,SG,M,WM,AE new
Orange boxes = changes relative to ฯ0.6. The VLM backbone + flow-matching action expert are unchanged; all new capability is routed through the prompt and the history/subgoal channels.
- VLM backbone. Gemma3-4B (same as ฯ0.6), including a 400M SigLIP vision encoder.
- History encoder (NEW). MEM-style video encoder. Takes up to 4 cameras (base + 2 wrist + optional rear) ร up to 6 history frames at 448ร448, plus up to 3 subgoal images (rear view omitted) processed through the same encoder; compresses spatially and temporally to a single frame's worth of tokens regardless of # frames. History stride = 1 s; full history dropped with p = 0.3, rear view dropped with p = 0.3.
- Proprioception (NEW). Unlike ฯ0.6 which used discretized text tokens, ฯ0.7 embeds the robot state via a linear projection to the backbone dimension; each history-state gets its own token.
- Action expert. 860M flow-matching transformer, 50-token chunk. Uses adaptive RMSNorm for timestep injection. Action tokens attend bidirectionally among themselves and can attend all VLM activations.
-
Attention (conditional, per Appendix B). ฯ0.7 uses different attention patterns depending on what's in the prompt:
- No image goals in prompt โ falls back to the same global bidirectional attention as ฯ0.5 (image and text embeddings all attend each other bidirectionally). This is the important correction over earlier wiki notes that claimed ฯ0.6 already used causal-on-text โ Appendix B explicitly says "in absence of image goals we use the same attention patterns as in ฯ0.5, with global bidirectional attention between embeddings for all".
- With image goals (subgoal images) โ switches to block-causal: observation tokens and subgoal-image tokens are bidirectional within themselves; subgoal-image tokens can additionally attend observations; the following text tokens (subtask, metadata) use causal attention. This is the change motivated by the new prompt modalities (subgoal images, metadata) โ bidirectional everywhere would let metadata leak across components and clash with the per-component dropout + CFG-on-metadata recipe.
- Action tokens (50 tokens, in both modes) attend bidirectionally among themselves and can attend all VLM activations.
- Total parameters. ~5B (ฯ0.7 policy) + 14B (BAGEL world model) + small high-level policy.
The VLM backbone is supervised with FAST tokens and web-style co-training; the action expert attends all VLM activations but gradients from the action expert do not flow back into the VLM. This keeps the VLM training signal stable (discrete cross-entropy) while the action expert learns continuous targets via flow matching.
$$\max_\theta \mathbb{E}{\mathcal D}, \log \pi\theta(a_{t:t+H} \mid o_{t-T:t}, C_t)$$
with
Each component of
- Subgoal images included in 25% of training examples; of those, 30% also drop the subtask text (subgoal alone substitutes for it).
- Metadata dropped entirely 15% of the time; additionally each field (speed/quality/mistake) dropped with 5% probability.
- Subtask text dropped per data source's annotation coverage.
- Control mode never dropped (always set).
At inference, a default prompt is Quality: 5, Mistake: false, Speed: 15th-percentile-per-task, plus subtask from the high-level policy and (optionally) generated subgoal images.
Apply CFG at each denoising step with unconditional prompt = metadata dropped:
During training, simulate inference delays of 0โ12 timesteps (up to 240 ms on a 50 Hz robot). At runtime, inference runs async; new chunks overlap the previous chunk's tail to handle latency and produce smooth trajectories.
A separate model gฯ โ initialized from BAGEL-14B, a mixture-of-transformers image editing/generation model with web-scale pretraining โ is trained with flow-matching to predict future subgoal images given the current observation and subtask. Because BAGEL is web-pretrained, the world model can generate plausible subgoals for unseen appliances (air fryer, bagel toaster) and feed them back into ฯ0.7's prompt. Trained on the subset of data with high-quality subtask labels; combined with sampling from real future frames (25% end-of-segment, 75% uniform 0โ4 s ahead) to mitigate train/test mismatch.
- Robot demonstrations: multiple robot platforms (static / mobile / bimanual / UR5e)
- Autonomous rollouts โ policy evaluation episodes, including ฯ*0.6 RL-specialist rollouts (this is new for the ฯ series).
- Failures and lower-quality demos.
- Human interventions during policy rollouts.
- Open-source: OXE + DROID (Franka arms).
- Egocentric human video.
- Web multimodal: VQA, detection, captioning, text-only.
| Platform | Type | Rate |
|---|---|---|
| Static bimanual (standard ฯ platform) | 6-DoF ร 2 | 50 Hz |
| BiPi (lightweight static bimanual) | 6-DoF ร 2 | 50 Hz |
| Mobile bimanual manipulator | 6-DoF ร 2 + base | 50 Hz |
| Bimanual UR5e (cross-embodiment) | 6-DoF ร 2, longer/heavier arms, Robotiq grippers | 20 Hz |
- ฯ0.5, ฯ0.6 โ prior ฯ series.
- ฯ*0.6 RL specialists โ per-task RL-finetuned policies from Physical Intelligence's November 2025 release.
- ฯ0.6 SFT specialists โ task-specific supervised fine-tunes.
-
ฯ0.7 ablations โ
no metadata,no eval data(excludes autonomous rollouts from training),w/o most diverse 20%,w/o random 20%.
5 Euler denoising steps; execute ฤค โ {15, 25} of 50 action tokens per chunk; async subgoal refresh every 4 s or when subtask changes.
On the four tasks ฯ*0.6 released with, a single ฯ0.7 matches or exceeds per-task specialists with no post-training:
| Task | Specialist baseline | ฯ0.7 outcome |
|---|---|---|
| Laundry (T-shirts + shorts) | ฯ*0.6 RL specialist | Matches success rate; higher normalized throughput |
| Laundry (diverse, hardest items) | ฯ*0.6 RL specialist | Matches success rate; higher throughput |
| Box building | ฯ*0.6 RL specialist | Matches success rate; higher throughput |
| Make espresso | ฯ*0.6 RL specialist | Matches success rate; matches throughput |
Additional dexterous tasks vs. ฯ0.6 SFT specialists (ฯ0.7 roughly matches each):
- Shirt inside-out ยท Make PB sandwich ยท Drive-through door
- Slice zucchini ยท Peel fruits & vegetables ยท Take out trash
Out-of-the-box, no fine-tuning:
- Swap 3 mugs ยท Find object ยท Scoop coffee ยท Window cleaning
ฯ0.7 matches or exceeds the MEM-paper specialists on all four.
14 instruction-following scenarios across 4 unseen kitchens + 2 unseen bedrooms; each scenario = 3โ6 open-ended instructions. ฯ0.7 significantly outperforms ฯ0.5 and ฯ0.6 on overall per-instruction success rate. Authors report "high absolute success rates" but do not publish a single headline percentage.
On the Office Desk Object Rearrangement suite split into standard and complex instructions:
- Standard instructions ("pick up the spoon"): all models succeed.
- Complex/referential instructions ("pick up the object I would use to eat soup", "pick up the fruit on the largest plate"): ฯ0.7 significantly outperforms ฯ0.5/ฯ0.6; adding subgoal images (ฯ0.7 GC) further boosts performance.
Two adversarial tasks where the correct instruction is the opposite of the dataset bias:
- Reverse Bussing: put trash in the dish bin, dishes in the trash. ฯ0.7 โซ ฯ0.5/ฯ0.6.
- Reverse Fridge โ Microwave: take food out of the microwave and put it in the fridge (reverse of training data). ฯ0.7 (GC) with world-model subgoals is critical for success โ language alone doesn't fully break the bias.
Rearrangement suite (bigger embodiment gap โ ฯ0.5 collapses, ฯ0.6 holds, ฯ0.7 leads):
| Task | Source robot | Target robot | Gap | Winner |
|---|---|---|---|---|
| Table Setting | Many robots (mobile/static/single-arm) | Static bimanual | Small | All methods strong |
| Bag In Backpack | UR5e bimanual (larger, heavier) | Static bimanual (smaller) | Medium | ฯ0.5 collapses, ฯ0.6 holds, ฯ0.7 strongest |
| Organize Tupperware | UR5e bimanual | Static bimanual | Medium | Same โ ฯ0.7 โฅ ฯ0.6 โซ ฯ0.5 |
| Shirt Bagging | Static bimanual (smaller) | Single-arm UR5e (bigger, heavier) | Large (one arm!) | ฯ0.7 significantly outperforms prior models |
Laundry folding โ the headline cross-embodiment result: Training data for folding was almost entirely from the lightweight static bimanual. Tested on the bimanual UR5e (never seen folding training data). ฯ0.7 adapts its strategy โ uses vertical grasps instead of the source robot's tilted grasps โ because UR5e kinematics make tilts bad.
Human subject study for context: 10 expert teleoperators, mean 375 hours of teleop experience across robots (top 2% by experience). They attempted shirt folding on UR5e for the first time (same zero-shot setting as ฯ0.7):
| Agent | Task progress | Success rate |
|---|---|---|
| Expert human teleoperators (first attempt on UR5e) | 90.9% | 80.6% |
| ฯ0.7 (zero-shot on UR5e) | 85.6% | 80.0% |
ฯ0.7 matches expert operators on their first attempt โ striking evidence of cross-embodiment transfer.
New short-horizon tasks, no training data collected for any of them, out-of-the-box:
- Pressing a French press plunger
- Scooping rice into a rice cooker
- Wiping office objects (ruler, headphones) with a cloth
- Spinning articulated items (gear set, desk fan)
ฯ0.7 performs at roughly equal levels when prompted by language or by generated subgoal images.
New long-horizon tasks via language coaching:
- Loading an air fryer (cook a sweet potato)
- Unloading an air fryer
- Toasting a bagel
Authors walk the robot through step-by-step ("grasp the air fryer handle", "pick up the sweet potato", "open the air fryer", โฆ). Prior models fail these because they don't follow coaching; ฯ0.7 succeeds. The paper states the robot data contained no training episodes for these tasks, "although similar appliances were seen in different contexts in human data and external datasets" โ i.e., the capability comes from remixing related skills plus web-pretraining priors rather than task-specific demonstrations.
Coaching โ autonomous policy: Use the coaching episodes as data to train the high-level subtask policy; ฯ0.7 then performs the full task autonomously. Tested on 5 tasks:
- Scoop rice into rice cooker
- Reverse fridge โ microwave
- Press French-press plunger
- Wiping office supplies
- Spinning articulated objects
Autonomous version โ coaching version on all 5 โ new autonomous capability acquired without any teleop data collected for the task.
On ฯ*0.6-release tasks:
- ฯ0.7 (no metadata): drops episode metadata from context โ significant throughput and success drops across the board.
- ฯ0.7 (no eval data): excludes autonomous rollouts from training โ significant drops.
Interpretation: both metadata and heterogeneous data matter โ neither is sufficient alone.
Laundry (T-shirts+shorts) split into 4 buckets by annotated quality & speed: top 30% / top 50% / top 80% / all data. For each bucket, train ฯ0.7 with and without metadata (8 models):
- ฯ0.7 without metadata: performance gets worse as more lower-quality data is added.
- ฯ0.7 with metadata: performance continues to improve even as average data quality drops.
This is the single strongest result in the paper โ it validates the whole thesis that metadata is what makes scaling on suboptimal data work.
On short-horizon unseen tasks from ยง6.6:
- ฯ0.7 (w/o most diverse 20%): remove the 20% of data with highest task diversity.
- ฯ0.7 (w/o random 20%): random 20% removed (data-controlled comparison).
ฯ0.7 โ ฯ0.7 (w/o random 20%) โซ ฯ0.7 (w/o most diverse 20%). Diverse data specifically, not just more data, drives compositional generalization.
End-effector control does not give noticeable gains over joint control with prior models; authors use joint-space control throughout.
- Zero-shot success rates still 60โ80% on unseen tasks, vs. >90% on in-distribution. Not deployment-ready for new tasks without at least coaching.
- "Seen vs. unseen" is ill-defined on large diverse datasets. With millions of episodes, "never seen" is nearly unfalsifiable โ the model may be achieving generalization primarily by remixing seen skills. The authors concede this but argue remixing is compositional generalization.
- Coaching and subgoal images still require a human to steer in the long-horizon cases โ fully autonomous long-horizon execution requires the follow-up high-level policy training step.
- Laundry comparisons are per-task throughput-normalized, not raw throughput โ raw numbers not released.
- No head-to-head against ICLR 2026 discrete-diffusion VLAs. The structural alternative (Discrete Diffusion VLA, Unified Diffusion VLA) is not evaluated here.
- BAGEL-14B is a large module in the inference path; the paper doesn't break down latency impact. Subgoal refresh is asynchronous (every 4 s), so it's not on the critical path per chunk, but VRAM cost is 3ร the policy.
- Metadata quality annotation is human-labeled and coarse. Scaling the metadata-conditioning thesis further would require either annotator models or self-labeling.
- Data-leak risk with distillation of ฯ*0.6 rollouts. ฯ0.7's "match RL specialists" claim could partly reflect direct distillation of those specialists' trajectories, not emergent generalization. The paper acknowledges this as "distillation of experience" but doesn't cleanly decompose the effect.
- Reproducibility. No open weights, no training code release. Like prior ฯ papers, this is a technical report more than an independently reproducible artifact.
- Compositional generalization claim is demonstrational, not statistical. Specific anecdotes (air fryer) are compelling, but the paper does not report success rates on a large held-out compositional-task suite.
- Prompt is the new integration surface. Instead of new architectures, package new capability as new prompt modalities with independent dropout. ฯ0.7's architectural delta over ฯ0.6 is small; the recipe delta is huge.
- Suboptimal data becomes signal. With metadata conditioning, failures and RL rollouts become useful training examples โ inverting the decade-old data-curation default.
- World models move inside the prompt. Subgoal-image generation from BAGEL is a cheaper form of "world model as teacher" than DreamGen-style in-imagination RL, but shares the thesis that web-pretrained generative priors have direct use in robot policies.
- First strong compositional generalization signal in the ฯ series. Up until ฯ0.7, the series gave in-distribution dexterity and some language following; ฯ0.7 is the first release with qualitatively new zero-shot capability (operating unfamiliar kitchen appliances).
- The RL-vs-BC debate softens. ฯ*0.6 RL specialists are now being distilled back into a BC generalist via the autonomous-data pathway; RL becomes an intermediate teacher rather than a deployment mechanism.
- How does ฯ0.7 compare directly to the ICLR 2026 unified discrete-diffusion line?
- Does prompt-expansion scale further โ can you add, say, force-torque traces, acoustic cues, per-step rewards as new prompt modalities without destabilizing training?
- Can the metadata annotator be self-trained (model labels its own episodes) to remove human-annotation as a bottleneck?
- What's the minimum BAGEL-scale world model that retains the subgoal-image benefit? Is 14B overkill?
- Is there a principled way to decompose "distillation of RL specialists" from "true compositional generalization" in the evaluation metrics?
- Paper PDF: https://www.pi.website/download/pi07.pdf
- Blog / project page: https://www.pi.website/blog/pi07
- ฯ series evolution (this wiki): pi-series-evolution
- Summary page (this wiki): PI-pi07
- Predecessor: ฯ0.6 (Nov 2025), ฯ*0.6 + RECAP (Nov 2025)
- Ancestor at CoRL 2025: ฯ0.5
- TechCrunch coverage (Apr 16, 2026): https://techcrunch.com/2026/04/16/physical-intelligence-a-hot-robotics-startup-says-its-new-robot-brain-can-figure-out-tasks-it-was-never-taught/
- OpenPi (ฯ0/ฯ0.5 reference implementation): https://github.com/Physical-Intelligence/openpi