A side-by-side comparison of Physical Intelligence's VLA releases through model, data, and training lenses. Each version is an additive delta on the previous one β the table makes the deltas legible.
Ο0 β Ο0.5: Hierarchical split (high-level subtask head + low-level action expert). Discrete FAST-token stage for pretraining.
Ο0.5 β Ο0.6: Backbone upgraded to Gemma3-4B, action expert grown to 860M, Knowledge Insulation stabilizes joint training.
Ο0.6 β Ο0.7: Keeps the same backbone/expert; adds MEM video-history encoder, subgoal-image conditioning, linear-projected proprioception, RTC, and CFG on metadata. Worldmodel (BAGEL-14B) is a separate module that feeds the prompt, not an architectural change to the policy itself.
Data axis
Ο0 β Ο0.5: Large-scale home data (~400 hrs MM Γ 100 homes), web multimodal co-training.
Ο0.5 β Ο0.6: More of the same, richer metadata annotation.
Ο0.6 β Ο0.7: Doubles down on heterogeneity β failures, autonomous rollouts, RL-specialist traces, egocentric human video, DROID. The novelty is not the sources but that metadata conditioning lets them be used productively instead of averaged.
Training axis
Ο0 β Ο0.5: Two-stage FASTβflow-matching; co-training on VQA/detection/subtask text.
Ο0.5 β Ο0.6: Knowledge Insulation (one-stage joint training with gradient-insulated VLM).
Ο0.6 β Ο*0.6: Advantage-conditioned RL (RECAP) β no log-prob RL needed (flow-matching policies give no action log-probs, so PPO/SAC are inapplicable). A multi-task distributional value function $V^{\pi_{ref}}(o_t,\ell)$ is trained on all data (demos + autonomous rollouts + teleoperated interventions); the policy then conditions on a binarized advantage indicator as a text token ("Advantage: positive/negative"), and at inference is run on the positive branch. Trained end-to-end with Knowledge Insulation. On the hardest tasks (laundry folding, box assembly, espresso) RECAP more than doubles task throughput and roughly halves the failure rate vs the base policy.
Ο0.6 β Ο0.7: Rich prompt conditioning + per-component dropout + CFG, distilling Ο*0.6 RL behavior into a generalist without RL at post-train time.
Key takeaway
The Ο series is evolving along a consistent thesis: push more capability into pre-training so post-training becomes optional. Ο0 needed per-task finetuning; Ο0.5 reduced it; Ο0.6 removed it for many tasks; Ο0.7 removes it even for tasks previously only solvable by RL specialists β and shows the first signs of compositional (not just in-distribution) generalization in the series (the Ο0.7 blog phrases it as "the first signs of compositional generalization," recombining skills across tasks).