PI RECAP - Heungwoo/research GitHub Wiki
Venue: RSS 2026 (Sydney, oral β VLA Models session; the conference's flagship VLA paper) Β· Physical Intelligence Β· arXiv: 2511.14759 (Nov 2025) Category: Baseline (RL for VLA) Trend tag: Trend 3 Β· RSS 2026 survey thread 1 (the VLA improvement loop)
flowchart TB
subgraph Train[Training]
R[Rollouts] --> O{Outcome label}
O -- success --> POS[Action labeled as POSITIVE]
O -- failure --> NEG[Action labeled as NEGATIVE]
POS --> M[Model: generate action<br/>conditioned on label]
NEG --> M
end
subgraph Deploy[Deployment]
Cond[Always condition on POSITIVE] --> Pol[Ο*0.6 policy]
Pol --> Good[Improved behavior<br/>steered away from past failures]
end

Figure 1 of the Ο*0.6 paper. Left: the heterogeneous training diet β diverse robotics data, sub-task commands, and multimodal web data. Center: the Ο*0.6 VLA (high-level + low-level with action expert) conditioned on language and an advantage token, alongside a separate VLM value function with a value head. Right: the RECAP loop β rollouts on the three deployment task families (assembling boxes, making espresso drinks, folding diverse laundry) flow back through interventions-and-labeling into RL training. The advantage token is the entire RL interface: no log-probs, no policy-gradient machinery.
Standard policy-gradient RL needs action log-probabilities. Flow-matching action experts (like the one in Ο0.6) learn a deterministic vector field β they don't expose log-probs. So you cannot apply the standard RL toolkit to improve a shipped flow-matching VLA from its own experience.
RECAP (RL with Experience and Corrections via Advantage-conditioned Policies) sidesteps the log-prob problem:
- Train a distributional value function (cross-entropy on discretized returns) and compute n-step advantages
A^Ο(o_t,a_t) = E[Ξ£r + V^Ο(o_{t+N})] β V^Ο(o_t)(Monte-Carlo episode returnsN=Tin pre-training;N=50lookahead in post-training). -
Binarize the advantage against a task-dependent threshold
Ξ΅_βto get an indicatorI_t = π(A > Ξ΅_β), then feed it to the policy as a text token (Advantage: positive/Advantage: negative). The flow-matching action expert learns to generate actions conditioned on this label. - At deployment, sample with
I_t = True(optionally with classifier-free guidance, Ξ²β[1.5, 2.5], to sharpen toward high-advantage behavior).
This is a supervised-style reformulation of advantage-weighted policy improvement that doesn't need explicit log-probs. RECAP ingests heterogeneous data β demonstrations, on-policy rollouts, and expert teleoperated interventions collected during autonomous execution (corrective actions are forced to I_t = True), which is the "Corrections" in the acronym.
Ο*0.6 demonstrates experience-driven improvement on real-robot tasks (folding diverse laundry, assembling boxes, making espresso with professional equipment), recovering from corner cases Ο0.6 used to fail on. On the hardest tasks RECAP more than doubles task throughput and roughly halves the failure rate; simple enough to drop into any flow-matching policy. The RSS 2026 camera-ready abstract sharpens the deployment claims: hours-long laundry folding in real homes, reliable factory box assembly, and espresso drinks on a professional machine β advantage conditioning applied in both pre-training and post-training so the policy ingests "highly heterogeneous real-world experience" (demonstrations, rollouts, online corrections).
Unblocks RL for flow-matching VLAs. Combined with PLD (residual RL) and VLA-RFT (world-model RFT), RECAP closes the "RL for VLA is impractical" objection. It also defines an axis on which discrete-diffusion VLAs could have a structural edge β they expose log-probs natively β but RECAP removes the need to switch architectures purely for RL access.
- arXiv: https://arxiv.org/html/2511.14759v1
- Project PDF: https://www.pi.website/download/pistar06.pdf
- Independent explainer (Federico Sarrocco): https://federicosarrocco.com/blog/pi-star-06-recap
- Ο0.6 (the base policy)
- RSS 2026 survey β RECAP anchors the conference's RL-from-experience thread (alongside RLux-VLA, continual RFT, and self-improving world-model policies)
- PLD (residual RL alternative)
- VLA-RFT (world-model RFT alternative)
- Qwen-VLA (ODEβSDE analytic log-prob β the third design point for flow-matching RL)
β Back to ICLR-2026 Β· RSS-2026-VLA-Manipulation-Survey Β· Topic: RL