PI RECAP - Heungwoo/research GitHub Wiki

Ο€*0.6 + RECAP β€” A VLA That Learns From Experience

Venue: RSS 2026 (Sydney, oral β€” VLA Models session; the conference's flagship VLA paper) Β· Physical Intelligence Β· arXiv: 2511.14759 (Nov 2025) Category: Baseline (RL for VLA) Trend tag: Trend 3 Β· RSS 2026 survey thread 1 (the VLA improvement loop)

Approach diagram

flowchart TB
  subgraph Train[Training]
    R[Rollouts] --> O{Outcome label}
    O -- success --> POS[Action labeled as POSITIVE]
    O -- failure --> NEG[Action labeled as NEGATIVE]
    POS --> M[Model: generate action<br/>conditioned on label]
    NEG --> M
  end
  subgraph Deploy[Deployment]
    Cond[Always condition on POSITIVE] --> Pol[Ο€*0.6 policy]
    Pol --> Good[Improved behavior<br/>steered away from past failures]
  end
Loading

Key figure

RECAP overview (Figure 1 of arXiv 2511.14759, Β© Physical Intelligence)

Figure 1 of the Ο€*0.6 paper. Left: the heterogeneous training diet β€” diverse robotics data, sub-task commands, and multimodal web data. Center: the Ο€*0.6 VLA (high-level + low-level with action expert) conditioned on language and an advantage token, alongside a separate VLM value function with a value head. Right: the RECAP loop β€” rollouts on the three deployment task families (assembling boxes, making espresso drinks, folding diverse laundry) flow back through interventions-and-labeling into RL training. The advantage token is the entire RL interface: no log-probs, no policy-gradient machinery.

Problem

Standard policy-gradient RL needs action log-probabilities. Flow-matching action experts (like the one in Ο€0.6) learn a deterministic vector field β€” they don't expose log-probs. So you cannot apply the standard RL toolkit to improve a shipped flow-matching VLA from its own experience.

Method

RECAP (RL with Experience and Corrections via Advantage-conditioned Policies) sidesteps the log-prob problem:

  1. Train a distributional value function (cross-entropy on discretized returns) and compute n-step advantages A^Ο€(o_t,a_t) = E[Ξ£r + V^Ο€(o_{t+N})] βˆ’ V^Ο€(o_t) (Monte-Carlo episode returns N=T in pre-training; N=50 lookahead in post-training).
  2. Binarize the advantage against a task-dependent threshold Ξ΅_β„“ to get an indicator I_t = πŸ™(A > Ξ΅_β„“), then feed it to the policy as a text token (Advantage: positive / Advantage: negative). The flow-matching action expert learns to generate actions conditioned on this label.
  3. At deployment, sample with I_t = True (optionally with classifier-free guidance, β∈[1.5, 2.5], to sharpen toward high-advantage behavior).

This is a supervised-style reformulation of advantage-weighted policy improvement that doesn't need explicit log-probs. RECAP ingests heterogeneous data β€” demonstrations, on-policy rollouts, and expert teleoperated interventions collected during autonomous execution (corrective actions are forced to I_t = True), which is the "Corrections" in the acronym.

Results

Ο€*0.6 demonstrates experience-driven improvement on real-robot tasks (folding diverse laundry, assembling boxes, making espresso with professional equipment), recovering from corner cases Ο€0.6 used to fail on. On the hardest tasks RECAP more than doubles task throughput and roughly halves the failure rate; simple enough to drop into any flow-matching policy. The RSS 2026 camera-ready abstract sharpens the deployment claims: hours-long laundry folding in real homes, reliable factory box assembly, and espresso drinks on a professional machine β€” advantage conditioning applied in both pre-training and post-training so the policy ingests "highly heterogeneous real-world experience" (demonstrations, rollouts, online corrections).

Significance

Unblocks RL for flow-matching VLAs. Combined with PLD (residual RL) and VLA-RFT (world-model RFT), RECAP closes the "RL for VLA is impractical" objection. It also defines an axis on which discrete-diffusion VLAs could have a structural edge β€” they expose log-probs natively β€” but RECAP removes the need to switch architectures purely for RL access.

Links

Related pages

  • Ο€0.6 (the base policy)
  • RSS 2026 survey β€” RECAP anchors the conference's RL-from-experience thread (alongside RLux-VLA, continual RFT, and self-improving world-model policies)
  • PLD (residual RL alternative)
  • VLA-RFT (world-model RFT alternative)
  • Qwen-VLA (ODEβ†’SDE analytic log-prob β€” the third design point for flow-matching RL)

← Back to ICLR-2026 Β· RSS-2026-VLA-Manipulation-Survey Β· Topic: RL

⚠️ **GitHub.com Fallback** ⚠️