RL - Heungwoo/research GitHub Wiki

Reinforcement Learning for VLA / Manipulation โ€” Topic Landing

Cross-venue topical index of RL papers covered in this wiki. Covers ICLR 2026 + CoRL 2025 + NeurIPS 2025 + closely-related contemporary work (e.g., Physical Intelligence's RL Tokens).

Why a separate RL section

2026 is the year RL stops being a side note for VLA work and becomes the main lever for closing the gap left by SFT. The literature now clusters into several distinct recipes that attack different blockers โ€” each is large enough to be its own subsection. The decision tree below sketches the main branches (frozen-backbone vs. full-policy); the numbered sections that follow expand these into the full set of recipes plus empirical studies and infrastructure.

flowchart TB
  Goal[Improve a pretrained VLA via RL] --> Q1{Backbone update?}
  Q1 -- freeze backbone --> Frozen[Frozen-backbone family]
  Q1 -- update full policy --> Full[Full-policy RL family]
  Frozen --> R1[Residual policy on top]
  Frozen --> R2[Compact RL-token interface]
  Frozen --> R3[Outcome conditioning, no log-probs]
  Full --> R4[World-model RFT]
  Full --> R5[Stage-aware reward shaping]
  Full --> R6[VLM-as-reward zero-shot]
  Full --> R7[Scaling infrastructure]

1. Residual RL on a frozen backbone

The dominant pattern: keep the big VLA frozen, train a tiny adapter with RL, distill back. Backbone never destabilizes; sample efficiency is good because the adapter has few parameters.

Paper Adapter Headline result
[PLD โ€” Probe, Learn, Distill](/Heungwoo/research/wiki/ICLR-2026-PLD) (ICLR 2026) Small residual policy + distillation back into base ~99% LIBERO, 100% on real Franka & YAM dexterous
[RFS โ€” Residual Flow Steering](/Heungwoo/research/wiki/ICLR-2026-RFS) (ICLR 2026) Residual flow-field correction on top of flow-matching base Stabilizes contact-rich dexterous tasks where pure imitation/RL fail

2. Compact "RL-token" interface (production-flavored residual RL)

Same idea as residual RL but engineered for real-robot deployment timelines (hours, not days), with a single learned token serving as the bridge between the frozen VLA and a tiny RL policy.

Paper Mechanism Headline result
๐Ÿ†• [RL Tokens (RLT, Physical Intelligence)](/Heungwoo/research/wiki/ICLR-2026-RL-Tokens) (PI tech report, March 2026) Add an "RL token" output to ฯ€0.6; tiny actor + critic consume it for online RL Up to 3ร— speedup on screwdriving / zip-tying / ethernet / power-cord; surpasses human teleop speed; converges in minutes

3. Outcome-conditioned policies (no log-probs needed)

Sidesteps the standard policy-gradient requirement of action log-probabilities โ€” relevant because flow-matching action experts (ฯ€0 family) don't expose them.

Paper Trick Headline result
[ฯ€*0.6 + RECAP](/Heungwoo/research/wiki/PI-RECAP) (PI tech report, Nov 2025) Label actions positive/negative by outcome; train both; condition on positive at deployment Real-robot improvement from experience without explicit log-probs

3a. Making flow-matching policies RL-trainable (log-prob-exact)

An alternative to the outcome-conditioned workaround: modify the flow-matching formulation itself so that it exposes exact action log-probabilities, allowing a standard PPO-style policy gradient.

Paper Mechanism Headline result
๐Ÿ†• [ReinFlow](/Heungwoo/research/wiki/NeurIPS-2025-ReinFlow) (NeurIPS 2025) Inject learnable noise into deterministic flow path โ†’ discrete-time Markov process with exact likelihoods +135% reward on legged, +40% manipulation; drops into ฯ€0 / ฯ€0.5 / GR00T-N1.5
[DSRL](/Heungwoo/research/wiki/CoRL-2025-DSRL) (CoRL 2025 Oral) RL in the initial-noise latent space of a frozen diffusion policy Ancestor of ReinFlow and the RL-on-pretrained-flow line
๐Ÿ†• [Flow Policy Optimization (FPO)](/Heungwoo/research/wiki/ICRA-2026-RFT-Flow) (ICRA 2026) Likelihood-FREE โ€” uses the per-sample change in the CFM objective as a PPO importance-ratio proxy (no log-prob estimate), leaves the flow sampler untouched ฯ€0 on LIBERO avg 87.2; the complementary design point to ReinFlow's exact-likelihood route and [Qwen-VLA](/Heungwoo/research/wiki/Review-Qwen-VLA)'s ODEโ†’SDE analytic log-prob

4. World-model RFT

Roll out RL inside a learned, controllable world model with verified rewards. Sidesteps real-rollout cost and the sim-to-real gap.

Paper Substrate Headline result
[VLA-RFT](/Heungwoo/research/wiki/ICLR-2026-VLA-RFT) (ICLR 2026) Data-driven controllable world model + verified rewards Improves policy quality at fraction of real-rollout cost
[Ctrl-World](/Heungwoo/research/wiki/ICLR-2026-Ctrl-World) (ICLR 2026) Generic controllable generative world model usable as RL env Substrate for VLA-RFT and WorldGym

5. Stage-aware reward shaping

Decompose manipulation into canonical stages so RL gets dense per-stage credit.

Paper Stages Headline result
[Stage-Aware RL](/Heungwoo/research/wiki/ICLR-2026-Stage-Aware-RL) (ICLR 2026) Reach โ†’ Grasp โ†’ Transport โ†’ Place dense rewards Stable RL where sparse-reward fails outright

6. Zero-shot reward modeling (VLM-as-reward)

Use a pretrained VLM at test time as the reward / value function โ€” removes the need to design rewards by hand or train a learned reward model on labeled data.

Paper Mechanism Headline result
[VITA](/Heungwoo/research/wiki/ICLR-2026-VITA) (ICLR 2026) Test-time-adapted VLM as zero-shot value function Trains policies on tasks with no hand-designed reward

7. RL applied to reasoning tokens, not raw actions

Apply R1-style RL not to action prediction but to the embodied chain-of-thought / pointing primitives that the VLA emits before actions.

Paper Intermediate Headline result
[Embodied-R1](/Heungwoo/research/wiki/ICLR-2026-Embodied-R1) (ICLR 2026) Pointing primitives REG/RRG/OFG/VTG, two-stage RFT curriculum Strong embodied-reasoning transfer across tasks and robots
๐Ÿ†• [ThinkAct](/Heungwoo/research/wiki/NeurIPS-2025-ThinkAct) (NeurIPS 2025, NVIDIA) MLLM reasoning plans rewarded by RL (goal completion + trajectory consistency); plan โ†’ visual latent โ†’ action Few-shot adaptation, long-horizon planning, self-correction
๐Ÿ†• Robot-R1 (NeurIPS 2025, arXiv 2506.00070) R1-style RL that reinforces reasoning traces leading to accurate keypoint predictions 7B model beats GPT-4o on low-level spatial reasoning

8. Scaling-focused RL infrastructure

Open-source frameworks built specifically to make VLA RL scale: parallel rollouts, multi-environment rendering, VLA-aware loss computation. Targets the question "what's the right infrastructure for industrial-scale VLA RL outside of a few private labs?"

Paper Infrastructure Headline result
๐Ÿ†• [SimpleVLA-RL](/Heungwoo/research/wiki/ICLR-2026-SimpleVLA-RL) (ICLR 2026, arXiv 2509.09674) veRL-based framework + VLA-specific sampling, parallelization, rendering, loss SOTA LIBERO; beats ฯ€0 on RoboTwin 1.0 & 2.0; introduces the "pushcut" phenomenon (RL discovers patterns not in SFT data)

9. Empirical studies & human-preference RL

Paper Contribution Headline result
๐Ÿ†• [What Can RL Bring to VLA Generalization?](/Heungwoo/research/wiki/NeurIPS-2025-What-Can-RL-Bring) (NeurIPS 2025, rlvla.github.io) Controlled head-to-head: PPO vs. DPO vs. GRPO on VLA semantic + execution-robustness generalization PPO wins both axes โ€” informs every ICLR 2026 RL-for-VLA paper's algorithm choice
๐Ÿ†• APO โ€” Action Preference Optimization (NeurIPS 2025, arXiv 2506.07127) Binary-signal adaptive reweighting RLHF-for-VLAs; learns from on-deployment failures collected via HRI Refines policies from binary intervention signals without reward engineering

Adjacent โ€” test-time composition (not RL but related)

Paper Mechanism
[Compose Your Policies!](/Heungwoo/research/wiki/ICLR-2026-Compose-Your-Policies) (ICLR 2026) Convex test-time composition of multiple diffusion / flow policies โ€” improves performance without any training

Adjacent โ€” evaluation infrastructure

RL papers depend critically on having benchmarks that don't saturate. Two ICLR 2026 contributions matter here:

Paper What it provides
[RoboArena โˆž](/Heungwoo/research/wiki/ICLR-2026-RoboArena) Real-to-sim auto-generated benchmarks for evaluating RL-tuned policies
[WorldGym](/Heungwoo/research/wiki/ICLR-2026-WorldGym) World-model-as-environment for closed-loop evaluation without a physics simulator

ICRA 2026 developments

ICRA 2026 (Vienna, June 1โ€“5) does not open a new RL recipe so much as confirm and harden the ones above โ€” the throughline of its ICRA RL/Data topic is that RL is now the standard post-training lever, not a from-scratch trainer, and the open problems have shifted from "does RL help?" to "is it safe and sample-efficient enough to run on the real robot, and can flow/diffusion policies actually be optimized from reward?"

On ยง3a (flow-matching RL), ICRA fills in the design space. FPO is the likelihood-free entry already rowed above โ€” using the per-sample CFM-loss differential as a PPO ratio proxy (ฯ€0 LIBERO avg 87.2), it sits opposite ReinFlow's learnable-noise exact log-probs and Qwen-VLA's ODEโ†’SDE analytic density as the third point on the "how do you policy-gradient a flow policy?" triangle. Complementing it from the pure-IL side, Dense-Jump Flow Matching (arXiv 2509.13574) diagnoses why adding integration steps degrades flow policies (late-time oversampling + non-Lipschitz velocity near t=1) and fixes it with U-shaped time scheduling for up to +23.7% โ€” a sampler-side correction that any flow-RL recipe inherits for free.

The more consequential shift is failure-aware, safety-gated real-world RL, the missing safety story for the residual/online recipes in ยง1โ€“2. Failure-Aware RL (arXiv 2601.07821) pairs a world-model safety critic with an offline-trained self-recovery policy and reports cutting intervention-requiring failures 73.1% while raising performance 11.3% during offline-to-online post-training, shipped with a FailureBench โ€” the cleanest answer this year to open question 6's deployment-safety gap. SHaRe-RL (arXiv 2509.13949) brings the same sample-efficiency/safety discipline to sub-mm contact-rich assembly by structuring skills into primitives and bounding interaction forces, and I-FailSense (arXiv 2509.16072) adds a lightweight VLM failure detector that generalizes zero-shot โ€” together a toolkit for the "last millimeter" problem RLT first surfaced.

Finally, ICRA reframes the sim-to-real and data axes that gate every RL pipeline as grounding rather than brute-force randomization. Phys2Real (arXiv 2510.11689) lets a VLM guess physical parameters (mass, CoM) then refines them from a few interactions, lifting a weighted-T-block push from 23%โ†’57% and 79%โ†’100% over domain randomization โ€” the online-adaptation flavor of sim-to-real. On the data side, real-to-sim and generative engines (ReยณSim's photorealistic reconstruction, AnchorDream's embodiment-anchored video diffusion at +36.4% sim / ~2ร— real) attack the rollout-cost problem from a different direction than ยง4's world-model RFT: instead of rolling RL inside a learned world model, they manufacture cheap training data for it. On preference RL, GRAPE (trajectory-level preference alignment, +51.8% in-domain / +58.2% unseen) extends ยง9's human-preference line to flow-matching VLAs. The net picture: the recipes are settled; ICRA's contribution is making them deployable.


Cross-cutting open questions

  1. Composability of recipes โ€” can residual RL (PLD / RL Tokens) be combined with world-model RFT (VLA-RFT) and outcome conditioning (RECAP) in a single pipeline?
  2. Sim-to-real transfer โ€” do world-model-trained policies transfer to real robots as well as PLD's residual RL or RLT's online RL do?
  3. Scaling laws โ€” SimpleVLA-RL is the first real attempt to study how RL gains scale with backbone size, demo count, env diversity. We need more.
  4. Architecture interaction โ€” discrete-diffusion VLAs (Discrete Diffusion VLA) expose token log-probs natively. Do they make RECAP-style workarounds unnecessary, and do they make residual RL more or less effective?
  5. What does RL discover? โ€” SimpleVLA-RL's "pushcut" finding (the policy learns patterns absent from SFT data) re-frames the question from "does RL help?" (settled: yes) to "what does RL discover?" (open).
  6. The "last millimeter" problem โ€” RLT shows the gap from "demo" to "deployment" is driven by sub-mm precision in contact-rich phases. Are residual-RL recipes the universal answer here, or is something else needed?

Reading order for someone new to VLA RL in 2026

  1. Read ฯ€0.6 first (the production baseline).
  2. Read What-Can-RL-Bring โ€” empirical "which RL algorithm?" answer (PPO).
  3. Read PLD โ€” the cleanest residual-RL story at benchmark level.
  4. Read RL Tokens โ€” the same idea engineered for real-robot deployment.
  5. Read ReinFlow โ€” how to make a flow-matching policy RL-trainable with exact log-probs.
  6. Read ThinkAct โ€” RL for reasoning plans, not raw actions.
  7. Read SimpleVLA-RL โ€” scaling + the pushcut finding.
  8. Read VLA-RFT โ€” the world-model alternative.
  9. Read RECAP โ€” outcome-conditioned alternative to exact log-probs.
  10. Then dive into the rest as needed.

Sources


๐Ÿ—“ State of the Field (updated Aug 2026)

Verdict: RL-from-experience went from "impractical" to production in twelve months โ€” advantage conditioning is the deployed recipe; the open problems moved to breadth, forgetting, and safety.

๐Ÿ“ˆ Trend

Stage Milestone
โ‰ค2025 "RL for VLA is impractical": flow policies expose no log-probs; on-robot exploration unsafe
Late 2025 Log-prob wall cracked 3 independent ways: advantage conditioning ([RECAP](/Heungwoo/research/wiki/PI-RECAP)), learnable-noise likelihoods ([ReinFlow](/Heungwoo/research/wiki/NeurIPS-2025-ReinFlow)), analytic ODEโ†’SDE ([Qwen-VLA](/Heungwoo/research/wiki/Review-Qwen-VLA))
RSS 2026 Production proof: ฯ€*0.6 runs hours-long laundry / box assembly / espresso at >2ร— throughput, ~ยฝ failures; ecosystem forms โ€” RLux-VLA framework, continual RFT, Robometer reward models, generatorโ€“verifier loops, BCโ†’Q extraction, TMRL exploration
ICML 2026 Reward/critic models & test-time RL mature: [VLAC](/Heungwoo/research/wiki/ICML-2026-VLAC) pairwise progress critic for dense real-world rewards; [LAGEA](/Heungwoo/research/wiki/ICML-2026-LAGEA) turns VLM failure reflections into shaped rewards; [ReLAM](/Heungwoo/research/wiki/ICML-2026-ReLAM) keypoint subgoals from action-free video; model-based loops arrive ([VLA-MBPO](/Heungwoo/research/wiki/ICML-2026-Towards-Practical-World-Model-based-Reinforcement); [VLAW](/Heungwoo/research/wiki/ICML-2026-VLAW) policyโ†”world-model co-improvement, +39.2% absolute); test-time critics ([VLA-ATTC](/Heungwoo/research/wiki/ICML-2026-VLA-ATTC) โˆ’50% failures on LIBERO-LONG over ฯ€0.5; [TapSampling](/Heungwoo/research/wiki/ICML-2026-TapSampling))

โš–๏ธ Approaches & trade-offs

Approach Mechanism Pros Cons
Advantage conditioning (RECAP) Binarized advantage as a token; condition "positive" at deployment No log-probs; ingests demos+rollouts+corrections uniformly; stable Coarse binary credit; needs a value function
Direct PG on flow (ODEโ†’SDE / ReinFlow) Stochasticize denoising โ†’ likelihoods โ†’ PPO Principled, standard toolkit More machinery; sim-heavy so far
Residual RL (PLD) Small RL policy corrects frozen base Safe, bolt-on Capped by base support
World-model RFT (VLA-RFT) Fine-tune in imagination No real-robot risk Inherits WM fidelity limits
Reward models (Robometer) Comparison-trained progress rewards Unlocks unlabeled/failed data Reward hacking unassessed at scale

โš ๏ธ Limitations & open problems

  • Every published loop is domain-narrow (single sites or SimplerEnv-class sims); no multi-site generalization of the improvement itself.
  • Continual learning is less dire than assumed โ€” ICML 2026's Oral Pretrained VLAs Resist Forgetting finds large pretrained VLAs barely forget, with simple Experience Replay reaching zero forgetting at small replay sizes โ€” but this is measured for supervised adaptation; forgetting under repeated RL remains open. Exploration safety is still procedural (human oversight), not methodological.
  • No shared metric for improvement efficiency (ฮ”success per robot-hour).

โ† Back to Home