RL - Heungwoo/research GitHub Wiki
Reinforcement Learning for VLA / Manipulation โ Topic Landing
Cross-venue topical index of RL papers covered in this wiki. Covers ICLR 2026 + CoRL 2025 + NeurIPS 2025 + closely-related contemporary work (e.g., Physical Intelligence's RL Tokens).
Why a separate RL section
2026 is the year RL stops being a side note for VLA work and becomes the main lever for closing the gap left by SFT. The literature now clusters into several distinct recipes that attack different blockers โ each is large enough to be its own subsection. The decision tree below sketches the main branches (frozen-backbone vs. full-policy); the numbered sections that follow expand these into the full set of recipes plus empirical studies and infrastructure.
flowchart TB
Goal[Improve a pretrained VLA via RL] --> Q1{Backbone update?}
Q1 -- freeze backbone --> Frozen[Frozen-backbone family]
Q1 -- update full policy --> Full[Full-policy RL family]
Frozen --> R1[Residual policy on top]
Frozen --> R2[Compact RL-token interface]
Frozen --> R3[Outcome conditioning, no log-probs]
Full --> R4[World-model RFT]
Full --> R5[Stage-aware reward shaping]
Full --> R6[VLM-as-reward zero-shot]
Full --> R7[Scaling infrastructure]
1. Residual RL on a frozen backbone
The dominant pattern: keep the big VLA frozen, train a tiny adapter with RL, distill back. Backbone never destabilizes; sample efficiency is good because the adapter has few parameters.
| Paper | Adapter | Headline result |
|---|---|---|
| [PLD โ Probe, Learn, Distill](/Heungwoo/research/wiki/ICLR-2026-PLD) (ICLR 2026) | Small residual policy + distillation back into base | ~99% LIBERO, 100% on real Franka & YAM dexterous |
| [RFS โ Residual Flow Steering](/Heungwoo/research/wiki/ICLR-2026-RFS) (ICLR 2026) | Residual flow-field correction on top of flow-matching base | Stabilizes contact-rich dexterous tasks where pure imitation/RL fail |
2. Compact "RL-token" interface (production-flavored residual RL)
Same idea as residual RL but engineered for real-robot deployment timelines (hours, not days), with a single learned token serving as the bridge between the frozen VLA and a tiny RL policy.
| Paper | Mechanism | Headline result |
|---|---|---|
| ๐ [RL Tokens (RLT, Physical Intelligence)](/Heungwoo/research/wiki/ICLR-2026-RL-Tokens) (PI tech report, March 2026) | Add an "RL token" output to ฯ0.6; tiny actor + critic consume it for online RL | Up to 3ร speedup on screwdriving / zip-tying / ethernet / power-cord; surpasses human teleop speed; converges in minutes |
3. Outcome-conditioned policies (no log-probs needed)
Sidesteps the standard policy-gradient requirement of action log-probabilities โ relevant because flow-matching action experts (ฯ0 family) don't expose them.
| Paper | Trick | Headline result |
|---|---|---|
| [ฯ*0.6 + RECAP](/Heungwoo/research/wiki/PI-RECAP) (PI tech report, Nov 2025) | Label actions positive/negative by outcome; train both; condition on positive at deployment | Real-robot improvement from experience without explicit log-probs |
3a. Making flow-matching policies RL-trainable (log-prob-exact)
An alternative to the outcome-conditioned workaround: modify the flow-matching formulation itself so that it exposes exact action log-probabilities, allowing a standard PPO-style policy gradient.
| Paper | Mechanism | Headline result |
|---|---|---|
| ๐ [ReinFlow](/Heungwoo/research/wiki/NeurIPS-2025-ReinFlow) (NeurIPS 2025) | Inject learnable noise into deterministic flow path โ discrete-time Markov process with exact likelihoods | +135% reward on legged, +40% manipulation; drops into ฯ0 / ฯ0.5 / GR00T-N1.5 |
| [DSRL](/Heungwoo/research/wiki/CoRL-2025-DSRL) (CoRL 2025 Oral) | RL in the initial-noise latent space of a frozen diffusion policy | Ancestor of ReinFlow and the RL-on-pretrained-flow line |
| ๐ [Flow Policy Optimization (FPO)](/Heungwoo/research/wiki/ICRA-2026-RFT-Flow) (ICRA 2026) | Likelihood-FREE โ uses the per-sample change in the CFM objective as a PPO importance-ratio proxy (no log-prob estimate), leaves the flow sampler untouched | ฯ0 on LIBERO avg 87.2; the complementary design point to ReinFlow's exact-likelihood route and [Qwen-VLA](/Heungwoo/research/wiki/Review-Qwen-VLA)'s ODEโSDE analytic log-prob |
4. World-model RFT
Roll out RL inside a learned, controllable world model with verified rewards. Sidesteps real-rollout cost and the sim-to-real gap.
| Paper | Substrate | Headline result |
|---|---|---|
| [VLA-RFT](/Heungwoo/research/wiki/ICLR-2026-VLA-RFT) (ICLR 2026) | Data-driven controllable world model + verified rewards | Improves policy quality at fraction of real-rollout cost |
| [Ctrl-World](/Heungwoo/research/wiki/ICLR-2026-Ctrl-World) (ICLR 2026) | Generic controllable generative world model usable as RL env | Substrate for VLA-RFT and WorldGym |
5. Stage-aware reward shaping
Decompose manipulation into canonical stages so RL gets dense per-stage credit.
| Paper | Stages | Headline result |
|---|---|---|
| [Stage-Aware RL](/Heungwoo/research/wiki/ICLR-2026-Stage-Aware-RL) (ICLR 2026) | Reach โ Grasp โ Transport โ Place dense rewards | Stable RL where sparse-reward fails outright |
6. Zero-shot reward modeling (VLM-as-reward)
Use a pretrained VLM at test time as the reward / value function โ removes the need to design rewards by hand or train a learned reward model on labeled data.
| Paper | Mechanism | Headline result |
|---|---|---|
| [VITA](/Heungwoo/research/wiki/ICLR-2026-VITA) (ICLR 2026) | Test-time-adapted VLM as zero-shot value function | Trains policies on tasks with no hand-designed reward |
7. RL applied to reasoning tokens, not raw actions
Apply R1-style RL not to action prediction but to the embodied chain-of-thought / pointing primitives that the VLA emits before actions.
| Paper | Intermediate | Headline result |
|---|---|---|
| [Embodied-R1](/Heungwoo/research/wiki/ICLR-2026-Embodied-R1) (ICLR 2026) | Pointing primitives REG/RRG/OFG/VTG, two-stage RFT curriculum | Strong embodied-reasoning transfer across tasks and robots |
| ๐ [ThinkAct](/Heungwoo/research/wiki/NeurIPS-2025-ThinkAct) (NeurIPS 2025, NVIDIA) | MLLM reasoning plans rewarded by RL (goal completion + trajectory consistency); plan โ visual latent โ action | Few-shot adaptation, long-horizon planning, self-correction |
| ๐ Robot-R1 (NeurIPS 2025, arXiv 2506.00070) | R1-style RL that reinforces reasoning traces leading to accurate keypoint predictions | 7B model beats GPT-4o on low-level spatial reasoning |
8. Scaling-focused RL infrastructure
Open-source frameworks built specifically to make VLA RL scale: parallel rollouts, multi-environment rendering, VLA-aware loss computation. Targets the question "what's the right infrastructure for industrial-scale VLA RL outside of a few private labs?"
| Paper | Infrastructure | Headline result |
|---|---|---|
| ๐ [SimpleVLA-RL](/Heungwoo/research/wiki/ICLR-2026-SimpleVLA-RL) (ICLR 2026, arXiv 2509.09674) | veRL-based framework + VLA-specific sampling, parallelization, rendering, loss | SOTA LIBERO; beats ฯ0 on RoboTwin 1.0 & 2.0; introduces the "pushcut" phenomenon (RL discovers patterns not in SFT data) |
9. Empirical studies & human-preference RL
| Paper | Contribution | Headline result |
|---|---|---|
| ๐ [What Can RL Bring to VLA Generalization?](/Heungwoo/research/wiki/NeurIPS-2025-What-Can-RL-Bring) (NeurIPS 2025, rlvla.github.io) | Controlled head-to-head: PPO vs. DPO vs. GRPO on VLA semantic + execution-robustness generalization | PPO wins both axes โ informs every ICLR 2026 RL-for-VLA paper's algorithm choice |
| ๐ APO โ Action Preference Optimization (NeurIPS 2025, arXiv 2506.07127) | Binary-signal adaptive reweighting RLHF-for-VLAs; learns from on-deployment failures collected via HRI | Refines policies from binary intervention signals without reward engineering |
Adjacent โ test-time composition (not RL but related)
| Paper | Mechanism |
|---|---|
| [Compose Your Policies!](/Heungwoo/research/wiki/ICLR-2026-Compose-Your-Policies) (ICLR 2026) | Convex test-time composition of multiple diffusion / flow policies โ improves performance without any training |
Adjacent โ evaluation infrastructure
RL papers depend critically on having benchmarks that don't saturate. Two ICLR 2026 contributions matter here:
| Paper | What it provides |
|---|---|
| [RoboArena โ](/Heungwoo/research/wiki/ICLR-2026-RoboArena) | Real-to-sim auto-generated benchmarks for evaluating RL-tuned policies |
| [WorldGym](/Heungwoo/research/wiki/ICLR-2026-WorldGym) | World-model-as-environment for closed-loop evaluation without a physics simulator |
ICRA 2026 developments
ICRA 2026 (Vienna, June 1โ5) does not open a new RL recipe so much as confirm and harden the ones above โ the throughline of its ICRA RL/Data topic is that RL is now the standard post-training lever, not a from-scratch trainer, and the open problems have shifted from "does RL help?" to "is it safe and sample-efficient enough to run on the real robot, and can flow/diffusion policies actually be optimized from reward?"
On ยง3a (flow-matching RL), ICRA fills in the design space. FPO is the likelihood-free entry already rowed above โ using the per-sample CFM-loss differential as a PPO ratio proxy (ฯ0 LIBERO avg 87.2), it sits opposite ReinFlow's learnable-noise exact log-probs and Qwen-VLA's ODEโSDE analytic density as the third point on the "how do you policy-gradient a flow policy?" triangle. Complementing it from the pure-IL side, Dense-Jump Flow Matching (arXiv 2509.13574) diagnoses why adding integration steps degrades flow policies (late-time oversampling + non-Lipschitz velocity near t=1) and fixes it with U-shaped time scheduling for up to +23.7% โ a sampler-side correction that any flow-RL recipe inherits for free.
The more consequential shift is failure-aware, safety-gated real-world RL, the missing safety story for the residual/online recipes in ยง1โ2. Failure-Aware RL (arXiv 2601.07821) pairs a world-model safety critic with an offline-trained self-recovery policy and reports cutting intervention-requiring failures 73.1% while raising performance 11.3% during offline-to-online post-training, shipped with a FailureBench โ the cleanest answer this year to open question 6's deployment-safety gap. SHaRe-RL (arXiv 2509.13949) brings the same sample-efficiency/safety discipline to sub-mm contact-rich assembly by structuring skills into primitives and bounding interaction forces, and I-FailSense (arXiv 2509.16072) adds a lightweight VLM failure detector that generalizes zero-shot โ together a toolkit for the "last millimeter" problem RLT first surfaced.
Finally, ICRA reframes the sim-to-real and data axes that gate every RL pipeline as grounding rather than brute-force randomization. Phys2Real (arXiv 2510.11689) lets a VLM guess physical parameters (mass, CoM) then refines them from a few interactions, lifting a weighted-T-block push from 23%โ57% and 79%โ100% over domain randomization โ the online-adaptation flavor of sim-to-real. On the data side, real-to-sim and generative engines (ReยณSim's photorealistic reconstruction, AnchorDream's embodiment-anchored video diffusion at +36.4% sim / ~2ร real) attack the rollout-cost problem from a different direction than ยง4's world-model RFT: instead of rolling RL inside a learned world model, they manufacture cheap training data for it. On preference RL, GRAPE (trajectory-level preference alignment, +51.8% in-domain / +58.2% unseen) extends ยง9's human-preference line to flow-matching VLAs. The net picture: the recipes are settled; ICRA's contribution is making them deployable.
Cross-cutting open questions
- Composability of recipes โ can residual RL (PLD / RL Tokens) be combined with world-model RFT (VLA-RFT) and outcome conditioning (RECAP) in a single pipeline?
- Sim-to-real transfer โ do world-model-trained policies transfer to real robots as well as PLD's residual RL or RLT's online RL do?
- Scaling laws โ SimpleVLA-RL is the first real attempt to study how RL gains scale with backbone size, demo count, env diversity. We need more.
- Architecture interaction โ discrete-diffusion VLAs (Discrete Diffusion VLA) expose token log-probs natively. Do they make RECAP-style workarounds unnecessary, and do they make residual RL more or less effective?
- What does RL discover? โ SimpleVLA-RL's "pushcut" finding (the policy learns patterns absent from SFT data) re-frames the question from "does RL help?" (settled: yes) to "what does RL discover?" (open).
- The "last millimeter" problem โ RLT shows the gap from "demo" to "deployment" is driven by sub-mm precision in contact-rich phases. Are residual-RL recipes the universal answer here, or is something else needed?
Reading order for someone new to VLA RL in 2026
- Read ฯ0.6 first (the production baseline).
- Read What-Can-RL-Bring โ empirical "which RL algorithm?" answer (PPO).
- Read PLD โ the cleanest residual-RL story at benchmark level.
- Read RL Tokens โ the same idea engineered for real-robot deployment.
- Read ReinFlow โ how to make a flow-matching policy RL-trainable with exact log-probs.
- Read ThinkAct โ RL for reasoning plans, not raw actions.
- Read SimpleVLA-RL โ scaling + the pushcut finding.
- Read VLA-RFT โ the world-model alternative.
- Read RECAP โ outcome-conditioned alternative to exact log-probs.
- Then dive into the rest as needed.
Sources
- Survey writeup: Survey: VLA & Manipulation ยง5
- Physical Intelligence: https://www.pi.website/research/rlt , https://arxiv.org/html/2511.14759v1
- SimpleVLA-RL: https://arxiv.org/abs/2509.09674 , https://github.com/PRIME-RL/SimpleVLA-RL
- PLD project: https://wenlixiao.com/self-improve-VLA-PLD
๐ State of the Field (updated Aug 2026)
Verdict: RL-from-experience went from "impractical" to production in twelve months โ advantage conditioning is the deployed recipe; the open problems moved to breadth, forgetting, and safety.
๐ Trend
| Stage | Milestone |
|---|---|
| โค2025 | "RL for VLA is impractical": flow policies expose no log-probs; on-robot exploration unsafe |
| Late 2025 | Log-prob wall cracked 3 independent ways: advantage conditioning ([RECAP](/Heungwoo/research/wiki/PI-RECAP)), learnable-noise likelihoods ([ReinFlow](/Heungwoo/research/wiki/NeurIPS-2025-ReinFlow)), analytic ODEโSDE ([Qwen-VLA](/Heungwoo/research/wiki/Review-Qwen-VLA)) |
| RSS 2026 | Production proof: ฯ*0.6 runs hours-long laundry / box assembly / espresso at >2ร throughput, ~ยฝ failures; ecosystem forms โ RLux-VLA framework, continual RFT, Robometer reward models, generatorโverifier loops, BCโQ extraction, TMRL exploration |
| ICML 2026 | Reward/critic models & test-time RL mature: [VLAC](/Heungwoo/research/wiki/ICML-2026-VLAC) pairwise progress critic for dense real-world rewards; [LAGEA](/Heungwoo/research/wiki/ICML-2026-LAGEA) turns VLM failure reflections into shaped rewards; [ReLAM](/Heungwoo/research/wiki/ICML-2026-ReLAM) keypoint subgoals from action-free video; model-based loops arrive ([VLA-MBPO](/Heungwoo/research/wiki/ICML-2026-Towards-Practical-World-Model-based-Reinforcement); [VLAW](/Heungwoo/research/wiki/ICML-2026-VLAW) policyโworld-model co-improvement, +39.2% absolute); test-time critics ([VLA-ATTC](/Heungwoo/research/wiki/ICML-2026-VLA-ATTC) โ50% failures on LIBERO-LONG over ฯ0.5; [TapSampling](/Heungwoo/research/wiki/ICML-2026-TapSampling)) |
โ๏ธ Approaches & trade-offs
| Approach | Mechanism | Pros | Cons |
|---|---|---|---|
| Advantage conditioning (RECAP) | Binarized advantage as a token; condition "positive" at deployment | No log-probs; ingests demos+rollouts+corrections uniformly; stable | Coarse binary credit; needs a value function |
| Direct PG on flow (ODEโSDE / ReinFlow) | Stochasticize denoising โ likelihoods โ PPO | Principled, standard toolkit | More machinery; sim-heavy so far |
| Residual RL (PLD) | Small RL policy corrects frozen base | Safe, bolt-on | Capped by base support |
| World-model RFT (VLA-RFT) | Fine-tune in imagination | No real-robot risk | Inherits WM fidelity limits |
| Reward models (Robometer) | Comparison-trained progress rewards | Unlocks unlabeled/failed data | Reward hacking unassessed at scale |
โ ๏ธ Limitations & open problems
- Every published loop is domain-narrow (single sites or SimplerEnv-class sims); no multi-site generalization of the improvement itself.
- Continual learning is less dire than assumed โ ICML 2026's Oral Pretrained VLAs Resist Forgetting finds large pretrained VLAs barely forget, with simple Experience Replay reaching zero forgetting at small replay sizes โ but this is measured for supervised adaptation; forgetting under repeated RL remains open. Exploration safety is still procedural (human oversight), not methodological.
- No shared metric for improvement efficiency (ฮsuccess per robot-hour).
โ Back to Home