IROS 2026 VLA RL - Heungwoo/research GitHub Wiki
IROS 2026 — VLA-RL: Masterful & General Manipulation with Scalable Reinforcement Learning
Venue: IROS 2026 (Pittsburgh) · paper #4707 · Tsinghua Univ. (Shenzhen IGS) · Nanyang Technological University (Lu, Guo, Zhang, Zhou, Jiang, Gao, Tang, Wang). Paper: arXiv 2505.18719 (May 2025) · HF. The scalable-RL datapoint of IROS 2026 — online RL that improves a pretrained autoregressive VLA by casting manipulation as a multi-turn conversation and supplying dense reward via a VLM process-reward model. Companions: RL for VLA · Multi-Task VLA · DyGRO-VLA · IROS 2026 survey.

1. Problem
High-capacity VLAs imitate human demonstrations well, but limited state coverage in offline data causes failures out of distribution. An exploration-based method that improves from online data at test time can close this gap — but making online RL compatible with autoregressive VLAs (sparse rewards, huge action-token spaces, unstable/slow training) is the obstacle.
2. Method
VLA-RL is a systematic framework to online-RL-finetune pretrained autoregressive VLAs:
- Trajectory-as-conversation — casts a manipulation trajectory as a multi-modal, multi-turn conversation, making RL optimization compatible with autoregressive VLAs.
- Robotic Process Reward Model (RPRM) — a pretrained VLM fine-tuned as a process reward model, trained with pseudo reward labels from automatically extracted task segments, to densify sparse task rewards.
- Systems for scale/stability — curriculum selection, GPU-balanced vectorized environments, batch decoding, and critic warmup; optimized via PPO with GAE.
3. Results
- OpenVLA-7B surpasses the strongest finetuned baseline by +4.5% on 40 challenging LIBERO tasks, and matches commercial π0-FAST.
- Under a unified real-world protocol, VLA-RL lifts OpenVLA success 60% → 90% within 3k interaction steps.
- Keeps improving with more test-time optimization — an early sign of inference-scaling laws in robotics (more RL compute → more skill).
4. Why it matters (RL / multi-task lens)
VLA-RL is IROS 2026's strongest evidence for the survey §3.2 "policy learning at scale" thread and a key entry in the Multi-Task VLA cluster C (optimization). It differs from DyGRO-VLA (which fights cross-task forgetting) by focusing on making online RL work at all for autoregressive VLAs at scale — the conversation reformulation + VLM process-reward are the enabling tricks. The inference-scaling observation (more RL budget keeps helping) is the notable forward-looking claim, echoing the LLM RL-scaling story. Contrast AtomVLA, which gets RL-quality post-training offline (WM critic, no rollouts) — VLA-RL pays for online rollouts but gets true exploration.
Limitations (reviewer): online RL needs many real/sim rollouts (3k steps real is cheap for RL but still a fleet cost); the RPRM's pseudo-reward quality bounds the signal; OpenVLA-7B + LIBERO is the main substrate; autoregressive-VLA-specific (not flow/diffusion heads).
5. Links
- Paper: arXiv 2505.18719 · HF · official program: IROS 2026 (#4707)
- Related: RL for VLA · Multi-Task VLA · DyGRO-VLA · AtomVLA · IROS 2026 survey
← Back to IROS 2026 survey · Home