IROS 2026 VLA RL - Heungwoo/research GitHub Wiki

IROS 2026 — VLA-RL: Masterful & General Manipulation with Scalable Reinforcement Learning

Venue: IROS 2026 (Pittsburgh) · paper #4707 · Tsinghua Univ. (Shenzhen IGS) · Nanyang Technological University (Lu, Guo, Zhang, Zhou, Jiang, Gao, Tang, Wang). Paper: arXiv 2505.18719 (May 2025) · HF. The scalable-RL datapoint of IROS 2026 — online RL that improves a pretrained autoregressive VLA by casting manipulation as a multi-turn conversation and supplying dense reward via a VLM process-reward model. Companions: RL for VLA · Multi-Task VLA · DyGRO-VLA · IROS 2026 survey.

VLA-RL framework — Rollout phase: N vectorized envs (curriculum-selected) run parallel rollouts of an OpenVLA+LoRA policy; each trajectory's sparse reward R is augmented with a dense Robotic Process Reward Model signal R^rprm from a fine-tuned VLM. Learning phase: GAE + a value/policy network update the policy via PPO over a replay buffer; actions are produced by OpenVLA (LoRA) → action detokenizer → (Δx, Δθ, grip) (framework figure from Lu et al., arXiv 2505.18719, © the authors)

1. Problem

High-capacity VLAs imitate human demonstrations well, but limited state coverage in offline data causes failures out of distribution. An exploration-based method that improves from online data at test time can close this gap — but making online RL compatible with autoregressive VLAs (sparse rewards, huge action-token spaces, unstable/slow training) is the obstacle.

2. Method

VLA-RL is a systematic framework to online-RL-finetune pretrained autoregressive VLAs:

  • Trajectory-as-conversation — casts a manipulation trajectory as a multi-modal, multi-turn conversation, making RL optimization compatible with autoregressive VLAs.
  • Robotic Process Reward Model (RPRM) — a pretrained VLM fine-tuned as a process reward model, trained with pseudo reward labels from automatically extracted task segments, to densify sparse task rewards.
  • Systems for scale/stability — curriculum selection, GPU-balanced vectorized environments, batch decoding, and critic warmup; optimized via PPO with GAE.

3. Results

  • OpenVLA-7B surpasses the strongest finetuned baseline by +4.5% on 40 challenging LIBERO tasks, and matches commercial π0-FAST.
  • Under a unified real-world protocol, VLA-RL lifts OpenVLA success 60% → 90% within 3k interaction steps.
  • Keeps improving with more test-time optimization — an early sign of inference-scaling laws in robotics (more RL compute → more skill).

4. Why it matters (RL / multi-task lens)

VLA-RL is IROS 2026's strongest evidence for the survey §3.2 "policy learning at scale" thread and a key entry in the Multi-Task VLA cluster C (optimization). It differs from DyGRO-VLA (which fights cross-task forgetting) by focusing on making online RL work at all for autoregressive VLAs at scale — the conversation reformulation + VLM process-reward are the enabling tricks. The inference-scaling observation (more RL budget keeps helping) is the notable forward-looking claim, echoing the LLM RL-scaling story. Contrast AtomVLA, which gets RL-quality post-training offline (WM critic, no rollouts) — VLA-RL pays for online rollouts but gets true exploration.

Limitations (reviewer): online RL needs many real/sim rollouts (3k steps real is cheap for RL but still a fleet cost); the RPRM's pseudo-reward quality bounds the signal; OpenVLA-7B + LIBERO is the main substrate; autoregressive-VLA-specific (not flow/diffusion heads).

5. Links

← Back to IROS 2026 survey · Home