ICLR 2026 SimpleVLA RL - Heungwoo/research GitHub Wiki

SimpleVLA-RL โ€” Scaling VLA Training via Reinforcement Learning

Venue: ICLR 2026 ยท arXiv: 2509.09674 ยท OpenReview: TQhSodCM4r Authors: Haozhan Li, Yuxin Zuo, โ€ฆ Bowen Zhou, Ning Ding et al. (PRIME-RL) โ€” Tsinghua University, Shanghai AI Lab, Shanghai Jiao Tong University, Peking University, The University of Hong Kong Code: github.com/PRIME-RL/SimpleVLA-RL Category: RL for VLA (scaling) Trend tag: Trend 3

Approach diagram

flowchart LR
  veRL[veRL infrastructure<br/>scalable RL framework] --> Adapt[VLA-specific adaptations]
  Adapt --> S1[VLA-specific trajectory sampling]
  Adapt --> S2[Scalable parallelization<br/>many envs in parallel]
  Adapt --> S3[Multi-environment rendering]
  Adapt --> S4[Optimized loss computation<br/>VLA-aware]
  S1 --> Train[Train RL on top of OpenVLA-OFT]
  S2 --> Train
  S3 --> Train
  S4 --> Train
  Train --> R[Results:<br/>SOTA LIBERO<br/>beats ฯ€0 on RoboTwin 1 & 2<br/>discovers 'pushcut' patterns]
Loading

See the SimpleVLA-RL GitHub README and the arXiv preprint (links below) for the authors' own diagrams of the framework and the pushcut analysis figures.

Problem

Vision-Language-Action models trained by supervised fine-tuning (SFT) face two fundamental limits:

  1. Data scarcity โ€” large-scale human-operated robot trajectories are expensive and slow to collect, so SFT scaling is rate-limited by teleop hours.
  2. Generalization โ€” SFT-only VLAs overfit the demonstration distribution and degrade on tasks involving distribution shift.

RL is the natural answer, but as of mid-2025 there was no general, efficient, scalable RL framework purpose-built for VLA models. Existing RL stacks (designed for game-playing or LLM RLHF) don't handle VLA's heavy multi-modal observations, action chunking, or the parallel-rollout requirements of robot envs.

Method

SimpleVLA-RL is an efficient RL framework tailored for VLA models. It builds on veRL (the open-source scalable RL infrastructure used for LLM RL) and adds VLA-specific machinery:

  • VLA-specific trajectory sampling โ€” handles action chunks and multi-modal observations natively
  • Scalable parallelization โ€” many environments rolled out in parallel for sample throughput
  • Multi-environment rendering โ€” efficient batched rendering for sim envs
  • Optimized loss computation โ€” VLA-aware advantage estimation and gradient updates

The online RL optimizer is GRPO (Group Relative Policy Optimization, Shao et al. 2024 โ€” normalizes advantages within trajectory groups with PPO-style clipping), with VLA-specific modifications and a sparse 0/1 task-success reward. The framework is applied on top of OpenVLA-OFT (it also supports vanilla OpenVLA) and trained with carefully tuned exploration-enhancing strategies that are necessary to avoid mode-collapse on robot tasks.

Results

  • State-of-the-art on LIBERO when applied to OpenVLA-OFT: ~91.6% โ†’ 99.1% average, with the largest gain on LIBERO-Long (86.5% โ†’ 98.5%, +12 pts).
  • Extreme data efficiency. With only a single demonstration per task (one-trajectory SFT cold start), LIBERO-Long jumps from 17.3% โ†’ 91.7% after RL โ€” surpassing the full-data SFT baseline, the headline evidence that RL substitutes for teleop data.
  • Outperforms ฯ€0 on RoboTwin 1.0 and 2.0 with the exploration-enhancing strategies (e.g. RoboTwin 1.0 โ‰ˆ70.4% vs ฯ€0 โ‰ˆ58.1%; RoboTwin 2.0 โ‰ˆ68.8% vs ฯ€0 โ‰ˆ49.2%) โ€” a notable comparison, since ฯ€0's flow-matching head is the production reference point.
  • Surpasses SFT in real-world tasks โ€” not just sim โ€” on AgileX Piper dual-arm hardware (โ‰ˆ17.5% โ†’ 38.5% averaged over four tasks), reducing dependence on large-scale teleop data and improving distribution-shift generalization.
  • Discovers the "pushcut" phenomenon โ€” during RL training the policy independently discovers strategies absent from the SFT demonstrations, e.g. directly pushing an object to its goal pose instead of the demonstrated grasp-move-place routine, showing genuine exploration rather than mere refinement of the demonstrated behavior. This is strong evidence that RL adds new capabilities to VLAs rather than just polishing existing ones.

Significance

SimpleVLA-RL is the scaling-focused counterpart to the residual-RL story told by PLD and the production-engineering story told by RL Tokens. Where PLD asks "how do we get the most out of residual RL on a frozen base?" and RLT asks "how do we ship online RL on a real robot in hours?", SimpleVLA-RL asks "what is the right open-source infrastructure for VLA RL at scale?"

Three things make the paper consequential for the field:

  1. Open infrastructure. veRL โ†’ SimpleVLA-RL is among the first credible open-source paths to do industrial-scale VLA RL outside Physical Intelligence and a few similar labs (the paper itself frames it as "one of the earliest systematic explorations of VLA online RL"). This will democratize RL fine-tuning of VLAs through 2026โ€“2027.
  2. The pushcut finding. "Pushcut" โ€” the policy discovering manipulation patterns beyond the SFT distribution โ€” is the empirical evidence the field needed that VLA RL is not just polishing but genuine capability discovery. This re-frames the question from "does RL help?" (yes, settled) to "what does RL discover?" (now an active research question).
  3. Beats ฯ€0 on RoboTwin. The strongest open-data result against the production reference point. Combined with PLD's 99% LIBERO and 100% real-world numbers, the "SFT is enough for VLAs" position no longer holds.

Open questions

  • How do the gains scale with backbone size, demonstration count, environment diversity?
  • Does pushcut show up on other tasks / other VLA backbones, or is it OpenVLA-OFT-specific?
  • Can SimpleVLA-RL's framework host the residual-RL pattern (PLD / RLT) as well as full-policy RL?

Links

Related pages

โ† Back to ICLR-2026 ยท Topic: RL

โš ๏ธ **GitHub.com Fallback** โš ๏ธ