NeurIPS 2025 What Can RL Bring - Heungwoo/research GitHub Wiki
Venue: NeurIPS 2025 ยท arXiv: 2505.19789 ยท Project: https://rlvla.github.io/ Category: RL for VLA
flowchart TB
Base[Pretrained VLA: OpenVLA] --> Fork{RL algorithm}
Fork -- PPO --> PPOres[โ
Best overall]
Fork -- DPO --> DPOres[โ Limited by sparse reward]
Fork -- GRPO --> GRPOres[โ Destabilized by non-stationary dynamics]
PPOres --> Gen[Big gains on semantic + execution;<br/>vision robustness โ SFT]
By NeurIPS 2025, the RL-for-VLA toolbox had exploded โ DPO, GRPO, RECAP, residual-RL, R1-style โ but nobody had run a controlled head-to-head asking: if you just want your VLA to generalize better, which RL algorithm actually works?
Fixed VLA backbone (OpenVLA), fixed evaluation suite (a ManiSkill-based setup with randomized training factors โ 16 tables, 16 objects, pose perturbations), swap the RL algorithm. Evaluates three generalization axes:
- Vision โ foreground/background changes, image noise.
- Semantics โ unseen objects, receptacles, instruction variations, new tasks.
- Execution โ object/receptacle position changes, robot-pose variations.
Tested algorithms: PPO, DPO, GRPO, each with careful hyperparameter search; SFT baseline.
- PPO drives the largest gains on Execution and Semantics; Vision robustness is comparable to SFT (RL does not noticeably help here).
- PPO consistently outperforms GRPO and DPO: GRPO is destabilized by the non-stationary dynamics of online RL, and DPO is limited by sparse rewards.
- RL matches a 16k-demonstration SFT model in-distribution but generalizes far better out-of-distribution (e.g., ~42.6% higher on unseen objects/tables per the project page).
- Efficient PPO-VLA recipe: shared actor-critic backbone (~45% less VRAM, ~53% faster), VLA warm-up (converges with ~50% fewer environment steps), and minimal PPO epochs (lower wall-clock time).
The first empirical answer to "what RL algorithm should I use for VLA?" Published into a landscape where GRPO had been riding the R1-wave in LLMs but hadn't been systematically tested on VLA generalization.
Informs every ICLR 2026 RL-for-VLA paper:
- SimpleVLA-RL picks PPO (and extends the recipe).
- Embodied-R1 uses GRPO-style for reasoning tokens specifically.
- VLA-RFT uses PPO in a world-model.
- RL Tokens uses PPO on a compact actor/critic.
- arXiv: https://arxiv.org/abs/2505.19789
- NeurIPS virtual page: https://neurips.cc/virtual/2025/loc/san-diego/poster/115842
- Project page: https://rlvla.github.io/
- SimpleVLA-RL ยท RL Tokens ยท VLA-RFT (PPO descendants)
- Embodied-R1 (GRPO-for-reasoning)
- RL for VLA
โ Back to NeurIPS-2025