NeurIPS 2025 What Can RL Bring - Heungwoo/research GitHub Wiki

What Can RL Bring to VLA Generalization? โ€” An Empirical Study

Venue: NeurIPS 2025 ยท arXiv: 2505.19789 ยท Project: https://rlvla.github.io/ Category: RL for VLA

Approach diagram

flowchart TB
  Base[Pretrained VLA: OpenVLA] --> Fork{RL algorithm}
  Fork -- PPO --> PPOres[โœ… Best overall]
  Fork -- DPO --> DPOres[โŒ Limited by sparse reward]
  Fork -- GRPO --> GRPOres[โŒ Destabilized by non-stationary dynamics]
  PPOres --> Gen[Big gains on semantic + execution;<br/>vision robustness โ‰ˆ SFT]
Loading

Problem

By NeurIPS 2025, the RL-for-VLA toolbox had exploded โ€” DPO, GRPO, RECAP, residual-RL, R1-style โ€” but nobody had run a controlled head-to-head asking: if you just want your VLA to generalize better, which RL algorithm actually works?

Method

Fixed VLA backbone (OpenVLA), fixed evaluation suite (a ManiSkill-based setup with randomized training factors โ€” 16 tables, 16 objects, pose perturbations), swap the RL algorithm. Evaluates three generalization axes:

  • Vision โ€” foreground/background changes, image noise.
  • Semantics โ€” unseen objects, receptacles, instruction variations, new tasks.
  • Execution โ€” object/receptacle position changes, robot-pose variations.

Tested algorithms: PPO, DPO, GRPO, each with careful hyperparameter search; SFT baseline.

Results

  • PPO drives the largest gains on Execution and Semantics; Vision robustness is comparable to SFT (RL does not noticeably help here).
  • PPO consistently outperforms GRPO and DPO: GRPO is destabilized by the non-stationary dynamics of online RL, and DPO is limited by sparse rewards.
  • RL matches a 16k-demonstration SFT model in-distribution but generalizes far better out-of-distribution (e.g., ~42.6% higher on unseen objects/tables per the project page).
  • Efficient PPO-VLA recipe: shared actor-critic backbone (~45% less VRAM, ~53% faster), VLA warm-up (converges with ~50% fewer environment steps), and minimal PPO epochs (lower wall-clock time).

Significance

The first empirical answer to "what RL algorithm should I use for VLA?" Published into a landscape where GRPO had been riding the R1-wave in LLMs but hadn't been systematically tested on VLA generalization.

Informs every ICLR 2026 RL-for-VLA paper:

Links

Related pages

โ† Back to NeurIPS-2025

โš ๏ธ **GitHub.com Fallback** โš ๏ธ