RSS 2026 RLux VLA - Heungwoo/research GitHub Wiki

RLux-VLA: A Unified and Efficient Framework for Reinforcement Learning of Vision-Language-Action Models

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: VLA Models · paper #89 Authors: Hongzhi Zang, Mingjie Wei, Si Xu, Yongji Wu, Zhen Guo, Yuanqing Wang, Hao Lin, Peihong Wang, Hua Yuan, Yixian Zhang, Liangzhi Shi, Yuqing Xie, Zhexuan Xu, Zhihao Liu, Kang Chen, Wenhao Tang, Quanlu Zhang, Weinan Zhang, Chao Yu, Yu Wang ...more arXiv: 2510.06710 · program page

Summary compiled from the arXiv paper (v2, 7 Feb 2026); all numbers quoted from the paper. The arXiv preprint is titled RLinf-VLA — the same system carries both names (the paper's own Fig. 1 labels the LIBERO panel "RLux-VLA Gain" and the RoboTwin panel "RLinf-VLA Gain"); the RSS program uses RLux-VLA. Trend context: RSS 2026 survey.

RLinf/RLux-VLA system overview (Figure 1 of arXiv 2510.06710, © the authors)

Fig. 1: A single unified interface spans multiple simulators (LIBERO, ManiSkill, RoboTwin), VLA architectures (OpenVLA, OpenVLA-OFT) and RL algorithms (PPO, GRPO). The middle band shows the three GPU-allocation strategies — collocated, disaggregated and the new hybrid mode — with throughput bars illustrating the 2.27× speedup; the bottom band shows per-task success-rate gains on LIBERO, ManiSkill and RoboTwin (solid = initial, dashed = RL gain).

Problem

Reinforcement learning is an increasingly important post-training paradigm for VLA models, but existing efforts are small-scale and fragmented: there is no unified platform for fair comparison across architectures, RL algorithms and simulators. Generic LLM-RL codebases (e.g. SimpleVLA-RL built on VeRL) also lack system-level optimizations for embodied settings, where simulators compete with model inference and training for GPU resources, producing pipeline bubbles and long training times.

Method

RLinf-VLA (RLux-VLA in the RSS program) provides a unified interface standardizing the integration of diverse VLA architectures (OpenVLA, OpenVLA-OFT), RL algorithms (PPO, GRPO) and simulators (LIBERO, ManiSkill, RoboTwin). It exposes three GPU execution modes — collocated, disaggregated, and a novel hybrid mode — using hybrid fine-grained pipelining for GPU-parallelized simulators and collocated execution for CPU-parallelized ones, plus algorithmic optimizations for training stability. The platform is released open-source.

Results

Hybrid fine-grained pipelining yields a 1.61×–1.88× training speedup, and up to 2.27× overall versus the baseline. Using a single unified model, RLinf-VLA reaches 98.11% success on 130 LIBERO tasks and 97.66% on 25 ManiSkill tasks, and an average 84.63% on six RoboTwin tasks (an average improvement of 63.75% on those tasks). Across the benchmarks, RL training gives consistent gains of roughly 20–85%.

Significance

Positioned as a foundational, open and reproducible system for RL-based VLA research, consolidating architectures, algorithms and simulators behind one interface. Connects to Review-VLA-Architecture and RL.

← Back to RSS 2026 survey · RSS-2026-Papers · Home