RSS 2026 RLux VLA - Heungwoo/research GitHub Wiki
RLux-VLA: A Unified and Efficient Framework for Reinforcement Learning of Vision-Language-Action Models
Venue: RSS 2026 (Sydney, Jul 13–17) · Session: VLA Models · paper #89 Authors: Hongzhi Zang, Mingjie Wei, Si Xu, Yongji Wu, Zhen Guo, Yuanqing Wang, Hao Lin, Peihong Wang, Hua Yuan, Yixian Zhang, Liangzhi Shi, Yuqing Xie, Zhexuan Xu, Zhihao Liu, Kang Chen, Wenhao Tang, Quanlu Zhang, Weinan Zhang, Chao Yu, Yu Wang ...more arXiv: 2510.06710 · program page
Summary compiled from the arXiv paper (v2, 7 Feb 2026); all numbers quoted from the paper. The arXiv preprint is titled RLinf-VLA — the same system carries both names (the paper's own Fig. 1 labels the LIBERO panel "RLux-VLA Gain" and the RoboTwin panel "RLinf-VLA Gain"); the RSS program uses RLux-VLA. Trend context: RSS 2026 survey.

Fig. 1: A single unified interface spans multiple simulators (LIBERO, ManiSkill, RoboTwin), VLA architectures (OpenVLA, OpenVLA-OFT) and RL algorithms (PPO, GRPO). The middle band shows the three GPU-allocation strategies — collocated, disaggregated and the new hybrid mode — with throughput bars illustrating the 2.27× speedup; the bottom band shows per-task success-rate gains on LIBERO, ManiSkill and RoboTwin (solid = initial, dashed = RL gain).
Problem
Reinforcement learning is an increasingly important post-training paradigm for VLA models, but existing efforts are small-scale and fragmented: there is no unified platform for fair comparison across architectures, RL algorithms and simulators. Generic LLM-RL codebases (e.g. SimpleVLA-RL built on VeRL) also lack system-level optimizations for embodied settings, where simulators compete with model inference and training for GPU resources, producing pipeline bubbles and long training times.
Method
RLinf-VLA (RLux-VLA in the RSS program) provides a unified interface standardizing the integration of diverse VLA architectures (OpenVLA, OpenVLA-OFT), RL algorithms (PPO, GRPO) and simulators (LIBERO, ManiSkill, RoboTwin). It exposes three GPU execution modes — collocated, disaggregated, and a novel hybrid mode — using hybrid fine-grained pipelining for GPU-parallelized simulators and collocated execution for CPU-parallelized ones, plus algorithmic optimizations for training stability. The platform is released open-source.
Results
Hybrid fine-grained pipelining yields a 1.61×–1.88× training speedup, and up to 2.27× overall versus the baseline. Using a single unified model, RLinf-VLA reaches 98.11% success on 130 LIBERO tasks and 97.66% on 25 ManiSkill tasks, and an average 84.63% on six RoboTwin tasks (an average improvement of 63.75% on those tasks). Across the benchmarks, RL training gives consistent gains of roughly 20–85%.
Significance
Positioned as a foundational, open and reproducible system for RL-based VLA research, consolidating architectures, algorithms and simulators behind one interface. Connects to Review-VLA-Architecture and RL.
← Back to RSS 2026 survey · RSS-2026-Papers · Home