ICML 2026 DyGRO VLA - Heungwoo/research GitHub Wiki
DyGRO-VLA: Cross-Task Scaling of Vision–Language–Action Models via Dynamic Grouped Residual Optimization — RL fine-tuning that scales VLAs across tasks instead of overfitting to one
Venue: ICML 2026 (Poster) Category: RL for VLA Affiliations: Sixu Lin, Yunpeng Qing, Litao Liu, Ming Zhou, Ruixing Jin, Xiaoyi Fan, Guiliang Liu Traction (2026-06): 1 citation (arXiv)

Problem
Reinforcement learning (RL) gives a principled way to optimize Vision-Language-Action (VLA) models, shifting them from pure trajectory imitation to active learning in the task environment. But the paper observes that most RL optimizers are task-specific: while they improve control precision on the trained task, they reduce a VLA from a generalist controller to a policy that overfits to a narrow set of tasks. The authors document two concrete failure modes. First, catastrophic forgetting: training a VLA on LIBERO-Spatial with reinforcement fine-tuning (RFT) causes success on unrelated LIBERO-Object tasks to drop sharply as training proceeds. Second, multi-task conflict: as the number of tasks grows, RFT becomes unstable and increasingly ineffective. A t-SNE study of the shared feature space (40 tasks, 10 samples each) shows that single-task RFT drives the tuned task's representations into an isolated cluster, drifting away from the shared feature space and degrading cross-task competency.
Method
DyGRO-VLA is a two-stage optimization framework:
-
Offline pre-training for cross-task representation. The authors argue that a prerequisite for VLA scalability is a shared feature space that encodes information useful across diverse tasks. This stage captures cross-task latent representations using information-theoretic principles, so the RL optimizer can later exploit task-relevant latent information without distorting the shared representation.
-
Online fine-tuning via a mixture of RL residuals. Policy optimization is refined dynamically through a mixture-of-experts (MoE) residual policy. Actions are formed as
a = Δa + a_base, whereΔais a learned residual on top of the base policy and aK-ensemble of criticsQ_{θk}(s_t, a_t)supplies value estimates. To prevent router instability and expert collapse, an entropy-based load-balancing regularizerL_LB(ψ)encourages balanced expert utilization over the per-task gating distribution.
Together this lets the optimizer exploit task-relevant latent structure while strategically mitigating adverse interference on the learned representations throughout training.

Results
On LIBERO under a 4-suite co-training protocol, DyGRO-VLA achieves leading or comparable success across all suites. Reinforcement fine-tuning lifts the average from the SFT base of 92.7% → 97.1%, with the largest gain on the hardest suite, LIBERO-Long (85.2% → 95.0%), indicating much stronger long-horizon robustness. Baselines trained under the same mixed-suite protocol include Octo, OpenVLA, SpatialVLA, π0-FAST*, π0, plus lightweight Diffusion Policy (56.5% avg from scratch) and MT-ACT. The method is further evaluated on RoboTwin2 and validated with real-world experiments, with consistent improvements over strong baselines under multi-task training and distribution shift.
Significance
DyGRO-VLA reframes RL fine-tuning of VLAs as a cross-task scaling problem rather than a per-task precision problem. By protecting a shared latent space and routing learning through balanced RL residuals, it directly targets the catastrophic-forgetting and multi-task-conflict pathologies that limit current RL optimizers for generalist robot policies.
Links
- In-depth review: DyGRO-VLA — in-depth · multi-task context: Multi-Task VLA
- arXiv: 2605.17486
← Back to ICML-2026