ICML 2026 DyGRO VLA - Heungwoo/research GitHub Wiki

DyGRO-VLA: Cross-Task Scaling of Vision–Language–Action Models via Dynamic Grouped Residual Optimization — RL fine-tuning that scales VLAs across tasks instead of overfitting to one

Venue: ICML 2026 (Poster) Category: RL for VLA Affiliations: Sixu Lin, Yunpeng Qing, Litao Liu, Ming Zhou, Ruixing Jin, Xiaoyi Fan, Guiliang Liu Traction (2026-06): 1 citation (arXiv)

DyGRO-VLA two-stage framework: offline cross-task representation pre-training feeding online mixture-of-RL-residuals fine-tuning (Figure 1 from Lin et al., 2026)

Problem

Reinforcement learning (RL) gives a principled way to optimize Vision-Language-Action (VLA) models, shifting them from pure trajectory imitation to active learning in the task environment. But the paper observes that most RL optimizers are task-specific: while they improve control precision on the trained task, they reduce a VLA from a generalist controller to a policy that overfits to a narrow set of tasks. The authors document two concrete failure modes. First, catastrophic forgetting: training a VLA on LIBERO-Spatial with reinforcement fine-tuning (RFT) causes success on unrelated LIBERO-Object tasks to drop sharply as training proceeds. Second, multi-task conflict: as the number of tasks grows, RFT becomes unstable and increasingly ineffective. A t-SNE study of the shared feature space (40 tasks, 10 samples each) shows that single-task RFT drives the tuned task's representations into an isolated cluster, drifting away from the shared feature space and degrading cross-task competency.

Method

DyGRO-VLA is a two-stage optimization framework:

  1. Offline pre-training for cross-task representation. The authors argue that a prerequisite for VLA scalability is a shared feature space that encodes information useful across diverse tasks. This stage captures cross-task latent representations using information-theoretic principles, so the RL optimizer can later exploit task-relevant latent information without distorting the shared representation.

  2. Online fine-tuning via a mixture of RL residuals. Policy optimization is refined dynamically through a mixture-of-experts (MoE) residual policy. Actions are formed as a = Δa + a_base, where Δa is a learned residual on top of the base policy and a K-ensemble of critics Q_{θk}(s_t, a_t) supplies value estimates. To prevent router instability and expert collapse, an entropy-based load-balancing regularizer L_LB(ψ) encourages balanced expert utilization over the per-task gating distribution.

Together this lets the optimizer exploit task-relevant latent structure while strategically mitigating adverse interference on the learned representations throughout training.

Gradient conflicts under naive multi-task RL fine-tuning, motivating the grouped-residual design (Figure 2 from Lin et al., 2026)

Results

On LIBERO under a 4-suite co-training protocol, DyGRO-VLA achieves leading or comparable success across all suites. Reinforcement fine-tuning lifts the average from the SFT base of 92.7% → 97.1%, with the largest gain on the hardest suite, LIBERO-Long (85.2% → 95.0%), indicating much stronger long-horizon robustness. Baselines trained under the same mixed-suite protocol include Octo, OpenVLA, SpatialVLA, π0-FAST*, π0, plus lightweight Diffusion Policy (56.5% avg from scratch) and MT-ACT. The method is further evaluated on RoboTwin2 and validated with real-world experiments, with consistent improvements over strong baselines under multi-task training and distribution shift.

Significance

DyGRO-VLA reframes RL fine-tuning of VLAs as a cross-task scaling problem rather than a per-task precision problem. By protecting a shared latent space and routing learning through balanced RL residuals, it directly targets the catastrophic-forgetting and multi-task-conflict pathologies that limit current RL optimizers for generalist robot policies.

Links

← Back to ICML-2026