Review DyGRO VLA - Heungwoo/research GitHub Wiki
Paper: "DyGRO-VLA: Cross-Task Scaling of Vision–Language–Action Models via Dynamic Grouped Residual Optimization" — arXiv 2605.17486 · ICML 2026 (Poster) · Sixu Lin, Yunpeng Qing, … Guiliang Liu. What it is: the RL-optimization answer to multi-task VLA — stop single-task RL fine-tuning from overfitting/forgetting by protecting a shared feature space and routing learning through balanced RL residuals. Clusters C + F in Multi-Task VLA §3. Venue entry: ICML-2026-DyGRO-VLA · companions RL for VLA · sibling MoE approach HiMoE-VLA.
- RL fine-tuning quietly de-generalizes a VLA. Most RL optimizers are task-specific: they sharpen the trained task but collapse the shared representation, turning a generalist into a narrow policy.
- Two measured pathologies. Catastrophic forgetting — RFT on LIBERO-Spatial sharply drops unrelated LIBERO-Object success as training proceeds. Multi-task conflict — RFT becomes unstable as task count grows; a t-SNE study (40 tasks × 10) shows single-task RFT drives the task into an isolated cluster, drifting off the shared feature space.
- Fix = two stages. (1) Offline cross-task representation pre-training builds an information-theoretic shared latent space; (2) online mixture-of-RL-residuals with dynamic task grouping refines the policy without distorting that space.
- Result: on LIBERO 4-suite co-training, RFT lifts 92.7% → 97.1% avg, with the biggest gain on the hardest suite LIBERO-Long 85.2% → 95.0%; validated on RoboTwin2 and real robots.
- It names an under-appreciated failure: RL itself causes multi-task collapse. While merging (MergeVLA) and MoE (HiMoE-VLA) fight interference in supervised multi-task training, DyGRO shows the same pathology bites during RL fine-tuning — and that naive per-task RFT is actively harmful to a generalist.
- Protect-then-optimize. The key move is ordering: first secure a cross-task latent space, then let RL exploit task-relevant structure through residuals rather than rewriting shared weights — a general recipe for "improve one task without forgetting the rest."
- Long-horizon is where it pays. The +9.8-point jump on LIBERO-Long suggests the shared-space protection matters most exactly where dense RFT is most brittle.
flowchart LR
subgraph S1[Stage 1 — offline]
P[Cross-task representation pre-training<br/>information-theoretic shared latent space]
end
subgraph S2[Stage 2 — online RL]
B[Base policy a_base] --> ADD["a = Δa + a_base"]
R[Mixture-of-RL-residuals Δa<br/>dynamic task grouping by feature similarity] --> ADD
Q[K-ensemble critics Q_θk] --> R
LB[Entropy load-balancing L_LB<br/>anti expert-collapse] -.-> R
end
P --> S2
ADD --> OUT[action]
- Stage 1 — cross-task latent capture. Argues a shared feature space is the prerequisite for scalability; captures cross-task latents via information-theoretic principles so later RL can exploit task-relevant info without distorting the shared representation.
-
Stage 2 — mixture-of-RL-residuals. Actions are
a = Δa + a_base(learned residual on a frozen-ish base). A K-ensemble of critics supplies value estimates; tasks are dynamically grouped by learned feature similarity and refined per group; an entropy-based load-balancing regularizerL_LB(ψ)prevents router instability / expert collapse over the per-task gating distribution.
- LIBERO (4-suite co-training): SFT base 92.7% → RFT 97.1% avg; LIBERO-Long 85.2% → 95.0% (largest gain → long-horizon robustness).
- Baselines (same mixed-suite protocol): Octo, OpenVLA, SpatialVLA, π0-FAST, π0, Diffusion Policy (56.5% from scratch), MT-ACT.
- Also: RoboTwin2 + real-world experiments — consistent gains over strong baselines under multi-task training and distribution shift.
Significance. DyGRO-VLA reframes RL fine-tuning of VLAs as a cross-task scaling problem (not per-task precision), and gives a concrete two-stage recipe — protect the shared latent, then optimize through balanced residuals — that directly targets forgetting and multi-task conflict.
Limitations (reviewer's).
- No explicit limitations section; robustness of the dynamic grouping at very large task counts is not stressed.
- Two-stage cost — the offline representation stage adds a pretraining phase before RL.
- Residual on a base policy assumes a competent base; gains where the base is weak are unclear.
- RL-side only — pairs with, but doesn't replace, architectural fixes (MoE/merging) for the supervised multi-task problem.
Design lesson. If you must RL-tune across tasks, don't let the optimizer rewrite the shared representation — freeze/protect it and learn grouped residuals on top.
- Paper: arXiv 2605.17486 · venue entry ICML-2026-DyGRO-VLA
- Multi-task context: Multi-Task VLA (clusters C + F) · sibling routing approach HiMoE-VLA · merging approach MergeVLA
- RL context: RL for VLA · continual: Stellar VLA · Pretrained VLAs Resist Forgetting