Review Multitask VLA - Heungwoo/research GitHub Wiki
In-Depth Review — Multi-Task VLA: why one policy struggles across tasks, and how to fix it
Scope: why a single Vision-Language-Action model degrades when asked to do many tasks at once, and the 2025–2026 research fixing it. Anchor case study: MergeVLA (CVPR 2026). Companions: VLA Architectures · VLM↔Action Connection · Stellar VLA (continual) · RL for VLA.
1. The problem — why one VLA ≠ many tasks
Naively training (or merging) a single VLA over many tasks reliably underperforms per-task specialists. Six documented failure modes:
- Negative transfer / gradient conflict. Conflicting per-task gradients push shared weights toward compromises. HiMoE-VLA shows this directly — co-training a dense π0 across mismatched action spaces (joint-angle vs EE) lowers its score. It worsens under LoRA's low-rank capacity limit.
- Non-mergeability of specialists. Even merging separately-finetuned experts fails: MergeVLA finds naive merging gives near-zero success, because (a) LoRA adapters diverge in task-specific directions and (b) action-expert self-attention feedback spreads task information across blocks, preventing modular recombination.
- Multi-task conflict + instability at scale. As task count grows, joint optimization destabilizes. DyGRO-VLA documents that RL fine-tuning becomes unstable and less effective as tasks grow, and a t-SNE study shows single-task tuning drifts the task into an isolated cluster, away from the shared feature space.
- Catastrophic forgetting. Optimizing one task erodes others — DyGRO: LIBERO-Spatial RFT sharply drops unrelated LIBERO-Object success; the continual setting is worse (Stellar VLA, Pretrained VLAs Resist Forgetting).
- Routing confusion / skill fragmentation. In MoE policies, routing on low-level signals (e.g. diffusion noise level) fragments reusable skills across experts at skill boundaries, hurting transfer (SMoDP).
- Instruction collapse / wrong-task selection. With many tasks sharing scenes, the model latches onto visual shortcuts and ignores the language, executing the wrong task — the "Information Collapse" of LangForce and the "observation leakage" of DISC.
Root tension: a fixed-capacity shared backbone must both specialize (per-task precision) and generalize (cross-task competency) — an information-theoretic trade-off ("capability and robustness cannot both be free").
2. Anchor case study — MergeVLA (CVPR 2026)
MergeVLA (arXiv 2511.18810; Fu et al., Univ. of Queensland) tackles the non-mergeability failure directly: compose per-task specialists into one generalist instead of jointly training.
Diagnosis — two barriers to merging:
- LoRA divergence — finetuning drives the VLM's LoRA adapters toward divergent task-specific directions beyond what merging can unify.
- Action-expert entanglement — self-attention feedback creates inter-block dependencies, so task information smears across layers and can't be recombined modularly.
Three fixes:
- Sparse task-masked LoRA — adapters are sparsely activated via task masks, keeping parameters consistent and reducing irreconcilable conflicts.
- Cross-attention-only action expert — replaces self-attention with cross-attention-only blocks, keeping each task's specialization localized and composable.
- Test-time task router — when the task is unknown, an unsupervised router picks the task mask + expert head from the initial observation.
Result: across LIBERO, LIBERO-Plus, RoboTwin, and a real SO101 arm, MergeVLA matches or exceeds individually finetuned experts — i.e. one merged checkpoint ≈ N specialists. The architectural lesson generalizes: avoid self-attention-driven entanglement in the action expert if you want composability.
3. The solution landscape
| Cluster | Idea | Exemplars |
|---|---|---|
| A. Model merging | train specialists, compose them into one checkpoint | [MergeVLA](/Heungwoo/research/wiki/CVPR-2026-MergeVLA) (task-mask LoRA + cross-attn-only + test-time router) |
| B. Mixture-of-Experts (route per task/skill) | give each task/skill its own expert; shared backbone only consolidates | [HiMoE-VLA](/Heungwoo/research/wiki/Review-HiMoE-VLA) (hierarchical AS-MoE/HB-MoE, 32 experts top-4), [SMoDP](/Heungwoo/research/wiki/RSS-2026-Semantically-Structured-Mixture-of-Experts-for-Compositional) (language-defined skill routing), [SMP](/Heungwoo/research/wiki/ICLR-2026-MoE-Diffusion-Skills) (orthogonal skill basis), [HALO](/Heungwoo/research/wiki/Review-HALO) (3-expert MoT) |
| C. Optimization / gradient control | resolve gradient conflict during training | [DyGRO-VLA](/Heungwoo/research/wiki/Review-DyGRO-VLA) (dynamic grouped residual optimization for cross-task RL), orthogonal gradient projection (2601.09684) |
| D. Skill decomposition / hierarchy | split a task into sub-skills, route per-phase | [SMoDP](/Heungwoo/research/wiki/RSS-2026-Semantically-Structured-Mixture-of-Experts-for-Compositional), [HALO](/Heungwoo/research/wiki/Review-HALO) (EM-CoT), [Move-Then-Operate](/Heungwoo/research/wiki/ICML-2026-Move-Then-Operate) (phase experts), π0.5 hierarchy |
| E. Instruction grounding (right-task selection) | force the policy to actually follow language | [DISC](/Heungwoo/research/wiki/RSS-2026-DISC) (decouple instruction from state), [LangForce](/Heungwoo/research/wiki/ICML-2026-LangForce) (penalize the vision shortcut), [InstructVLA](/Heungwoo/research/wiki/ICLR-2026-InstructVLA) (instruction tuning) |
| F. Continual / knowledge-structured | add tasks over time without forgetting | [Stellar VLA](/Heungwoo/research/wiki/Review-Stellar-VLA) (Dirichlet-Process self-evolving knowledge space), [Pretrained VLAs Resist Forgetting](/Heungwoo/research/wiki/ICML-2026-VLA-Forgetting) |
| G. Modular adapters / plug-ins | isolate task-specific capacity in small modules | [DECO](/Heungwoo/research/wiki/ICML-2026-DECO) (plug-in tactile adapter), [GuidedVLA](/Heungwoo/research/wiki/RSS-2026-GuidedVLA) (specialized attention heads), DexVLA plug-in expert |
4. How the clusters relate — a decision guide
- You already have strong per-task specialists? → merge them (A, MergeVLA) — cheapest path to one checkpoint, no joint retraining.
- Training one model on many tasks from scratch, heterogeneous action spaces/embodiments? → routed MoE (B, HiMoE-VLA) so mismatched tasks don't average out; add skill-semantic routing (SMoDP) to stop fragmentation.
- RL fine-tuning across tasks? → gradient/optimization control (C, DyGRO-VLA) to avoid the isolated-cluster forgetting.
- Long-horizon / compositional tasks? → skill decomposition (D) — route per sub-skill.
- Model executes the wrong task / ignores language? → instruction grounding (E, DISC/LangForce).
- Tasks arrive over time? → continual, knowledge-structured (F, Stellar VLA).
The unifying insight: every fix is a way to give each task its own parameters/pathway while sharing a common substrate — whether by masking (MergeVLA), routing (MoE), projecting gradients (DyGRO), or structuring a knowledge space (Stellar). Dense parameter sharing is exactly what causes the interference; structured, sparse specialization is the cure.
5. Open questions
- Merge vs jointly-train vs route — which wins at matched compute? No head-to-head across MergeVLA (merge), HiMoE-VLA (MoE), and dense co-training on the same task suite.
- How many experts/masks before the router itself becomes the bottleneck? Test-time task inference (MergeVLA) and top-k gating (HiMoE) are unproven at hundreds of tasks.
- Does skill-semantic routing (SMoDP) generalize to unseen skill compositions, or only seen skills?
- Capability↔robustness bound — is the information-theoretic trade-off fundamental, or can structured specialization sidestep it?
- Cross-embodiment × multi-task — HiMoE handles action-space heterogeneity; combining with single-checkpoint multi-robot deployment is open.
6. Links
- Anchor: MergeVLA (2511.18810, CVPR 2026, project)
- MoE / routing: HiMoE-VLA · SMoDP · SMP · HALO
- Optimization / continual: DyGRO-VLA · Stellar VLA · Pretrained VLAs Resist Forgetting
- IROS 2026 🆕: VLA-RL (scalable RL for masterful general manipulation) · LAR-MoE (latent-aligned routing MoE in imitation — cluster B) · AtomVLA (subtask grounding + offline GRPO — cluster C) · Responsibility-Induced Specialized Experts (MoE-VLA for humanoid loco-manip) — context: IROS 2026 survey
- Instruction grounding: DISC · LangForce · InstructVLA
- Adapters / decomposition: DECO · GuidedVLA · Move-Then-Operate
- Taxonomy: VLA Architectures · VLM↔Action Connection