Review Multitask VLA - Heungwoo/research GitHub Wiki

In-Depth Review — Multi-Task VLA: why one policy struggles across tasks, and how to fix it

Scope: why a single Vision-Language-Action model degrades when asked to do many tasks at once, and the 2025–2026 research fixing it. Anchor case study: MergeVLA (CVPR 2026). Companions: VLA Architectures · VLM↔Action Connection · Stellar VLA (continual) · RL for VLA.


1. The problem — why one VLA ≠ many tasks

Naively training (or merging) a single VLA over many tasks reliably underperforms per-task specialists. Six documented failure modes:

  1. Negative transfer / gradient conflict. Conflicting per-task gradients push shared weights toward compromises. HiMoE-VLA shows this directly — co-training a dense π0 across mismatched action spaces (joint-angle vs EE) lowers its score. It worsens under LoRA's low-rank capacity limit.
  2. Non-mergeability of specialists. Even merging separately-finetuned experts fails: MergeVLA finds naive merging gives near-zero success, because (a) LoRA adapters diverge in task-specific directions and (b) action-expert self-attention feedback spreads task information across blocks, preventing modular recombination.
  3. Multi-task conflict + instability at scale. As task count grows, joint optimization destabilizes. DyGRO-VLA documents that RL fine-tuning becomes unstable and less effective as tasks grow, and a t-SNE study shows single-task tuning drifts the task into an isolated cluster, away from the shared feature space.
  4. Catastrophic forgetting. Optimizing one task erodes others — DyGRO: LIBERO-Spatial RFT sharply drops unrelated LIBERO-Object success; the continual setting is worse (Stellar VLA, Pretrained VLAs Resist Forgetting).
  5. Routing confusion / skill fragmentation. In MoE policies, routing on low-level signals (e.g. diffusion noise level) fragments reusable skills across experts at skill boundaries, hurting transfer (SMoDP).
  6. Instruction collapse / wrong-task selection. With many tasks sharing scenes, the model latches onto visual shortcuts and ignores the language, executing the wrong task — the "Information Collapse" of LangForce and the "observation leakage" of DISC.

Root tension: a fixed-capacity shared backbone must both specialize (per-task precision) and generalize (cross-task competency) — an information-theoretic trade-off ("capability and robustness cannot both be free").


2. Anchor case study — MergeVLA (CVPR 2026)

MergeVLA (arXiv 2511.18810; Fu et al., Univ. of Queensland) tackles the non-mergeability failure directly: compose per-task specialists into one generalist instead of jointly training.

Diagnosis — two barriers to merging:

  • LoRA divergence — finetuning drives the VLM's LoRA adapters toward divergent task-specific directions beyond what merging can unify.
  • Action-expert entanglement — self-attention feedback creates inter-block dependencies, so task information smears across layers and can't be recombined modularly.

Three fixes:

  1. Sparse task-masked LoRA — adapters are sparsely activated via task masks, keeping parameters consistent and reducing irreconcilable conflicts.
  2. Cross-attention-only action expert — replaces self-attention with cross-attention-only blocks, keeping each task's specialization localized and composable.
  3. Test-time task router — when the task is unknown, an unsupervised router picks the task mask + expert head from the initial observation.

Result: across LIBERO, LIBERO-Plus, RoboTwin, and a real SO101 arm, MergeVLA matches or exceeds individually finetuned experts — i.e. one merged checkpoint ≈ N specialists. The architectural lesson generalizes: avoid self-attention-driven entanglement in the action expert if you want composability.


3. The solution landscape

Cluster Idea Exemplars
A. Model merging train specialists, compose them into one checkpoint [MergeVLA](/Heungwoo/research/wiki/CVPR-2026-MergeVLA) (task-mask LoRA + cross-attn-only + test-time router)
B. Mixture-of-Experts (route per task/skill) give each task/skill its own expert; shared backbone only consolidates [HiMoE-VLA](/Heungwoo/research/wiki/Review-HiMoE-VLA) (hierarchical AS-MoE/HB-MoE, 32 experts top-4), [SMoDP](/Heungwoo/research/wiki/RSS-2026-Semantically-Structured-Mixture-of-Experts-for-Compositional) (language-defined skill routing), [SMP](/Heungwoo/research/wiki/ICLR-2026-MoE-Diffusion-Skills) (orthogonal skill basis), [HALO](/Heungwoo/research/wiki/Review-HALO) (3-expert MoT)
C. Optimization / gradient control resolve gradient conflict during training [DyGRO-VLA](/Heungwoo/research/wiki/Review-DyGRO-VLA) (dynamic grouped residual optimization for cross-task RL), orthogonal gradient projection (2601.09684)
D. Skill decomposition / hierarchy split a task into sub-skills, route per-phase [SMoDP](/Heungwoo/research/wiki/RSS-2026-Semantically-Structured-Mixture-of-Experts-for-Compositional), [HALO](/Heungwoo/research/wiki/Review-HALO) (EM-CoT), [Move-Then-Operate](/Heungwoo/research/wiki/ICML-2026-Move-Then-Operate) (phase experts), π0.5 hierarchy
E. Instruction grounding (right-task selection) force the policy to actually follow language [DISC](/Heungwoo/research/wiki/RSS-2026-DISC) (decouple instruction from state), [LangForce](/Heungwoo/research/wiki/ICML-2026-LangForce) (penalize the vision shortcut), [InstructVLA](/Heungwoo/research/wiki/ICLR-2026-InstructVLA) (instruction tuning)
F. Continual / knowledge-structured add tasks over time without forgetting [Stellar VLA](/Heungwoo/research/wiki/Review-Stellar-VLA) (Dirichlet-Process self-evolving knowledge space), [Pretrained VLAs Resist Forgetting](/Heungwoo/research/wiki/ICML-2026-VLA-Forgetting)
G. Modular adapters / plug-ins isolate task-specific capacity in small modules [DECO](/Heungwoo/research/wiki/ICML-2026-DECO) (plug-in tactile adapter), [GuidedVLA](/Heungwoo/research/wiki/RSS-2026-GuidedVLA) (specialized attention heads), DexVLA plug-in expert

4. How the clusters relate — a decision guide

  • You already have strong per-task specialists? → merge them (A, MergeVLA) — cheapest path to one checkpoint, no joint retraining.
  • Training one model on many tasks from scratch, heterogeneous action spaces/embodiments? → routed MoE (B, HiMoE-VLA) so mismatched tasks don't average out; add skill-semantic routing (SMoDP) to stop fragmentation.
  • RL fine-tuning across tasks? → gradient/optimization control (C, DyGRO-VLA) to avoid the isolated-cluster forgetting.
  • Long-horizon / compositional tasks? → skill decomposition (D) — route per sub-skill.
  • Model executes the wrong task / ignores language? → instruction grounding (E, DISC/LangForce).
  • Tasks arrive over time? → continual, knowledge-structured (F, Stellar VLA).

The unifying insight: every fix is a way to give each task its own parameters/pathway while sharing a common substrate — whether by masking (MergeVLA), routing (MoE), projecting gradients (DyGRO), or structuring a knowledge space (Stellar). Dense parameter sharing is exactly what causes the interference; structured, sparse specialization is the cure.


5. Open questions

  1. Merge vs jointly-train vs route — which wins at matched compute? No head-to-head across MergeVLA (merge), HiMoE-VLA (MoE), and dense co-training on the same task suite.
  2. How many experts/masks before the router itself becomes the bottleneck? Test-time task inference (MergeVLA) and top-k gating (HiMoE) are unproven at hundreds of tasks.
  3. Does skill-semantic routing (SMoDP) generalize to unseen skill compositions, or only seen skills?
  4. Capability↔robustness bound — is the information-theoretic trade-off fundamental, or can structured specialization sidestep it?
  5. Cross-embodiment × multi-task — HiMoE handles action-space heterogeneity; combining with single-checkpoint multi-robot deployment is open.

6. Links

← Back to Reviews · Home