CVPR 2026 MergeVLA - Heungwoo/research GitHub Wiki
MergeVLA — Cross-Skill Model Merging Toward a Generalist VLA
Venue: CVPR 2026 · arXiv: 2511.18810 (Nov 2025, rev. Mar 2026) · project: mergevla.github.io Authors: Yuxia Fu, Zhizhen Zhang, Yuqi Zhang, Zijian Wang, Zi Huang, Yadan Luo (University of Queensland) The anchor case study in Multi-Task VLA — compose per-task specialists into one generalist instead of jointly training.
Problem
VLA models fine-tune well on a single task/embodiment but degrade in multi-skill settings, and — the paper's focus — directly merging VLA experts trained on different tasks yields near-zero success. Two sources of non-mergeability:
- LoRA divergence in the VLM — "finetuning drives LoRA adapters in the VLM backbone toward divergent, task-specific directions beyond the capacity of existing merging methods to unify."
- Action-expert entanglement — "action experts develop inter-block dependencies through self-attention feedback, causing task information to spread across layers and preventing modular recombination."
Method
Three architectural changes make specialists composable:
- Sparse task-masked LoRA — adapters are sparsely activated via task masks, retaining consistent parameters and reducing irreconcilable conflicts in the VLM.
- Cross-attention-only action expert — replaces self-attention with cross-attention-only blocks so each task's specialization stays localized and composable (no cross-block smearing).
- Test-time task router — when the task is unknown, adaptively selects the task mask + expert head from the initial observation, enabling unsupervised task inference.
Results
Across LIBERO, LIBERO-Plus, RoboTwin, and a real SO101 robotic arm, MergeVLA achieves performance comparable to or exceeding individually finetuned experts — one merged checkpoint ≈ N specialists — with robust generalization across tasks, embodiments, and environments. (The abstract reports relative parity rather than absolute per-benchmark numbers.)
Significance
MergeVLA reframes multi-task VLA as a merging problem and gives a transferable architectural lesson: self-attention in the action expert entangles tasks; cross-attention-only + task-masked LoRA keeps them composable. See the failure-mode taxonomy and the full solution landscape in Multi-Task VLA.
Links
- arXiv: 2511.18810 · project: mergevla.github.io
- Review: Multi-Task VLA · VLA Architectures · VLM↔Action Connection
← Back to Multi-Task VLA · Home