CVPR 2026 MergeVLA - Heungwoo/research GitHub Wiki

MergeVLA — Cross-Skill Model Merging Toward a Generalist VLA

Venue: CVPR 2026 · arXiv: 2511.18810 (Nov 2025, rev. Mar 2026) · project: mergevla.github.io Authors: Yuxia Fu, Zhizhen Zhang, Yuqi Zhang, Zijian Wang, Zi Huang, Yadan Luo (University of Queensland) The anchor case study in Multi-Task VLA — compose per-task specialists into one generalist instead of jointly training.

Problem

VLA models fine-tune well on a single task/embodiment but degrade in multi-skill settings, and — the paper's focus — directly merging VLA experts trained on different tasks yields near-zero success. Two sources of non-mergeability:

  1. LoRA divergence in the VLM — "finetuning drives LoRA adapters in the VLM backbone toward divergent, task-specific directions beyond the capacity of existing merging methods to unify."
  2. Action-expert entanglement — "action experts develop inter-block dependencies through self-attention feedback, causing task information to spread across layers and preventing modular recombination."

Method

Three architectural changes make specialists composable:

  1. Sparse task-masked LoRA — adapters are sparsely activated via task masks, retaining consistent parameters and reducing irreconcilable conflicts in the VLM.
  2. Cross-attention-only action expert — replaces self-attention with cross-attention-only blocks so each task's specialization stays localized and composable (no cross-block smearing).
  3. Test-time task router — when the task is unknown, adaptively selects the task mask + expert head from the initial observation, enabling unsupervised task inference.

Results

Across LIBERO, LIBERO-Plus, RoboTwin, and a real SO101 robotic arm, MergeVLA achieves performance comparable to or exceeding individually finetuned experts — one merged checkpoint ≈ N specialists — with robust generalization across tasks, embodiments, and environments. (The abstract reports relative parity rather than absolute per-benchmark numbers.)

Significance

MergeVLA reframes multi-task VLA as a merging problem and gives a transferable architectural lesson: self-attention in the action expert entangles tasks; cross-attention-only + task-masked LoRA keeps them composable. See the failure-mode taxonomy and the full solution landscape in Multi-Task VLA.

Links

← Back to Multi-Task VLA · Home