NeurIPS 2025 ChatVLA 2 - Heungwoo/research GitHub Wiki
Venue: NeurIPS 2025 · Authors: Midea Group + East China Normal University · arXiv: 2505.21906 Category: VLA Architecture / CoT Reasoning
flowchart LR
V[Vision] --> VLM[VLM backbone]
L[Language] --> VLM
VLM --> MoE[Dynamic MoE:<br/>reasoning experts ⊕ action experts]
MoE -- reasoning path --> R[VLM-preserved reasoning output]
MoE -- action path --> AE[Action decoder]
AE --> A[Actions]
note[2-stage pipeline:<br/>1. co-train image-text ⊕ robot data<br/>2. freeze VLM, train action expert only]
Fine-tuning a VLM into a VLA nearly always destroys open-world reasoning — the model gets good at its training tasks but loses the broad commonsense / math / language-following that its VLM pretraining provided. VLM4VLA documents the phenomenon; ChatVLA-2 attacks the root cause.
Built on DexVLA as the foundational architecture, with Qwen2-VL as the core VLM. Because Qwen2-VL's LLM lacks native MoE support, the authors insert a dynamic mixture-of-experts inside the VLM backbone to disentangle multimodal-understanding vs. robotic-action feature spaces without disrupting the pretrained structure:
- 8 experts total, top-2 selected per inference via an adaptive gating network conditioned on the visual/textual input;
- Some experts capture shared multimodal/spatial-reasoning features, others specialize on task-specific (manipulation) features;
- A separate action expert (inherited from DexVLA) decodes actions, aligned with the model's internal reasoning.
Two-stage training pipeline (paper §3.3):
- Open-world reasoning stage — co-train on image-text data ⊕ robot data, simultaneously training robotic actions, to preserve pretrained multimodal knowledge and establish reasoning↔action connections.
- Reasoning-following stage — freeze the entire VLM and train only the action expert, so open-world reasoning is preserved while instruction/reasoning-following in action execution is enhanced.
- In-domain manipulation: ChatVLA-2 performs comparably to strong imitation baselines (11/13 vs DexVLA 12/13 and π0 12/13 on the math-matching task; all near-saturated). The paper is explicit that it "does not significantly outperform models like π0 and DexVLA" in-domain.
- Open-world is where it wins decisively. On the math-matching game (robot reads whiteboard equations and picks matching number cards) ChatVLA-2 reaches an 82.7% manipulation success rate (43/52) with OCR score 3.58 and math-reasoning score 1.73, while baselines (Octo, Diffusion Policy, OpenVLA, DexVLA, π0) score 0–10/52 — these reasoning/OCR abilities were never explicitly trained in the VLA. A separate toy-placement task shows analogous open-world spatial-reasoning gains over unseen objects/directions.
- Ablation (paper Table 4): removing Stage 2 collapses open-world control to ~23%; removing Stage 1 drives open-world reasoning to near-zero — both stages are necessary (reasoning is generated in Stage 1 but only injected into action execution in Stage 2).
One of the strongest published demonstrations that MoE is an effective mechanism for the reasoning-vs-action feature conflict — and that simply freezing the VLM in a second stage suffices to preserve open-world reasoning while injecting it into control. Seeds the ICLR 2026 MoE cluster (HiMoE-VLA, AdaMoE). Note it is a single-system, MoE-in-backbone design (not an explicit fast/slow dual-system like Fast-in-Slow or ThinkAct), though it is frequently grouped with them as a reasoning-preserving VLA.
- arXiv: https://arxiv.org/abs/2505.21906
- NeurIPS virtual page: https://neurips.cc/virtual/2025/poster/120199
- Project: https://chatvla-2.github.io/
- Fast-in-Slow · ThinkAct (dual-system siblings)
- HiMoE-VLA (ICLR 2026 MoE descendant)
- Review: VLA Architectures — §5.F Hierarchical / dual-system / MoE
← Back to NeurIPS-2025