NeurIPS 2025 ChatVLA 2 - Heungwoo/research GitHub Wiki

ChatVLA-2 — VLA with Open-World Embodied Reasoning

Venue: NeurIPS 2025 · Authors: Midea Group + East China Normal University · arXiv: 2505.21906 Category: VLA Architecture / CoT Reasoning

Approach diagram

flowchart LR
  V[Vision] --> VLM[VLM backbone]
  L[Language] --> VLM
  VLM --> MoE[Dynamic MoE:<br/>reasoning experts ⊕ action experts]
  MoE -- reasoning path --> R[VLM-preserved reasoning output]
  MoE -- action path --> AE[Action decoder]
  AE --> A[Actions]
  note[2-stage pipeline:<br/>1. co-train image-text ⊕ robot data<br/>2. freeze VLM, train action expert only]
Loading

Problem

Fine-tuning a VLM into a VLA nearly always destroys open-world reasoning — the model gets good at its training tasks but loses the broad commonsense / math / language-following that its VLM pretraining provided. VLM4VLA documents the phenomenon; ChatVLA-2 attacks the root cause.

Built on DexVLA as the foundational architecture, with Qwen2-VL as the core VLM. Because Qwen2-VL's LLM lacks native MoE support, the authors insert a dynamic mixture-of-experts inside the VLM backbone to disentangle multimodal-understanding vs. robotic-action feature spaces without disrupting the pretrained structure:

  • 8 experts total, top-2 selected per inference via an adaptive gating network conditioned on the visual/textual input;
  • Some experts capture shared multimodal/spatial-reasoning features, others specialize on task-specific (manipulation) features;
  • A separate action expert (inherited from DexVLA) decodes actions, aligned with the model's internal reasoning.

Two-stage training pipeline (paper §3.3):

  1. Open-world reasoning stage — co-train on image-text data ⊕ robot data, simultaneously training robotic actions, to preserve pretrained multimodal knowledge and establish reasoning↔action connections.
  2. Reasoning-following stage — freeze the entire VLM and train only the action expert, so open-world reasoning is preserved while instruction/reasoning-following in action execution is enhanced.
  • In-domain manipulation: ChatVLA-2 performs comparably to strong imitation baselines (11/13 vs DexVLA 12/13 and π0 12/13 on the math-matching task; all near-saturated). The paper is explicit that it "does not significantly outperform models like π0 and DexVLA" in-domain.
  • Open-world is where it wins decisively. On the math-matching game (robot reads whiteboard equations and picks matching number cards) ChatVLA-2 reaches an 82.7% manipulation success rate (43/52) with OCR score 3.58 and math-reasoning score 1.73, while baselines (Octo, Diffusion Policy, OpenVLA, DexVLA, π0) score 0–10/52 — these reasoning/OCR abilities were never explicitly trained in the VLA. A separate toy-placement task shows analogous open-world spatial-reasoning gains over unseen objects/directions.
  • Ablation (paper Table 4): removing Stage 2 collapses open-world control to ~23%; removing Stage 1 drives open-world reasoning to near-zero — both stages are necessary (reasoning is generated in Stage 1 but only injected into action execution in Stage 2).

Significance

One of the strongest published demonstrations that MoE is an effective mechanism for the reasoning-vs-action feature conflict — and that simply freezing the VLM in a second stage suffices to preserve open-world reasoning while injecting it into control. Seeds the ICLR 2026 MoE cluster (HiMoE-VLA, AdaMoE). Note it is a single-system, MoE-in-backbone design (not an explicit fast/slow dual-system like Fast-in-Slow or ThinkAct), though it is frequently grouped with them as a reasoning-preserving VLA.

Links

Related pages

← Back to NeurIPS-2025

⚠️ **GitHub.com Fallback** ⚠️