ICLR 2026 VLM4VLA - Heungwoo/research GitHub Wiki

VLM4VLA — Revisiting Vision-Language Models in VLA

Venue: ICLR 2026 Category: VLA Architecture — Analysis Trend tag: Surprising finding (Trend 7) Authors: Jianke Zhang, Xiaoyu Chen, Yanjiang Guo, Yucheng Hu, Jianyu Chen, et al. (incl. Qwen team)

Approach diagram

flowchart TB
  subgraph Setup[Controlled comparison]
    VLM_A[VLM A<br/>strong general/VQA capability] --> H[Minimal adaptation pipeline<br/>same VLA harness, same hyperparams]
    VLM_B[VLM B<br/>weak general/VQA capability] --> H
    VLM_C[VLM C] --> H
  end
  H --> M[Measure downstream<br/>manipulation success]
  M --> R[Correlation: VLM capability ↔ manipulation success]
  R --> F[Result: ≈ ZERO]
Loading

Problem

The dominant heuristic in VLA research has been "use the strongest available VLM backbone." But there was no published systematic comparison of how VLM choice affects downstream manipulation under matched training conditions. The field has been assuming without evidence.

Method

VLM4VLA is a minimal adaptation pipeline that converts a general-purpose VLM into a VLA policy using only a small set of new learnable parameters, so backbones can be compared fairly under identical conditions. Nine open-source VLM backbones are swapped into the same VLA training pipeline, all other hyperparameters held constant: Qwen2.5VL-3B, Qwen2.5VL-7B, Qwen3VL-2B, Qwen3VL-4B, Qwen3VL-8B, Qwen3VL-30B-A3B, PaliGemma-1, PaliGemma-2, and Kosmos-2. Each backbone's general VLM capability is measured separately, then compared against downstream manipulation success on CALVIN ABC-D, SimplerEnv (Bridge), and LIBERO-Long (simulation only; no real-robot eval).

Results

No obvious positive correlation between VLM general capability and VLA performance. VLM initialization still beats training from scratch, but the strongest VLM by general/VQA standards is not the strongest VLA — some smaller VLMs outperform much larger ones on robot tasks. Concrete examples from the main tables (Tables 1–2):

  • On CALVIN ABC-D (avg. completed tasks per sequence, max 5), the small Qwen3VL-2B scores 4.142 — the best of all backbones, ahead of the larger Qwen2.5VL-7B (4.057), Qwen3VL-8B (4.035), and even Qwen3VL-30B-A3B (4.075).
  • On LIBERO-10, Qwen3VL-2B reaches 55.8% success, again topping its larger Qwen3VL siblings (4B = 44.4, 8B = 46.2).
  • The smallest model, Kosmos-2 (1.7B), ties for the highest SimplerEnv-Bridge score (60.4), underscoring that scale/general-capability is not predictive of control.

Two further results sharpen the diagnosis:

  • Embodied-skill enhancement doesn't transfer. Fine-tuning VLMs on seven auxiliary embodied tasks (embodied QA, visual pointing, depth estimation, etc.) does not guarantee better downstream control — better embodied-benchmark scores ≠ better action policy.
  • The vision module, not the language module, is the bottleneck. Modality-level ablations isolate the VLM's visual encoder as the primary performance bottleneck. Injecting control-relevant supervision into the vision encoder yields consistent gains even when the encoder is kept frozen during downstream fine-tuning (Table 4: on SimplerEnv-Bridge with Qwen3VL-4B, action-info VLM fine-tuning with an unfrozen vision encoder lifts the policy from 56.3 to 59.4, +3.1; the same supervision applied while the encoder is frozen gives a much larger +18.1 in the frozen-VLA regime), pointing to a persistent domain gap between VLM pretraining objectives and embodied action-planning.

Significance

The most practically consequential "surprising finding" in 2026 VLA research. Invalidates the default VLM-selection heuristic and motivates:

  • Action-aware backbone selection
  • Action-aware VLM pretraining objectives
  • New proxy benchmarks that actually predict downstream control performance

📖 In-depth review

For a long-form review with all 9 benchmarked VLM backbones, full ablation tables, and limitations: In-Depth Review of VLM4VLA.

Links

Related pages

  • π0.6 (uses Gemma3-4B — informative reference point)
  • UniVLA (action-as-first-class-modality alternative)

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️