Review VLM4VLA - Heungwoo/research GitHub Wiki

In-Depth Review โ€” VLM4VLA: Revisiting Vision-Language Models in VLA

Paper: VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models ยท ICLR 2026 (poster) ยท arXiv: 2601.03309 (v1 Jan 6, 2026 ยท v2 May 30, 2026) Affiliation: Tsinghua IIIS (Jianke Zhang, Xiaoyu Chen, Yanjiang Guo, Yucheng Hu, Jianyu Chen) ร— Qwen Team, Alibaba (Qiuyue Wang, Mingsheng Li, Jiajun Zhang, Shuai Bai, Junyang Lin) OpenReview: https://openreview.net/forum?id=tc2UsBeODW Related summary: VLM4VLA

This is the long-form review page complementing the VLM4VLA summary. All numbers and claims below are verified against the latest revision (arXiv 2601.03309v2, May 2026); the only substantive v1โ†’v2 change is a corrected ฯ€0 Calvin Task-3 value (0.786 โ†’ 0.686, fixing an internal sum inconsistency โ€” the 3.509 total was always correct).


1. TL;DR

Everyone assumed the recipe for a better VLA was "use the strongest available VLM." VLM4VLA runs the controlled experiment and finds: general VLM benchmark scores are a poor predictor of downstream manipulation performance. Kosmos-2 (1.7B) matches or beats Qwen3VL-30B-A3B (31B) on SimplerEnv. Fine-tuning VLMs on supposedly-helpful embodied auxiliary tasks (Robopoint, Vica, BridgeVQA, RoboBrain2) consistently hurts downstream VLA performance. Ablations locate the bottleneck: it's the vision encoder, not the language component. Freezing the vision encoder during VLA training is catastrophic (โˆ’42 points on SimplerEnv for Paligemma-1), while freezing word embeddings barely matters (ยฑ0.2 points). Fine-tuning the VLM's vision encoder on action-labeled real-world data closes the gap by +18 points. Bottom line: VLM competence is necessary but not sufficient for VLA competence, and the visual representations optimized for VQA are not the ones needed for low-level control.

2. Motivation โ€” the assumption nobody checked

"The majority of existing VLA methods have focused on developing more advanced network architectures โ€ฆ Limited attention has been given to a fundamental question at the core of VLA: how do the choice and specific capabilities of the underlying VLM affect the performance of VLA policies?"

Practitioners have been picking VLM backbones heuristically โ€” "use the newest/largest/best-MMMU-score model." This paper is the first systematic controlled study asking whether that heuristic is even valid.

3. Representative diagrams

Figure 1 from the paper โ€” evaluation pipeline

VLM4VLA framework (Figure 1 from Zhang et al., 2026)

Figure 1 of the VLM4VLA paper (Zhang et al., ICLR 2026, CC-BY-4.0). Left: the evaluation pipeline, which fine-tunes different VLM backbones on robot data and evaluates on downstream tasks, plus an optional fine-tuning stage on auxiliary embodied tasks. Top right: systematic investigation of three axes influencing VLM-to-VLA transfer โ€” the choice of VLM backbone, the impact of fine-tuning on auxiliary embodied tasks, and the influence of different training strategies (Frozen vs. fine-tuned different VLM modules). Bottom right: visualization of inconsistent performance of various VLM backbones across downstream tasks.

Figure 2 from the paper โ€” VLM4VLA network

VLM4VLA network architecture (Figure 2 from Zhang et al., 2026)

Figure 2 of the VLM4VLA paper. The minimal adaptation harness: a learnable action query token is appended after vision and text tokens; the VLM backbone processes the sequence; the action query output is decoded by a small MLP into an action chunk. All VLM parameters are tuned (vision encoder, LLM, word embeddings), and only <1% new parameters are introduced โ€” deliberately minimal so comparisons isolate the effect of the VLM backbone itself.

Our reconstruction as mermaid

flowchart TB
  subgraph Setup[Controlled matched-harness study]
    V1[Qwen2.5VL-3B]
    V2[Qwen2.5VL-7B]
    V3[Qwen3VL-2B / 4B / 8B / 30B-A3B]
    V4[Paligemma-1 / Paligemma-2]
    V5[Kosmos-2]
  end
  V1 & V2 & V3 & V4 & V5 --> H[VLM4VLA harness<br/>< 1% new params<br/>action query token + MLP head]
  H --> T{Train on robot data<br/>identical hyperparameters}
  T --> B1[Calvin ABC-D<br/>long-horizon task chains]
  T --> B2[SimplerEnv Bridge<br/>real-to-sim]
  T --> B3[Libero-Long<br/>10 long-horizon tasks]
  B1 & B2 & B3 --> R[Compare VLM's VQA score<br/>vs. VLA success rate]
  R --> F[Finding: rโ‰ˆ0.84 on Calvin,<br/>rโ‰ˆ-0.36 SimplerEnv,<br/>rโ‰ˆ-0.19 Libero]

  classDef result fill:#ffe8c2,stroke:#b47820,color:#000
  class F result
Loading

4. Method โ€” the VLM4VLA harness (<1% new params)

4.1 Minimal adaptation pipeline

To compare VLMs fairly, the authors deliberately avoid architectural bells & whistles. The harness introduces:

  • One learnable action query token, appended after the image and language tokens.
  • A small MLP policy head that decodes the VLM's output at the action-query position into actions.
  • No proprioceptive state โ€” visual + language inputs only (fair, reproducible, and isolates VLM effects).
  • No diffusion / flow matching โ€” they deliberately avoid this because stochastic sampling inflates evaluation variance.

Action decoding: $$a = \text{MLP}\left(\text{VLM}([\langle\text{img}\rangle \ldots \langle\text{text}\rangle \langle\text{ActionQuery}\rangle])\right)$$

4.2 Loss (Equation 1)

$$\mathcal{L} = \frac{1}{|\mathcal B|} \sum_{\mathcal B} \left(|a^{\text{pos}} - \hat a^{\text{pos}}|_2^2 + \text{BCE}(a^{\text{end}}, \hat a^{\text{end}})\right)$$

  • $a^{\text{pos}}$: relative end-effector position (regression).
  • $a^{\text{end}}$: discrete gripper open/close (BCE).
  • Pure maximum-likelihood imitation learning.

4.3 Parameter budget

  • All VLM parameters fine-tuned (vision encoder, LLM, word embeddings).
  • New learnable parameters added by the harness: <1% of the VLM's parameter count.
  • Single-view images at 224ร—224 (resized to model's native resolution if needed).
  • One unified hyperparameter set across all VLMs to guarantee fair comparison.

4.4 Benchmarks

Benchmark Setup Train steps Metric
Calvin ABC-D Train on scenes A, B, C; test on unseen D; 1000 sequences of 5 sub-tasks 30k Average tasks completed per sequence (0โ€“5)
SimplerEnv Bridge Train on real BridgeV2; test real-to-sim in 4 scenes (Pick Carrot, Pick Eggplant, Pick Spoon, Stack Cube); 24 trials/scene 50k Success rate
Libero-Long (Libero-10) 10 long-horizon tasks; 50 trials each 50k Success rate

4.5 VLMs tested

The abstract frames the study as evaluating "24 different VLMs that are either zero-shot or fine-tuned" โ€” i.e. the 9 distinct base backbones below plus their auxiliary-task fine-tuned variants (ยง6.2). The main Calvin/SimplerEnv/Libero comparison uses these 9 open-source base backbones spanning 1.7Bโ€“31B parameters:

Model Params
Qwen2.5VL-3B 3.8B
Qwen2.5VL-7B 8.3B
Qwen3VL-2B 2.1B
Qwen3VL-4B 4.4B
Qwen3VL-8B 8.8B
Qwen3VL-30B-A3B (MoE) 31.1B
Paligemma-1 2.9B
Paligemma-2 3.0B
Kosmos-2 1.7B

4.6 Expert baselines

  • OpenVLA (Llama2-7B, discrete action tokens)
  • ฯ€0 (Paligemma-1 + flow-matching action expert; modified to remove proprioception, ~3.1B)
  • ThinkAct (Qwen2.5VL-7B, RL-enhanced; uses proprioception โ€” a confound in direct comparisons)

5. Results โ€” the surprising numbers

5.1 Calvin ABC-D (Table 1) โ€” avg tasks completed per 5-step sequence

Model Task-1 Task-2 Task-3 Task-4 Task-5 Calvin โ†‘
OpenVLA* 0.792 0.644 0.499 0.368 0.245 2.548
ฯ€0* 0.896 0.785 0.686 0.610 0.532 3.509
Paligemma-1 0.914 0.813 0.692 0.599 0.488 3.506
Paligemma-2 0.901 0.775 0.669 0.575 0.486 3.406
Kosmos-2 0.878 0.721 0.591 0.498 0.408 3.096
Qwen2.5VL-3B 0.922 0.842 0.766 0.700 0.626 3.856
Qwen2.5VL-7B 0.935 0.864 0.807 0.758 0.693 4.057
Qwen3VL-2B 0.943 0.882 0.831 0.776 0.710 ๐Ÿ† 4.142
Qwen3VL-4B 0.933 0.857 0.790 0.719 0.644 3.943
Qwen3VL-8B 0.940 0.868 0.797 0.746 0.684 4.035
Qwen3VL-30B-A3B 0.939 0.877 0.820 0.757 0.682 4.075

Headline: Qwen3VL-2B (smallest Qwen3) is the overall winner at 4.142, beating the 31B MoE (4.075), 8B (4.035), and 7B (4.057). Paligemma-1 + VLM4VLA harness (3.506) matches ฯ€0's full flow-matching pipeline (3.509) โ€” strong sign that architectural sophistication matters less than backbone choice.

5.2 SimplerEnv Bridge & Libero-Long (Table 2) โ€” success rate (%)

Model Carrot Eggplant Spoon Cube Simpler โ†‘ Libero โ†‘
OpenVLA 4.2 0.0 0.0 12.5 4.2 53.7
ฯ€0* 62.5 100.0 54.2 25.0 60.4 46.0
ThinkAct (uses state) 37.5 70.8 58.3 8.7 43.8 ๐Ÿ† 70.9
Qwen2.5VL-3B 20.8 91.7 79.2 0.0 48.0 43.0
Qwen2.5VL-7B 12.5 100.0 75.0 0.0 46.8 45.0
Qwen3VL-2B 20.8 95.8 79.2 0.0 49.0 55.8
Qwen3VL-4B 54.2 95.8 75.0 0.0 56.3 44.4
Qwen3VL-8B 58.3 95.8 79.2 0.0 58.3 46.2
Qwen3VL-30B-A3B 29.2 79.2 70.8 0.0 44.8 46.8
Paligemma-1 50.0 91.7 75.0 4.2 55.3 44.2
Paligemma-2 75.0 75.0 79.2 0.0 57.3 46.2
Kosmos-2 (1.7B) 37.5 100.0 75.0 29.2 ๐Ÿ† 60.4 55.0

Headline: Kosmos-2 (the smallest model at 1.7B) wins SimplerEnv at 60.4%, tied with ฯ€0. Qwen3VL-30B-A3B (31B, MoE) lags at 44.8%. Libero-Long is won by ThinkAct (70.9%), but ThinkAct cheats by using proprioceptive state; without state, Qwen3VL-2B leads at 55.8%.

5.3 Correlation: general VLM score vs. VLA performance (Figure 3)

Linear regression across all 9 VLMs:

Benchmark Pearson r Rยฒ Interpretation
Calvin ABC-D +0.839 0.703 High correlation (scene variation โ‰ˆ QA-like)
SimplerEnv Bridge โˆ’0.358 0.128 No correlation, slightly negative
Libero-Long โˆ’0.194 0.038 No correlation

This is the core finding distilled into three numbers: strong VQA scores do not predict strong VLA performance on the benchmarks that measure actual manipulation.

5.4 From-scratch baseline (Table 8, Appendix)

Model Calvin (pretrained) Calvin (from scratch) ฮ” Simpler (pretrained) Simpler (from scratch) ฮ”
Qwen2.5VL-3B 3.856 1.381 โˆ’2.475 48.00 15.75 โˆ’32.25
Qwen2.5VL-7B 4.057 1.769 โˆ’2.288 46.75 18.20 โˆ’28.55
Paligemma-1 3.506 1.129 โˆ’2.377 55.25 14.50 โˆ’40.75

Pretraining is necessary โ€” training from scratch collapses performance. So VLM pretraining is necessary; it's just not sufficient, and stronger pretraining isn't strictly better.


6. Ablations โ€” locating the bottleneck

6.1 Vision vs. language (Table 3)

Freeze the vision encoder (or word embeddings) during VLA training; retrain; compare:

Model Full fine-tune + Freeze vision Vision ฮ” + Freeze word embeds Embed ฮ”
Qwen2.5VL-3B (Calvin) 3.856 2.855 โˆ’1.001 3.849 โˆ’0.007
Qwen2.5VL-3B (SimplerEnv) 48.00 23.95 โˆ’24.05 46.88 โˆ’1.12
Qwen2.5VL-7B (Calvin) 4.057 2.823 โˆ’1.234 3.874 โˆ’0.183
Qwen2.5VL-7B (SimplerEnv) 46.75 25.50 โˆ’21.25 48.96 +2.21
Paligemma-1 (Calvin) 3.506 0.495 โˆ’3.011 3.485 โˆ’0.021
Paligemma-1 (SimplerEnv) 55.25 13.25 โˆ’42.00 52.25 โˆ’3.00

Interpretation: Freezing the vision encoder is catastrophic (21โ€“42 points lost on SimplerEnv). Freezing word embeddings is essentially free (ยฑ0.2 points). The vision encoder is the primary bottleneck โ€” and fine-tuning the vision encoder matters more than just increasing the LLM's trainable parameters (frozen-vision 7.6B-tunable Qwen2.5VL-7B loses to fully-tuned 3.8B Qwen2.5VL-3B).

6.2 Auxiliary embodied task fine-tuning (Figure 4 / Table 10)

Fine-tune each VLM on a supposedly-helpful embodied auxiliary task before running the VLM4VLA harness; measure Calvin delta:

Auxiliary task Data ฮ” on Calvin
Robopoint (2D pointing) 1.43M samples โˆ’0.073
Vica-332k (spatial QA: size/distance/depth) 332k โˆ’0.009
BridgeVQA (spatial QA, real robot) โ€” โˆ’0.091
Robo2VLM (action-oriented VQA) 667k VQA / 176k traj โˆ’0.096
RoboBrain2 (embodied VQA at scale, official FT) โ€” โˆ’0.170
Omni-Generation (image+depth+seg gen + VQA) โ€” โˆ’0.181
VQA-Mix (general + embodied) โ€” โˆ’0.079

Every auxiliary task hurts. Best case (VQA-Mix) is โˆ’0.079. Worst (Omni-Generation) is โˆ’0.181. Most show increased variance on top of the degradation. The intuitive "embodied pretraining helps VLA" assumption fails empirically.

6.3 Vision-encoder fine-tuning with action supervision (Table 4)

Fine-tune Qwen3VL-4B's vision encoder on real-world BridgeV2 images with action information encoded as FAST tokens, then run VLM4VLA:

Condition Freeze vision during VLA Unfreeze vision during VLA
Baseline (no VLM fine-tune) 27.6 56.3
VLM FT, freeze vision during FT 28.0 (+0.4) 56.3 (+0.0)
VLM FT, unfreeze vision during FT 45.7 (+18.1) 59.4 (+3.1)

Interpretation: Fine-tuning VLM without touching the vision encoder doesn't help. Fine-tuning that does touch the vision encoder gains +18.1 points when VLA-time vision is frozen and +3.1 points when VLA-time vision is unfrozen. The gap between "VLM" and "VLA" is not about visual domain (sim vs. real) โ€” it's about which visual features are aligned with low-level control.

6.4 Image resolution (Table 7)

Raising input resolution from 224 โ†’ 512 โ†’ 768 doesn't improve Calvin scores when vision is unfrozen, and widens the frozen-vision gap (โˆ’1.00 โ†’ โˆ’1.12 โ†’ โˆ’1.31 for Qwen2.5VL-3B). Suggests frozen high-resolution vision leaks more spurious correlations.


7. Limitations (authors' own + critical reading)

7.1 Authors' stated limitations

  1. No real-world robot experiments. "A limitation of our work is the absence of experiments on physical robots." They argue fairness+reproducibility on shared sim benchmarks outweighs this, and note their Table-4 experiments with real-world Bridge images show that sim-vs-real visual gap is not the primary bottleneck.
  2. Simulation-only. This prevents other researchers from directly comparing physical-robot VLAs to VLM4VLA findings.
  3. VLM coverage is limited to 1Bโ€“10B parameter open-source models plus one MoE (Qwen3VL-30B-A3B). No proprietary frontier models, no larger dense models.
  4. Auxiliary task sampling is not exhaustive. 7 tasks tested, findings may be dataset-specific.
  5. Unified hyperparameters. Required for fairness, but may not be individually optimal for each architecture.

7.2 Reviewer's concerns worth flagging (not in the paper)

  • Regression loss + no diffusion simplifies evaluation variance but also handicaps every VLM equally against ฯ€0 (which has a flow-matching expert). The takeaway "Kosmos-2 matches ฯ€0" is harness-conditional โ€” with diffusion/flow heads the ranking may shift.
  • Deterministic MSE + BCE loss is probably underfit for multimodal action distributions. The paper implicitly argues the ranking is still informative, which is plausible but not proven.
  • Calvin benefits QwenVL, SimplerEnv benefits Paligemma/Kosmos-2 โ€” implicit evidence that benchmarks themselves select for different VLM capabilities. The paper frames this as a negative finding; a more positive reading is that different VLMs are specialists for different embodied skill sets, and the real research question is how to identify which VLM suits which benchmark.
  • No vision-encoder architecture comparison. Kosmos-2 uses CLIP; Paligemma uses SigLIP; Qwen uses a custom ViT. The paper ablates tuning but doesn't compare vision encoder families directly.
  • No study of how many action-conditioned vision-encoder tokens are needed to close the gap (the +18.1 result uses FAST tokens but the amount is not quantified cleanly in the body).
  • The paper's diagnostic is strong; the prescription is weak. It correctly identifies the bottleneck but only gestures at solutions (action-aware VLM pretraining). Nobody has delivered a backbone trained end-to-end for VLA use that outperforms repurposed VQA VLMs.

8. Takeaways โ€” what VLM4VLA changes about the field

  1. The VLM-selection heuristic is broken. "Pick the biggest/best-MMMU model" does not optimize VLA performance. Actually benchmark each candidate on your manipulation task.
  2. Vision encoder, not language, is the lever. Freezing word embeddings is nearly free; freezing vision is catastrophic. If you're cost-cutting, freeze the LLM, not the ViT.
  3. Embodied VQA pretraining is a red herring. Robopoint, Vica, BridgeVQA, Robo2VLM, RoboBrain2, Omni-Generation โ€” every auxiliary task tested hurts downstream VLA performance. Training a VLM on "robot-adjacent" data without touching the vision encoder with action signal is counterproductive.
  4. Action-supervised vision-encoder fine-tuning is the missing recipe. +18.1 points is the largest positive effect in the paper, and the only intervention that consistently helps.
  5. Counterintuitive best-in-class: Kosmos-2 (1.7B, 2023-vintage) matches the ฯ€0 production system on SimplerEnv. Something about its visual representation is better-aligned with low-level control than strong modern VLMs. The paper doesn't explain why; this is an open research question.

9. Open questions

  • What does a VLM trained for VLA look like? The paper argues current VLM pretraining objectives are misaligned with action grounding, but doesn't propose a concrete alternative objective.
  • Is there a cheap proxy benchmark that actually predicts downstream VLA performance? VQA scores don't; the paper doesn't suggest a replacement.
  • Does the result hold under diffusion / flow-matching heads? The MSE harness isolates VLM effects but may not translate directly to production flow-matching VLAs.
  • Does the result hold at frontier scale (100B+ VLMs)? Only sub-31B tested.
  • Why does Kosmos-2 punch above its weight? Possible hypotheses: it was explicitly trained for grounding (bounding-box tokens), it has denser visual tokens, its architecture is smaller and less distractible. Needs follow-up work.

10. Relation to this wiki

  • Qwen Team's VLA Program โ€” cross-paper review reading this paper as the diagnostic foundation of the Qwen VLA line (Qiuyue Wang / Mingsheng Li / Shuai Bai later authored Qwen-VLA)
  • VLM4VLA summary โ€” short version of this page
  • ManipBench โ€” the earlier CoRL 2025 paper arguing the same thing via a VLM benchmark on low-level manipulation reasoning
  • UniVLA โ€” action-as-first-class-modality direction, a possible answer
  • ฯ€0.6 โ€” uses Gemma3-4B (not in VLM4VLA's comparison set); informative reference point for production VLAs
  • Survey: VLA & Manipulation (ICLR 2026) โ€” where VLM4VLA is flagged as the "surprising finding" (Trend 7)

11. Links

โ† Back to ICLR-2026-VLM4VLA ยท ICLR-2026 ยท Home

โš ๏ธ **GitHub.com Fallback** โš ๏ธ