Review VLM4VLA - Heungwoo/research GitHub Wiki
Paper: VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models ยท ICLR 2026 (poster) ยท arXiv: 2601.03309 (v1 Jan 6, 2026 ยท v2 May 30, 2026) Affiliation: Tsinghua IIIS (Jianke Zhang, Xiaoyu Chen, Yanjiang Guo, Yucheng Hu, Jianyu Chen) ร Qwen Team, Alibaba (Qiuyue Wang, Mingsheng Li, Jiajun Zhang, Shuai Bai, Junyang Lin) OpenReview: https://openreview.net/forum?id=tc2UsBeODW Related summary: VLM4VLA
This is the long-form review page complementing the VLM4VLA summary. All numbers and claims below are verified against the latest revision (arXiv 2601.03309v2, May 2026); the only substantive v1โv2 change is a corrected ฯ0 Calvin Task-3 value (0.786 โ 0.686, fixing an internal sum inconsistency โ the 3.509 total was always correct).
Everyone assumed the recipe for a better VLA was "use the strongest available VLM." VLM4VLA runs the controlled experiment and finds: general VLM benchmark scores are a poor predictor of downstream manipulation performance. Kosmos-2 (1.7B) matches or beats Qwen3VL-30B-A3B (31B) on SimplerEnv. Fine-tuning VLMs on supposedly-helpful embodied auxiliary tasks (Robopoint, Vica, BridgeVQA, RoboBrain2) consistently hurts downstream VLA performance. Ablations locate the bottleneck: it's the vision encoder, not the language component. Freezing the vision encoder during VLA training is catastrophic (โ42 points on SimplerEnv for Paligemma-1), while freezing word embeddings barely matters (ยฑ0.2 points). Fine-tuning the VLM's vision encoder on action-labeled real-world data closes the gap by +18 points. Bottom line: VLM competence is necessary but not sufficient for VLA competence, and the visual representations optimized for VQA are not the ones needed for low-level control.
"The majority of existing VLA methods have focused on developing more advanced network architectures โฆ Limited attention has been given to a fundamental question at the core of VLA: how do the choice and specific capabilities of the underlying VLM affect the performance of VLA policies?"
Practitioners have been picking VLM backbones heuristically โ "use the newest/largest/best-MMMU-score model." This paper is the first systematic controlled study asking whether that heuristic is even valid.

Figure 1 of the VLM4VLA paper (Zhang et al., ICLR 2026, CC-BY-4.0). Left: the evaluation pipeline, which fine-tunes different VLM backbones on robot data and evaluates on downstream tasks, plus an optional fine-tuning stage on auxiliary embodied tasks. Top right: systematic investigation of three axes influencing VLM-to-VLA transfer โ the choice of VLM backbone, the impact of fine-tuning on auxiliary embodied tasks, and the influence of different training strategies (Frozen vs. fine-tuned different VLM modules). Bottom right: visualization of inconsistent performance of various VLM backbones across downstream tasks.

Figure 2 of the VLM4VLA paper. The minimal adaptation harness: a learnable action query token is appended after vision and text tokens; the VLM backbone processes the sequence; the action query output is decoded by a small MLP into an action chunk. All VLM parameters are tuned (vision encoder, LLM, word embeddings), and only <1% new parameters are introduced โ deliberately minimal so comparisons isolate the effect of the VLM backbone itself.
flowchart TB
subgraph Setup[Controlled matched-harness study]
V1[Qwen2.5VL-3B]
V2[Qwen2.5VL-7B]
V3[Qwen3VL-2B / 4B / 8B / 30B-A3B]
V4[Paligemma-1 / Paligemma-2]
V5[Kosmos-2]
end
V1 & V2 & V3 & V4 & V5 --> H[VLM4VLA harness<br/>< 1% new params<br/>action query token + MLP head]
H --> T{Train on robot data<br/>identical hyperparameters}
T --> B1[Calvin ABC-D<br/>long-horizon task chains]
T --> B2[SimplerEnv Bridge<br/>real-to-sim]
T --> B3[Libero-Long<br/>10 long-horizon tasks]
B1 & B2 & B3 --> R[Compare VLM's VQA score<br/>vs. VLA success rate]
R --> F[Finding: rโ0.84 on Calvin,<br/>rโ-0.36 SimplerEnv,<br/>rโ-0.19 Libero]
classDef result fill:#ffe8c2,stroke:#b47820,color:#000
class F result
To compare VLMs fairly, the authors deliberately avoid architectural bells & whistles. The harness introduces:
- One learnable action query token, appended after the image and language tokens.
- A small MLP policy head that decodes the VLM's output at the action-query position into actions.
- No proprioceptive state โ visual + language inputs only (fair, reproducible, and isolates VLM effects).
- No diffusion / flow matching โ they deliberately avoid this because stochastic sampling inflates evaluation variance.
Action decoding:
-
$a^{\text{pos}}$ : relative end-effector position (regression). -
$a^{\text{end}}$ : discrete gripper open/close (BCE). - Pure maximum-likelihood imitation learning.
- All VLM parameters fine-tuned (vision encoder, LLM, word embeddings).
- New learnable parameters added by the harness: <1% of the VLM's parameter count.
- Single-view images at 224ร224 (resized to model's native resolution if needed).
- One unified hyperparameter set across all VLMs to guarantee fair comparison.
| Benchmark | Setup | Train steps | Metric |
|---|---|---|---|
| Calvin ABC-D | Train on scenes A, B, C; test on unseen D; 1000 sequences of 5 sub-tasks | 30k | Average tasks completed per sequence (0โ5) |
| SimplerEnv Bridge | Train on real BridgeV2; test real-to-sim in 4 scenes (Pick Carrot, Pick Eggplant, Pick Spoon, Stack Cube); 24 trials/scene | 50k | Success rate |
| Libero-Long (Libero-10) | 10 long-horizon tasks; 50 trials each | 50k | Success rate |
The abstract frames the study as evaluating "24 different VLMs that are either zero-shot or fine-tuned" โ i.e. the 9 distinct base backbones below plus their auxiliary-task fine-tuned variants (ยง6.2). The main Calvin/SimplerEnv/Libero comparison uses these 9 open-source base backbones spanning 1.7Bโ31B parameters:
| Model | Params |
|---|---|
| Qwen2.5VL-3B | 3.8B |
| Qwen2.5VL-7B | 8.3B |
| Qwen3VL-2B | 2.1B |
| Qwen3VL-4B | 4.4B |
| Qwen3VL-8B | 8.8B |
| Qwen3VL-30B-A3B (MoE) | 31.1B |
| Paligemma-1 | 2.9B |
| Paligemma-2 | 3.0B |
| Kosmos-2 | 1.7B |
- OpenVLA (Llama2-7B, discrete action tokens)
- ฯ0 (Paligemma-1 + flow-matching action expert; modified to remove proprioception, ~3.1B)
- ThinkAct (Qwen2.5VL-7B, RL-enhanced; uses proprioception โ a confound in direct comparisons)
| Model | Task-1 | Task-2 | Task-3 | Task-4 | Task-5 | Calvin โ |
|---|---|---|---|---|---|---|
| OpenVLA* | 0.792 | 0.644 | 0.499 | 0.368 | 0.245 | 2.548 |
| ฯ0* | 0.896 | 0.785 | 0.686 | 0.610 | 0.532 | 3.509 |
| Paligemma-1 | 0.914 | 0.813 | 0.692 | 0.599 | 0.488 | 3.506 |
| Paligemma-2 | 0.901 | 0.775 | 0.669 | 0.575 | 0.486 | 3.406 |
| Kosmos-2 | 0.878 | 0.721 | 0.591 | 0.498 | 0.408 | 3.096 |
| Qwen2.5VL-3B | 0.922 | 0.842 | 0.766 | 0.700 | 0.626 | 3.856 |
| Qwen2.5VL-7B | 0.935 | 0.864 | 0.807 | 0.758 | 0.693 | 4.057 |
| Qwen3VL-2B | 0.943 | 0.882 | 0.831 | 0.776 | 0.710 | ๐ 4.142 |
| Qwen3VL-4B | 0.933 | 0.857 | 0.790 | 0.719 | 0.644 | 3.943 |
| Qwen3VL-8B | 0.940 | 0.868 | 0.797 | 0.746 | 0.684 | 4.035 |
| Qwen3VL-30B-A3B | 0.939 | 0.877 | 0.820 | 0.757 | 0.682 | 4.075 |
Headline: Qwen3VL-2B (smallest Qwen3) is the overall winner at 4.142, beating the 31B MoE (4.075), 8B (4.035), and 7B (4.057). Paligemma-1 + VLM4VLA harness (3.506) matches ฯ0's full flow-matching pipeline (3.509) โ strong sign that architectural sophistication matters less than backbone choice.
| Model | Carrot | Eggplant | Spoon | Cube | Simpler โ | Libero โ |
|---|---|---|---|---|---|---|
| OpenVLA | 4.2 | 0.0 | 0.0 | 12.5 | 4.2 | 53.7 |
| ฯ0* | 62.5 | 100.0 | 54.2 | 25.0 | 60.4 | 46.0 |
| ThinkAct (uses state) | 37.5 | 70.8 | 58.3 | 8.7 | 43.8 | ๐ 70.9 |
| Qwen2.5VL-3B | 20.8 | 91.7 | 79.2 | 0.0 | 48.0 | 43.0 |
| Qwen2.5VL-7B | 12.5 | 100.0 | 75.0 | 0.0 | 46.8 | 45.0 |
| Qwen3VL-2B | 20.8 | 95.8 | 79.2 | 0.0 | 49.0 | 55.8 |
| Qwen3VL-4B | 54.2 | 95.8 | 75.0 | 0.0 | 56.3 | 44.4 |
| Qwen3VL-8B | 58.3 | 95.8 | 79.2 | 0.0 | 58.3 | 46.2 |
| Qwen3VL-30B-A3B | 29.2 | 79.2 | 70.8 | 0.0 | 44.8 | 46.8 |
| Paligemma-1 | 50.0 | 91.7 | 75.0 | 4.2 | 55.3 | 44.2 |
| Paligemma-2 | 75.0 | 75.0 | 79.2 | 0.0 | 57.3 | 46.2 |
| Kosmos-2 (1.7B) | 37.5 | 100.0 | 75.0 | 29.2 | ๐ 60.4 | 55.0 |
Headline: Kosmos-2 (the smallest model at 1.7B) wins SimplerEnv at 60.4%, tied with ฯ0. Qwen3VL-30B-A3B (31B, MoE) lags at 44.8%. Libero-Long is won by ThinkAct (70.9%), but ThinkAct cheats by using proprioceptive state; without state, Qwen3VL-2B leads at 55.8%.
Linear regression across all 9 VLMs:
| Benchmark | Pearson r | Rยฒ | Interpretation |
|---|---|---|---|
| Calvin ABC-D | +0.839 | 0.703 | High correlation (scene variation โ QA-like) |
| SimplerEnv Bridge | โ0.358 | 0.128 | No correlation, slightly negative |
| Libero-Long | โ0.194 | 0.038 | No correlation |
This is the core finding distilled into three numbers: strong VQA scores do not predict strong VLA performance on the benchmarks that measure actual manipulation.
| Model | Calvin (pretrained) | Calvin (from scratch) | ฮ | Simpler (pretrained) | Simpler (from scratch) | ฮ |
|---|---|---|---|---|---|---|
| Qwen2.5VL-3B | 3.856 | 1.381 | โ2.475 | 48.00 | 15.75 | โ32.25 |
| Qwen2.5VL-7B | 4.057 | 1.769 | โ2.288 | 46.75 | 18.20 | โ28.55 |
| Paligemma-1 | 3.506 | 1.129 | โ2.377 | 55.25 | 14.50 | โ40.75 |
Pretraining is necessary โ training from scratch collapses performance. So VLM pretraining is necessary; it's just not sufficient, and stronger pretraining isn't strictly better.
Freeze the vision encoder (or word embeddings) during VLA training; retrain; compare:
| Model | Full fine-tune | + Freeze vision | Vision ฮ | + Freeze word embeds | Embed ฮ |
|---|---|---|---|---|---|
| Qwen2.5VL-3B (Calvin) | 3.856 | 2.855 | โ1.001 | 3.849 | โ0.007 |
| Qwen2.5VL-3B (SimplerEnv) | 48.00 | 23.95 | โ24.05 | 46.88 | โ1.12 |
| Qwen2.5VL-7B (Calvin) | 4.057 | 2.823 | โ1.234 | 3.874 | โ0.183 |
| Qwen2.5VL-7B (SimplerEnv) | 46.75 | 25.50 | โ21.25 | 48.96 | +2.21 |
| Paligemma-1 (Calvin) | 3.506 | 0.495 | โ3.011 | 3.485 | โ0.021 |
| Paligemma-1 (SimplerEnv) | 55.25 | 13.25 | โ42.00 | 52.25 | โ3.00 |
Interpretation: Freezing the vision encoder is catastrophic (21โ42 points lost on SimplerEnv). Freezing word embeddings is essentially free (ยฑ0.2 points). The vision encoder is the primary bottleneck โ and fine-tuning the vision encoder matters more than just increasing the LLM's trainable parameters (frozen-vision 7.6B-tunable Qwen2.5VL-7B loses to fully-tuned 3.8B Qwen2.5VL-3B).
Fine-tune each VLM on a supposedly-helpful embodied auxiliary task before running the VLM4VLA harness; measure Calvin delta:
| Auxiliary task | Data | ฮ on Calvin |
|---|---|---|
| Robopoint (2D pointing) | 1.43M samples | โ0.073 |
| Vica-332k (spatial QA: size/distance/depth) | 332k | โ0.009 |
| BridgeVQA (spatial QA, real robot) | โ | โ0.091 |
| Robo2VLM (action-oriented VQA) | 667k VQA / 176k traj | โ0.096 |
| RoboBrain2 (embodied VQA at scale, official FT) | โ | โ0.170 |
| Omni-Generation (image+depth+seg gen + VQA) | โ | โ0.181 |
| VQA-Mix (general + embodied) | โ | โ0.079 |
Every auxiliary task hurts. Best case (VQA-Mix) is โ0.079. Worst (Omni-Generation) is โ0.181. Most show increased variance on top of the degradation. The intuitive "embodied pretraining helps VLA" assumption fails empirically.
Fine-tune Qwen3VL-4B's vision encoder on real-world BridgeV2 images with action information encoded as FAST tokens, then run VLM4VLA:
| Condition | Freeze vision during VLA | Unfreeze vision during VLA |
|---|---|---|
| Baseline (no VLM fine-tune) | 27.6 | 56.3 |
| VLM FT, freeze vision during FT | 28.0 (+0.4) | 56.3 (+0.0) |
| VLM FT, unfreeze vision during FT | 45.7 (+18.1) | 59.4 (+3.1) |
Interpretation: Fine-tuning VLM without touching the vision encoder doesn't help. Fine-tuning that does touch the vision encoder gains +18.1 points when VLA-time vision is frozen and +3.1 points when VLA-time vision is unfrozen. The gap between "VLM" and "VLA" is not about visual domain (sim vs. real) โ it's about which visual features are aligned with low-level control.
Raising input resolution from 224 โ 512 โ 768 doesn't improve Calvin scores when vision is unfrozen, and widens the frozen-vision gap (โ1.00 โ โ1.12 โ โ1.31 for Qwen2.5VL-3B). Suggests frozen high-resolution vision leaks more spurious correlations.
- No real-world robot experiments. "A limitation of our work is the absence of experiments on physical robots." They argue fairness+reproducibility on shared sim benchmarks outweighs this, and note their Table-4 experiments with real-world Bridge images show that sim-vs-real visual gap is not the primary bottleneck.
- Simulation-only. This prevents other researchers from directly comparing physical-robot VLAs to VLM4VLA findings.
- VLM coverage is limited to 1Bโ10B parameter open-source models plus one MoE (Qwen3VL-30B-A3B). No proprietary frontier models, no larger dense models.
- Auxiliary task sampling is not exhaustive. 7 tasks tested, findings may be dataset-specific.
- Unified hyperparameters. Required for fairness, but may not be individually optimal for each architecture.
- Regression loss + no diffusion simplifies evaluation variance but also handicaps every VLM equally against ฯ0 (which has a flow-matching expert). The takeaway "Kosmos-2 matches ฯ0" is harness-conditional โ with diffusion/flow heads the ranking may shift.
- Deterministic MSE + BCE loss is probably underfit for multimodal action distributions. The paper implicitly argues the ranking is still informative, which is plausible but not proven.
- Calvin benefits QwenVL, SimplerEnv benefits Paligemma/Kosmos-2 โ implicit evidence that benchmarks themselves select for different VLM capabilities. The paper frames this as a negative finding; a more positive reading is that different VLMs are specialists for different embodied skill sets, and the real research question is how to identify which VLM suits which benchmark.
- No vision-encoder architecture comparison. Kosmos-2 uses CLIP; Paligemma uses SigLIP; Qwen uses a custom ViT. The paper ablates tuning but doesn't compare vision encoder families directly.
- No study of how many action-conditioned vision-encoder tokens are needed to close the gap (the +18.1 result uses FAST tokens but the amount is not quantified cleanly in the body).
- The paper's diagnostic is strong; the prescription is weak. It correctly identifies the bottleneck but only gestures at solutions (action-aware VLM pretraining). Nobody has delivered a backbone trained end-to-end for VLA use that outperforms repurposed VQA VLMs.
- The VLM-selection heuristic is broken. "Pick the biggest/best-MMMU model" does not optimize VLA performance. Actually benchmark each candidate on your manipulation task.
- Vision encoder, not language, is the lever. Freezing word embeddings is nearly free; freezing vision is catastrophic. If you're cost-cutting, freeze the LLM, not the ViT.
- Embodied VQA pretraining is a red herring. Robopoint, Vica, BridgeVQA, Robo2VLM, RoboBrain2, Omni-Generation โ every auxiliary task tested hurts downstream VLA performance. Training a VLM on "robot-adjacent" data without touching the vision encoder with action signal is counterproductive.
- Action-supervised vision-encoder fine-tuning is the missing recipe. +18.1 points is the largest positive effect in the paper, and the only intervention that consistently helps.
- Counterintuitive best-in-class: Kosmos-2 (1.7B, 2023-vintage) matches the ฯ0 production system on SimplerEnv. Something about its visual representation is better-aligned with low-level control than strong modern VLMs. The paper doesn't explain why; this is an open research question.
- What does a VLM trained for VLA look like? The paper argues current VLM pretraining objectives are misaligned with action grounding, but doesn't propose a concrete alternative objective.
- Is there a cheap proxy benchmark that actually predicts downstream VLA performance? VQA scores don't; the paper doesn't suggest a replacement.
- Does the result hold under diffusion / flow-matching heads? The MSE harness isolates VLM effects but may not translate directly to production flow-matching VLAs.
- Does the result hold at frontier scale (100B+ VLMs)? Only sub-31B tested.
- Why does Kosmos-2 punch above its weight? Possible hypotheses: it was explicitly trained for grounding (bounding-box tokens), it has denser visual tokens, its architecture is smaller and less distractible. Needs follow-up work.
- Qwen Team's VLA Program โ cross-paper review reading this paper as the diagnostic foundation of the Qwen VLA line (Qiuyue Wang / Mingsheng Li / Shuai Bai later authored Qwen-VLA)
- VLM4VLA summary โ short version of this page
- ManipBench โ the earlier CoRL 2025 paper arguing the same thing via a VLM benchmark on low-level manipulation reasoning
- UniVLA โ action-as-first-class-modality direction, a possible answer
- ฯ0.6 โ uses Gemma3-4B (not in VLM4VLA's comparison set); informative reference point for production VLAs
- Survey: VLA & Manipulation (ICLR 2026) โ where VLM4VLA is flagged as the "surprising finding" (Trend 7)
- arXiv: https://arxiv.org/abs/2601.03309
- arXiv HTML: https://arxiv.org/html/2601.03309v1
- OpenReview: https://openreview.net/forum?id=tc2UsBeODW
- ICLR 2026 poster: https://iclr.cc/virtual/2026/poster/10006964
โ Back to ICLR-2026-VLM4VLA ยท ICLR-2026 ยท Home