Review Qwen VLA - Heungwoo/research GitHub Wiki

In-Depth Review — Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments

Paper: Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments Authors: Qiuyue Wang*, Mingsheng Li*, Jian Guan*, Jinhui Ye, Sicheng Xie, Yitao Liu, Junhao Chen, Zhixuan Liang, Jie Zhang, Xintong Hu, Xuhong Huang, Pei Lin, Junyang Lin, Dayiheng Liu, Shuai Bai† (corresponding), Jingren Zhou, + 24 contributors Affiliation: Qwen Team, Alibaba Group (Tongyi Lab) arXiv: 2605.30280 · v1 May 28, 2026 · v2 Jun 1, 2026 (34 pages, cs.RO) — this review is verified against v2; the substantive v1→v2 additions are the no-T2A baseline (60.9%, making T2A worth +10.2 pp) and the T2A action-representation clarification in §4.2 Code: github.com/QwenLM/Qwen-VLA · Blog: qwen.ai/blog?id=qwenvla Weights / License: model weights and license terms not stated in the paper as of this review

This is the long-form companion to the per-paper summary. Companion reviews to read alongside: π series evolution · GR00T series · LBM Co-training · Knowledge Insulation · VLA Architectures · VLM↔Action Connection.


1. TL;DR

  1. The Qwen team's first dedicated VLA. A Qwen3.5-4B vision-language backbone with early multimodal fusion is paired with a separate 1.15B DiT flow-matching action expert. Architecturally this is the same family as π0.6 / π0.7 (Category B in Review-VLA-Architecture) — not a same-stack MoE or a cross-attended dual-system head like GR00T.
  2. One generalist, many embodiments. A single set of weights handles WidowX, Google Robot, Franka Panda, ARX5, Fourier GR-1, Mobile ALOHA, AgiBot A2-D, Galaxea R1, AIRBOT MMK2, TienKung, and human MANO hands. The platform is selected by embodiment-aware prompt conditioning (plain text describing arms / waist / mobile base / control frequency / chunk size) — no per-embodiment heads, no architectural switches.
  3. A four-stage training recipe built around a "compression" thesis: (I) T2A — text-only DiT pretraining with images suppressed, (II) CPT — joint multimodal continued pretraining, (III) SFT in two parallel branches (multi-task / real-robot), (IV) PPO RL in SimplerEnv with sparse binary rewards and an analytic flow-matching log-probability via an ODE→SDE conversion.
  4. Headline numbers as a single generalist: 97.9% LIBERO, 73.7% Simpler-WidowX, 86.1 / 87.2% RoboTwin-Easy / Hard, 56.7% RoboCasa-GR1, 69.0% OSR / 57.5% SR on R2R, 59.6% SR on RxR, 76.9% average OOD success on ALOHA real-robot, 26.6% zero-shot SR on DOMINO dynamic manipulation. The DOMINO number beats every published fine-tuned baseline including PUMA (17.2%) without any DOMINO data.
  5. Recipe verdicts that align with TRI's LBM study. VL co-training helps fine-grained-recognition benchmarks (+4.9 pp RoboCasa-GR1, +4.6 pp RoboTwin) and never interferes; an ablation shows a pretrained DiT outperforms a from-scratch DiT throughout SFT. Explicit proprioceptive state in either the VLM prompt or the DiT yields at most +1.3 pp — the embodiment text prompt is sufficient.

2. Why this paper matters in the 2026 landscape

By Q2 2026 the VLA field has three publicly comparable "VLM-developer-as-VLA-author" lineages:

Lineage VLM developer Their VLA Architecture family
Google DeepMind Gemini Gemini Robotics 1.5 (Sep 2025) Closed dual-system
AllenAI Molmo MolmoAct (2508.07917) Reasoning-augmented (Cat G)
Alibaba Qwen Qwen3-VL / Qwen3.5 Qwen-VLA (May 2026, this paper) Flow-matching expert (Cat B)

Until May 2026 the Qwen3-VL backbone had been a building block for other people's VLAs — GR00T N1.7 uses Qwen3-VL-2B as Cosmos-Reason2-2B; NORA uses Qwen2.5-VL-3B; the StarVLA family uses Qwen2.5/3-VL across multiple action-head variants. Qwen-VLA is the first time the Qwen team itself ships a VLA, on a 4B Qwen3.5 (not Qwen3-VL) backbone with native multimodal early-fusion. This matters for two practical reasons:

  • Backbone identity is no longer ambiguous. Reading the paper, the Qwen team picks the Qwen3.5 family — a hybrid of gated linear attention and grouped-query softmax attention (Bai et al., Qwen3-VL technical report, arXiv 2511.21631) — rather than Qwen3-VL. The 4B size matches what PI and TRI have settled on for their LBM-class models (Gemma3-4B for π0.6/π0.7 and PaliGemma2-3B for LBM). The convergence is striking: three independent teams, three different VLM lineages, similar ~4B size, all paired with similar-sized flow-matching action experts.
  • Open-source posture is at least partially open. Code is published at github.com/QwenLM/Qwen-VLA. The paper itself does not state the weight license, and as of this review the repository content is documentation rather than full weights, so the practical openness story is still developing. This puts Qwen-VLA somewhere between fully-closed (π series, Gemini Robotics) and fully-open (GR00T N1.7 under Apache 2.0).

The other reason it matters is the DOMINO 26.6% zero-shot SR. DOMINO is a 2026 dynamic-manipulation benchmark (Fang et al., 2026) that other models fine-tune on — and Qwen-VLA-Instruct, evaluated zero-shot with current-frame observations only, surpasses the best fine-tuned baseline PUMA (17.2%) by 9.4 percentage points. That single number is the most defensible "the joint-pretraining bet paid off" claim in the paper.


3. Architecture

flowchart TB
  subgraph PROMPT["Embodiment-aware prompt + task instruction"]
    P1["The robot is {robot_tag} with {single/dual arms}[, waist][, and mobile base]."]
    P2["The control frequency is {FPS} Hz."]
    P3["Please predict the next {chunk_size} control actions to execute the following task: {ori_instruction}."]
  end

  subgraph IMG["Multi-view observations with view-tag tokens"]
    V1["ego camera"]
    V2["cam_left_wrist"]
    V3["cam_right_wrist"]
    VTAGS["wrapped as <|tag_start|> image <|tag_end|>"]
  end

  PROMPT --> VLM["Qwen3.5 (4B) — natively multimodal<br/>ViT with spatial merging · interleaved visual + text tokens<br/>Gated linear attention (majority of layers)<br/>+ GQA softmax attention at intervals<br/>+ M-RoPE"]
  IMG --> VLM

  VLM -- hidden states --> CAT["Concatenate VLM hidden states with noisy action chunk"]
  NOISE["Noisy action Y_tau in R^HxK"] --> CAT
  CAT --> DIT["DiT-style flow-matching action expert · 1.15B<br/>16 blocks · joint self-attention<br/>AdaLN timestep · multi-section RoPE aligned with backbone"]

  TIME["timestep tau"] --> ADALN["AdaLN modulation"]
  ADALN --> DIT
  DIT --> VEL["velocity field v_theta"]
  VEL --> EULER["Few Euler integration steps<br/>tau = 1 to 0"]
  EULER --> ACT["Action chunk Y_0 in R^HxK<br/>(zero-padded, mask-aware)"]

  VLM -. next-token CE on auxiliary text .-> LMHEAD["LM head"]
  LMHEAD --> VLLOSS["L_vl"]
  VEL --> FMLOSS["L_act"]

  classDef vlm fill:#bbdefb,stroke:#1565c0,color:#000
  classDef act fill:#c8e6c9,stroke:#2e7d32,color:#000
  classDef inp fill:#fff9c4,stroke:#f57f17,color:#000
  class VLM vlm
  class DIT,ADALN,VEL,EULER,ACT act
  class PROMPT,IMG inp
Loading

3.1 Backbone

  • Qwen3.5-4B (the paper cites Team, 2026 — pointing at the still-in-progress Qwen3.5 line, not the Nov 2025 Qwen3-VL technical report).
  • Natively multimodal with early fusion. Visual tokens are produced by a ViT with spatial merging and interleaved directly into the text token stream. There is no separate vision encoder bolted on after the fact.
  • Hybrid attention. Majority of layers use gated linear attention (cf. Qwen3.5's Gated DeltaNet choice); at regular intervals layers use grouped-query (GQA) softmax attention for full global reasoning. This is the Qwen3.5 standard recipe, inherited unchanged.
  • Multi-section RoPE (M-RoPE) for the image, text, and action token sub-sequences.
  • No separate vision encoder named (no SigLIP / DINOv2 / Eagle citation) — the ViT is part of the Qwen3.5 stack.

3.2 Action expert

  • Single-stream DiT (Esser et al., 2024) with flow-matching objective (Lipman et al., 2023).
  • Parameter budget — exact breakdown from §2.2 of the paper:
Component Parameters
16 DiT blocks (70.8M each) 1.13B
Action projection MLPs (raw action dim ↔ DiT latent) 4.9M
VLM hidden states → DiT channel linear 3.9M
Timestep embedding 2.8M
Output AdaLN modulation 4.7M
Total action expert ~1.15B
  • Connection mechanism: the action expert concatenates VLM hidden states with the noisy action chunk into one sequence and processes them via joint self-attention with AdaLN-injected timestep conditioning. There is no cross-attention into the VLM; there is no shared transformer either. This sits between Cat B (π-style separate-expert-with-attention-routing) and Cat C (GR00T-style cross-attention-into-VLM-features), and the paper itself frames it as a decoupled design that "lets the action expert specialize in fine-grained action generation".
  • Inference: Few Euler integration steps from τ = 1 to τ = 0. The paper does not give a single canonical step count, but its DOMINO discussion describes "coherent action chunks" so the recipe matches the small-K (≈5) regime of π0.6/π0.7 rather than the K=50 of original π0.

3.3 Unified action and trajectory representation (the cross-embodiment trick)

The single design choice that lets one DiT handle 11+ embodiments is:

  • Fixed tensor interface Y ∈ R^(H×K). H is a fixed prediction horizon; K is a fixed channel dimension shared across all control modes.
  • Active channels and zero padding. A control mode uses c ≤ K channels. The c task-relevant values live in the leading c dims; the remaining K − c are zero-padded.
  • Per-channel binary mask M ∈ {0,1}^(H×K): M_{h,k} = 1 iff k < c and h < H_task. The mask is used to exclude padded entries from the gradient.
  • No embodiment-specific output head. One DiT, one set of weights, switched only by the embodiment prompt + dataset-specific quantile normalization.

This is the same family of "shared latent, per-embodiment projection" idea that GR00T implements with MultiEmbodimentActionEncoder + CategorySpecificMLP, but Qwen-VLA simplifies further: the §5.2.2 ablation shows that zero-padding with a single shared MLP performs within 1.2 pp of per-embodiment Multi-MLP or Concatenation projections on Bridge + Robocasa, so the team picks zero-padding as the default.

3.4 VLM↔Action wiring — taxonomic placement

In the Review-VLA-Architecture taxonomy this is Category B (Flow-matching action expert), but with a concatenation attention routing rather than the prefix-KV / same-stack-MoE routing of the π series. Specifically:

  • π0.6 / π0.7: same-stack MoE — action expert is a parallel branch within the same transformer attending to a prefix-KV cache from the VLM.
  • GR00T N1 → N1.7: cross-attention from a separate DiT into truncated VLM hidden states (Cat C).
  • TRI LBM: adaLN-conditioned 8-layer flow expert reading a single observation token aggregated from the last 4 VLM layers.
  • Qwen-VLA: concatenate VLM hidden states + noisy action tokens → joint self-attention in a 16-block DiT. No cross-attention into VLM; no shared backbone transformer; one-direction information flow (VLM features feed in, no action-gradient back-prop into VLM during T2A; backbone unfrozen during CPT).

This makes Qwen-VLA's wiring a third-way that is closest in spirit to the simpler "DiT receives VLM features as context tokens" pattern that DiT4DiT and several 2026 VAM-class systems use, applied to a fully VLA setting.

3.5 Inference details

The paper does not state a single canonical inference setting. From the experimental sections:

  • Action chunk H = 16 for all sim manipulation benchmarks (LIBERO, RoboCasa-GR1, Simpler-WidowX, RoboTwin 2.0).
  • Action chunk H = 8 waypoints for navigation (R2R, RxR).
  • RL stage uses H = 16 and 128 parallel environments.
  • Temperature: τ = 1.0 during PPO rollouts, τ = 0.6 at evaluation time, to sharpen the action distribution.
  • No published latency / ms-per-chunk number — a gap relative to π0.6 (63 ms / chunk on H100) and GR00T N1 (63.9 ms / 16-action chunk on L40).

4. Training methodology

4.1 The compression view of action learning

The paper's training recipe is built on an explicit thesis:

"A language instruction such as 'pick up the red cup' together with an embodiment prompt compactly encodes the task intent in a handful of tokens, yet the corresponding action trajectory may span hundreds of high-dimensional joint-position values. Bridging this dimensionality gap is a structured decompression problem."

The four-stage recipe is built so each stage closes one specific gap from the one before it:

flowchart LR
  S1["Stage I — T2A<br/>VLM FROZEN<br/>Train DiT only<br/>Text + embodiment prompt → action<br/>No images (deliberately)"]
  S2["Stage II — CPT<br/>VLM UNFROZEN<br/>Joint multimodal training<br/>Heterogeneous mixture (Table 1)<br/>Both sim + real-robot data"]
  S3a["Stage III-a — Multi-task SFT<br/>VLM + DiT unfrozen<br/>VQA + spatial grounding<br/>+ manipulation + navigation<br/>Embodiment- and task-balanced"]
  S3b["Stage III-b — Real-robot SFT<br/>In-house ALOHA teleop<br/>Tests CPT-to-hardware transfer"]
  S4["Stage IV — RL<br/>PPO + GAE<br/>Sparse binary rewards<br/>SimplerEnv rollouts only<br/>→ Qwen-VLA-Instruct"]
  S1 --> S2
  S2 --> S3a
  S2 --> S3b
  S3a --> S4
  classDef ph fill:#fff9c4,stroke:#f57f17,color:#000
  class S1,S2,S3a,S3b,S4 ph
Loading

4.2 Stage I — T2A (text-to-action) DiT pretraining

The most distinctive design choice. The VLM is frozen; only the DiT trains; images are deliberately suppressed. The decoder must reconstruct action distributions from language + embodiment text alone. This is the "decompression prior" — the DiT learns how language indexes regions of action space before any visual grounding is available.

The §5.2.1 ablations are unusually thorough and produce four specific recommendations:

  • T2A itself is worth +10.2 pp (added in v2): the no-T2A baseline scores 60.9%, vs 71.1% for the best T2A configuration — the most direct quantification of the warm-start's value.
  • Data composition. Pure synthetic gets 64.1% downstream SFT, pure real gets 51.0%; the optimum is ~20% synthetic + 80% real, reaching 71.1%. Real anchors the prior in plausible dynamics; synthetic broadens language–action coverage.
  • Full-sequence prediction beats chunk prediction. +4.9 pp at 10% synthetic data, +2.9 pp at 0% synthetic. The decoder needs to see trajectory-level coherence; chunks fragment it.
  • Sigmoid-Normal timestep distribution at T2A is +5.7 pp over Beta. Without visual conditioning, intermediate noise levels carry the most learning signal for the language-action prior. Beta is then used at CPT/SFT where rich VLM conditioning is available.
  • T2A converges fast: 2,000 steps is optimal. Performance plateaus through 10k steps (67.5% / 67.2%); at 40k steps overfitting kicks in (60.4%, vs. 71.1% at 2k).
  • Action representation is identical between T2A and downstream stages (clarified in v2): actions are delta end-effector displacements relative to the first frame of each chunk, with the embodiment prompt carrying the platform, normalisation convention, and prediction horizon — so the language-only prior transfers directly into the visually-conditioned stages.

The "compression" framing is original — no other major VLA paper organizes the warm-start this way. It is functionally related to the language-as-warm-start arguments in OpenVLA-OFT and the action-prior pretraining of GR00T-N1's pseudo-action latent stage, but it is more aggressive: Qwen-VLA forces the prior to be purely language-conditioned.

4.3 Stage II — Continued Pretraining (CPT)

  • Both VLM and DiT unfrozen.
  • Heterogeneous data mixture (Table 1 of the paper):
Source Proportion
Robot Manipulation Trajectories 74.2%
Navigation Trajectories 7.5%
Human Egocentric Trajectories 6.0%
Synthetic Simulation Trajectories 3.7%
General Vision-Language Data 3.4%
Spatial Grounding (2D) 2.5%
Autonomous Driving VQA 2.4%
Fine-Grained Embodied Action Caption 0.2%

This is the data-axis identity of Qwen-VLA: heavily robot-trajectory-weighted (74.2%), with VL co-training as a minority but explicitly-purposeful contributor.

4.4 Stage III — Supervised fine-tuning (two parallel branches)

  • Multi-task SFT. Joint fine-tune on VQA + spatial grounding + manipulation + navigation under embodiment-balanced + task-balanced sampling. SFT loss weights: 0.1 for vision-language next-token prediction, 1.0 for both manipulation and navigation action prediction. The 10× ratio gives gradient capacity to action while preserving language grounding.
  • Real-robot SFT. In-house ALOHA teleop. Tests whether CPT's cross-domain priors transfer to physical hardware.

4.5 Stage IV — Reinforcement learning

This is the second distinctive contribution after T2A. Key details:

  • Algorithm: PPO with GAE (γ = 0.99, λ = 0.95, ϵ = 0.2). Four optimization epochs per rollout batch.
  • Value head: A lightweight value head is attached directly to the VLM backbone. The head mean-pools all VLM hidden states and maps them to a scalar via a linear projection. Stop-gradient on the VLM hidden states before they enter the value head — value-function gradients do not back-prop through the pretrained backbone. The value head uses LR = 10⁻⁴, ≈20× the actor LR of 5×10⁻⁶, so it converges fast while the policy updates remain conservative.
  • Log-probability under flow matching. The headline RL contribution: the deterministic probability-flow ODE is converted into a corresponding SDE by injecting controlled noise at each Euler denoising step (Song et al., 2021). Each transition then becomes an explicit Gaussian whose log-probability can be computed analytically without numerical ODE integration. During rollout the intermediate denoising states are stored; at the PPO update the velocity field is re-evaluated under the current parameters and Gaussian log-probability is recomputed. By default a single denoising step is randomly sampled per rollout for the log-prob estimate, requiring only one extra DiT forward pass during recomputation.
  • Reward: sparse binary — R = 1 if the task goal is achieved at episode end, 0 otherwise. No learned reward model.
  • Rollouts: 128 parallel environments, 8 rollout epochs × 128 environment steps per iteration = 8,192 transition chunks per iteration (with chunk length H = 16).
  • Training distribution: RL rollouts collected exclusively in SimplerEnv.

This is a meaningful contribution. Flow-matching policies have notoriously awkward log-probabilities — the closest 2025–2026 publications are ReinFlow (learnable noise injection for exact likelihoods) and PI's RECAP (advantage-conditioned RL that avoids needing a log-prob). Qwen-VLA picks a third path: a per-step Gaussian-stochasticization that lets standard PPO clip the importance ratio without modifying the policy parameterization.

4.6 Loss formulation

The flow-matching action loss with the per-channel mask M (Eq. 1–2 of the paper):

$$\ell_k = \frac{\sum_{h=1}^{H} M_{h,k} \left[ v_\theta(Y_\tau, \tau \mid o_{1:t}, x, e, z) - (Y_1 - Y_0) \right]_{h,k}^2}{\sum_{h=1}^{H} M_{h,k}}, \quad \mathcal{L}_{act} = \mathbb{E}_{\tau, Y_0, Y_1}\left[\frac{1}{c}\sum_{k=0}^{c-1} \ell_k\right]$$

The two-level averaging (first across timesteps for each active channel, then across active channels) ensures each control dimension contributes equally regardless of how many channels a given embodiment uses, and padded positions are fully excluded.

The next-token vision-language loss (Eq. 3):

$$\mathcal{L}_{vl} = -\sum_i \log p_\theta(w_i \mid w_{&lt;i}, o_{1:t})$$

Joint loss (Eq. 4):

$$\mathcal{L} = \lambda_{act} \mathcal{L}_{act} + \lambda_{vl} \mathcal{L}_{vl}$$

The exact pretraining λ values are not stated. The SFT-stage weights are λ_vl = 0.1, λ_act = 1.0 (action ≫ VL by 10×).

4.7 Hyperparameters — what is and isn't stated

Detail What the paper says
Optimizer Not explicitly stated (AdamW assumed from context)
Peak LR RL stage: 5×10⁻⁶ (actor), 1×10⁻⁴ (value head). Pretraining/SFT LRs not stated as a single number; the paper says "cosine-decayed learning rates with separate group-wise schedules for VLM backbone and action decoder"
Total steps Not stated. T2A converges in 2k steps; SFT ablation curves go up to 80k steps
Batch size RL: 8,192 transition chunks per iteration. Pretraining/SFT BS not stated
Hardware Not stated
GPU-hours Not stated
Inference latency (ms/chunk) Not stated

This is the biggest reproducibility gap. The paper is unusually detailed on the RL stage and the T2A ablations but does not give a single canonical pretrain compute budget.

4.8 Pretraining data — concrete sources

Section 3.2 of the paper lists every source explicitly. Distilled:

  • Real-robot public datasets (74.2% mixture, including all robot-trajectory tiers): RobotSet, Galaxea, AgiBot World, RoboCOIN, RoboMIND V1/V2, RDT-1B, DROID, BridgeData V2, RH20T, RT-1, BC-Z — over 10,000 hours total.
  • In-house real-robot: over 1,000 hours, ~20% of the total mixture.
  • Simulation trajectories (3.7% of mixture, "over 8M synthetic" in total): InternData-A1 + GR00T-X-Embodiment-Sim (public sim) + the team's own RoboInf vision-conditioned synthesis pipeline (random-placement scene generation: 20 tabletop scenes × 10 init configurations = 200 base scene configs, 450 tasks, 300 trajectories each, plus randomization over 3K backgrounds and 1K table textures), which yields 359,848 full successful trajectories including subtask segments (§3.2.3). The "over 8M synthetic" headline figure is the total synthetic count and is dominated by the 7.2M language-only trajectories below, not by RoboInf vision-conditioned data.
  • Stage-I-specific language-only action data: six task templates (pick-and-place, push, pull, rotation with reposition, rotation toward viewpoint, swap) × six robots (Franka Panda, UR10e, UR5e, Kinova Gen3, TM12, xArm7) × ~200k trajectories each = ~7.2M trajectories, >14,000 hours of motion-planned simulated robot trajectories at 50 Hz, without any rendering or physics simulation.
  • Egocentric human (6.0%): Ego4D + EPIC-KITCHENS subsets processed by VITRA (atomic manipulation segments + fine-grained captions + 3D hand & camera trajectories); EgoDex (829h dexterous, Apple Vision Pro, 194 tasks); EgoVerse (1,300+ hours, 1,965 tasks, 240 scenes); Xperience (depth + hand + body mocap + hierarchical annotations). Hand articulation is compressed to 10 PCA eigengrasp coefficients per hand from the 45-DoF axis-angle MANO pose; per-hand wrist motion is 6 SE(3) parameters → 32 dims per time step for ego-human data.
  • Navigation (7.5%): instruction-following (4.3%) + object searching (2.3%) + target tracking (1.0%); ~2 FPS video sampling; 3-DoF mobile robot (translation + heading).
  • Auxiliary VL (8.5% combined): general VL (3.4%) + spatial grounding (2.5%) + autonomous driving VQA (2.4%) + fine-grained embodied action caption (0.2%; ~48k video-caption pairs annotated along 13 dimensions by a two-stage Qwen3.6-plus pipeline).

The autonomous-driving VQA subset is particularly notable — it includes LingoQA, DriveAction, MMAU, Impromptu-VLA, nuScenes-QA, nuScenes-MQA, MapLM, WaymoQA, CODA-LM, Talk2Car, DrivingVQA, DriveLM, W3DA, GRAID, Bench2Drive-VL, DriveGPT4, OmniDrive, Senna, NAVSIM-RecogDrive, NAVSIM-Traj. The motivation is "viewpoint-robust localization", "temporal scene understanding", "language-grounded localization", and "planning-aware reasoning" — i.e. AD data is being recycled as embodied-perception co-training, not because Qwen-VLA targets driving.

4.9 Embodiment-aware prompt — the cross-embodiment interface

Every training sample is prefixed with:

The robot is {robot_tag} with {single arm / dual arms}[, waist][, and mobile base]. The control frequency is {FPS} Hz. Please predict the next {chunk_size} control actions to execute the following task: {ori_instruction}.

Eleven representative embodiments are listed in Table 2 of the paper:

Robot Arms Action type
WidowX Single ΔEEF + G
Google Robot Single ΔEEF + G
Franka Panda Single / Dual ΔEEF + G; Abs Joint + G
ARX5 Dual ΔEEF + G
Fourier GR-1 Dual ΔEEF + G
Mobile ALOHA Dual ΔEEF + G; Abs Joint + G
AgiBot A2-D Dual Abs Joint + G; Abs Joint + DH
Galaxea R1 Dual Abs Joint + G
AIRBOT MMK2 Dual Abs Joint + DH
TienKung Dual Abs Joint + G; Abs Joint + DH
Real Human Dual ΔEEF (from MANO)

Action values are normalized per-dataset using 1st / 99th percentile quantile mapping to [−1, 1] (Eq. 5 of the paper). Each dataset preserves its native control convention; the embodiment prompt + quantile normalization carries the platform-specific semantics.


6. Comparative analysis

This section answers the central question — where does Qwen-VLA sit relative to the leading 2026 VLA recipes?

6.1 Backbone choice

Axis Qwen-VLA π0.6 / π0.7 TRI LBM GR00T N1.7
Backbone Qwen3.5-4B (natively multimodal, gated-linear + GQA hybrid, M-RoPE) Gemma3-4B (SigLIP 400M + Gemma3-4B LLM) PaliGemma2-3B (SigLIP + Gemma2 LLM) Qwen3-VL-2B / Cosmos-Reason2-2B
Backbone params ~4B ~4.4B (4B LLM + 400M SigLIP) ~3B 2.44B
Vision encoder ViT with spatial merging, integrated into backbone SigLIP 400M, frozen-then-unfrozen SigLIP, frozen Qwen3 native vision encoder (truncated at last layer)
Backbone status during action training Frozen at T2A, unfrozen at CPT (joint training) Unfrozen but gradient-insulated via Knowledge Insulation Frozen throughout action training Top-N LLM layers unfrozen (tune_top_llm_layers=4 for N1.6, inherited at N1.7)
Attention recipe Gated linear attention majority + GQA at intervals Full-attention transformer Full-attention transformer Full-attention transformer + vlln + vl_self_attention (N1.7)
Native multimodal early fusion Yes (image tokens interleaved in text stream) No (SigLIP → projection → LLM) No (PaliGemma2 standard fusion) No (Qwen3-VL native vision; not fully early-fused)
Cross-embodiment design ethos Single set of weights + embodiment text prompt Cross-embodiment training, no per-embodiment head Per-target-platform specialization in phase 3 MultiEmbodimentActionEncoder + CategorySpecificMLP

Two non-trivial observations:

  1. The Qwen team chose Qwen3.5 over Qwen3-VL. This is consequential: Qwen3.5 uses a different attention recipe (gated linear attention + GQA intervals) than Qwen3-VL (pure GQA + QK-Norm — see ML-Attention table). Whether this matters for embodied tasks is not ablated in the paper.
  2. Qwen-VLA is the only one of the four that explicitly froze the backbone at the warm-start stage (T2A) — π series freezes nothing (KI gradient-routes instead); LBM keeps PaliGemma2 frozen throughout; GR00T does partial top-N unfreeze. The Qwen-VLA recipe sits between "all gradients allowed everywhere" and "VLM is forever frozen".

6.2 VLM↔Action connection

Three competing patterns are now well-documented in the field; Qwen-VLA picks a fourth:

Pattern Exemplar Mechanism
Same-stack MoE + prefix-KV π0.6 / π0.7 Action expert is a parallel branch within the same transformer; attends to prefix-KV from VLM
Cross-attention into truncated VLM GR00T N1 → N1.7 Separate DiT cross-attends to VLM hidden states; vlln + vl_self_attention preprocessing in N1.7
AdaLN-conditioned on aggregated VLM features TRI LBM 8-layer ActionFT consumes a single observation token built from last 4 VLM layers, gated by adaLN
Concatenation + joint self-attention Qwen-VLA Noisy action chunk concatenated with VLM hidden states into one sequence; DiT processes them through joint self-attention with AdaLN timestep injection

The Qwen-VLA choice is closer to the DiT-style image diffusion lineage (Esser et al., 2024 — Stable Diffusion 3 / SD3 family) than to any 2025 VLA. The action expert is not a separate cross-attender; it is a transformer reading a concatenated [VLM_features | noisy_action] sequence. This is the simplest possible "feed in the VLM features and let attention sort it out" approach, and the §5.2.2 ablation (pretrained DiT outperforms from-scratch DiT throughout SFT) is the empirical justification.

The cost of this choice: every action-expert forward pass re-attends over the full VLM hidden state. Compared to π0.6/π0.7's prefix-KV cache, this is in principle a higher inference cost per chunk. The paper does not publish latency, so this is unverified.

6.3 Action head & objective

System Action head class Objective Tokenizer / discretization
OpenVLA / π0-FAST Cat A (AR tokens) Cross-entropy FAST DCT tokens (vocab 2048)
DDVLA / dVLA Cat D (discrete diffusion in VLM) Masked cross-entropy Discrete action vocabulary
π0.6 / π0.7 / LBM / GR00T Cat B (separate flow-matching expert) Flow matching MSE on continuous targets Continuous (no action tokenizer) — π series adds FAST as VLM CE target via KI
Qwen-VLA Cat B (separate flow-matching expert) Flow matching MSE + standard NTP on language Continuous (no action tokenizer); auxiliary CE only on language

Qwen-VLA is unambiguously Category B in the Review-VLA-Architecture taxonomy. There is no FAST head, no VQ-VAE head, no latent action token head, no auxiliary CE on discretized actions — only the flow-matching loss on continuous targets and the standard next-token CE on auxiliary language data. This aligns with TRI LBM's finding that discrete action tokens (FAST, VQ-VAE) provide no significant benefit at LBM scale and sometimes hurt — Qwen-VLA never tried them. Whether this is because the team independently arrived at the same conclusion or because they read TRI's Feb 2026 study is not addressed.

6.4 Pretraining data composition

System Robot teleop Cross-embodiment OXE Human ego-video Web VL Auxiliary Total scale
Qwen-VLA >1,000 h in-house + >10,000 h public (74.2% mix) Same OXE bucket — not separated ~3,000+ h (Ego4D, EPIC-KITCHENS via VITRA + 829 h EgoDex + 1,300 h EgoVerse + Xperience) at 6.0% 3.4% + spatial grounding 2.5% + AD VQA 2.4% + caption 0.2% = 8.5% "over 8M" synthetic total = 359,848 RoboInf vision-conditioned + 7.2M language-only (3.7% mix) >10,000 h robot + ~3,000+ h human + 8.5% VL
π0.6 / π0.7 PI internal (very large, not disclosed) OXE + DROID + community Egocentric human (introduced in π0.7) Web multimodal + detection + captioning RL rollouts + failures (π0.7) Not separately disclosed
TRI LBM 523 h TRI-Ramen (target) 1,150 h OXE-Ramen (12 robots, 924 tasks, 466k demos) 2,271 h (Ego4D + EgoDex + Sth-Sth V2 + Epic Kitchen + HoloAssist) 50M VL samples (RoboPoint + RefSpatial) GPT-5 robot captions + GPT-5 human captions 523 + 1,150 + 2,271 h + 50M VL
GR00T N1.7 Several thousand hours teleop (YAM + AGIBot Genie1 + Galaxea R1 Pro + Unitree G1 + BEHAVIOR sim + DROID) Cross-embodiment throughout, all stages 20,854 h EgoScale (Apple Vision Pro EgoDex 829 h + in-the-wild ~20,000 h) (less prominent) DexMimicGen 780k sim + DreamGen neural trajectories ~3,000+ h robot + ~20,854 h human

Where Qwen-VLA sits on the scale axis: comparable to π / GR00T on robot trajectories (>10,000 h), comparable to LBM on human-video (~3,000 h is between LBM's 2,271 h and N1.7's 20,854 h), and lower than LBM on VL data (LBM's 50M RoboPoint + RefSpatial is the entire mixture; Qwen-VLA's 8.5% combined VL share is closer to π series' "co-training as side channel" framing). The data thesis is robot-trajectory-heavy with VL as regularizer, which is closest to π series rather than to LBM.

6.5 Training methodology details

Axis Qwen-VLA π0.6 / π0.7 TRI LBM GR00T N1.6 / 1.7
Number of phases 4 (T2A → CPT → SFT-multi-task ∥ SFT-real-robot → RL) 1 (joint with KI) 3 (pure co-train → joint specialization → target specialization) 1 (joint cross-embodiment)
Backbone freezing Frozen at T2A, unfrozen at CPT Unfrozen, gradient-insulated (KI) Frozen throughout Top-N LLM layers unfrozen (N=4 reported in N1.6)
Action-head warm-start T2A: text-only DiT pretraining Joint from scratch Joint from scratch (PaliGemma frozen, ActionFT random) Joint from scratch
Auxiliary CE objective NTP on language only (no FAST / discrete action target) FAST tokens CE as VLM training signal (via KI) NTP on language (RoboPoint/RefSpatial QA, GPT-5 captions, human captions); FAST/VQ-VAE/latent-video heads ablated and not used in the final recipe Co-train with various VLM losses + FLARE (N1.5 only) world-model latent alignment auxiliary
VL co-training role "Stabilizes language grounding and prevents catastrophic forgetting" Co-training for VLM-side supervision; KI keeps stable "Largest single-modality contribution; restores MMBench/GQA scores robot-only training degrades" Web data via SigLIP pretrain; less explicit VL co-train at action time
Loss weights SFT: λ_vl = 0.1, λ_act = 1.0 Not disclosed; KI structurally separates the two w_CE = 0.02 for CE on language; rest is flow matching Not disclosed
RL stage Yes — PPO + GAE on flow-matching log-prob via ODE→SDE, SimplerEnv-only RECAP (advantage-conditioned, separate π*0.6 release) No (pure supervised) DreamGen-style synthetic data, no canonical RL stage in the public N1.7 recipe

The biggest methodological points of comparison:

  • Phase count. Qwen-VLA at four phases is the longest pipeline of the four (π is single-phase, LBM is three, GR00T is single-phase). The complexity buys two things: a language-conditioned action prior (T2A) and a task-success refinement (RL), neither of which the other three pursue together.
  • Backbone freezing strategy. Qwen-VLA is the only one of the four that uses a time-varying freezing policy (frozen → unfrozen), and the only one that does it both for the warm-start and for the RL value head (the value head uses stop-gradient on VLM hidden states throughout RL). It is closer in spirit to LBM than to π / GR00T on this axis.
  • No discrete action token target. Both Qwen-VLA and the LBM study's final recipe drop discrete action tokens; π series (KI) and π0-FAST use them. This is the cleanest 2026 alignment with LBM's empirical finding.

6.6 Cross-embodiment & data scaling philosophy

Lab Philosophy
PI Internal-data-first; OXE and community data are background; cross-embodiment via shared continuous-action space + heterogeneous training
TRI LBM OXE used in phase 1 and phase 2; drop OXE in phase 3 to specialize the target dual-Franka; cross-embodiment as warm-start, not as final spec
GR00T Cross-embodiment throughout all stages; MultiEmbodimentActionEncoder + CategorySpecificMLP per-embodiment heads; data pyramid spans 7 ego-video datasets, sim, real teleop, neural trajectories
Qwen-VLA Cross-embodiment throughout (11 embodiments + human MANO in CPT mixture, all kept through SFT, no platform-specialization phase). Embodiment-aware text prompt as the sole interface; zero-padding projection ablated to be on par with per-embodiment Multi-MLP

Qwen-VLA is the most extreme "single weights, all embodiments, embodiment by prompt" philosophy of the four. Where LBM treats cross-embodiment as a phase-1 supplement and drops it for the target, Qwen-VLA explicitly aims for a generalist policy that serves all 11 platforms simultaneously — and its Table 4 demonstrates this by evaluating a single Qwen-VLA-Instruct on LIBERO + RoboCasa-GR1 + Simpler-WidowX + RoboTwin-Easy + RoboTwin-Hard without per-benchmark adaptation and comparing it head-to-head with specialist baselines that were fine-tuned per-benchmark. The single generalist beats the specialists on RoboCasa-GR1 (56.7 vs π0.5's 37.0), Simpler-WidowX (73.7 vs StarVLA-OFT's 64.6), and RoboTwin-Easy / Hard (86.1 / 87.2 vs ABot-M0's 86.0 / 85.0).

6.7 Empirical positioning

The most direct head-to-head numbers from Table 4:

Benchmark π0 π0.5 StarVLA-OFT GR00T N1.6 ABot-M0 Being-H0.5 Qwen-VLA-Base Qwen-VLA-Instruct
LIBERO 94.4 97.6 96.6 97.2 98.6 97.6 90.8 97.9
RoboCasa-GR1 — 37.0 48.8 49.9 58.3 53.3 40.4 56.7
Simpler-WidowX — 46.9 64.6 63.2 — — 64.3 73.7
RoboTwin-Easy 65.9 82.7 50.4 47.6 86.0 — 64.3 86.1
RoboTwin-Hard 58.4 76.8 — — 85.0 — 66.4 87.2

Qwen-VLA-Instruct does not beat ABot-M0 on LIBERO (98.6 vs 97.9) but matches it on RoboTwin and surpasses every model published as of May 2026 on RoboCasa-GR1 and Simpler-WidowX as a single generalist. The Base → Instruct delta (+7.1 LIBERO, +16.3 RoboCasa, +9.4 Simpler, +21.8 RoboTwin-Easy, +20.8 RoboTwin-Hard) is the empirical evidence that "instruction tuning yields substantial gains" on top of large-scale pretraining.

For real-world ALOHA (Tables 5–6 of the paper), the Qwen-VLA-aloha-with-pretrain variant outperforms π0.5 (Black et al. 2025) on the in-domain six-task average (83.6% vs 71.6%) and on every one of the five OOD axes, with the most striking gap being instruction generalization (84.6% vs π0.5's 42.3%) and background generalization (80.8% vs 26.9%). GR00T N1.6 trails far behind (28.6% in-domain, 25.4% OOD average), though it should be noted GR00T was not specifically fine-tuned for the ALOHA tasks — the comparison is "single model evaluation across the same suite", not a per-platform fairness adjustment.

Two additional empirical points worth flagging:

  • SimplerEnv-OOD (Table 8). Qwen-VLA-Instruct 32.0% > π0.5 12.6% on average across 6 OOD task types where fine-tuning was restricted to simple Bridge pick-and-place. π0.5 fails completely on MoveRight (0.0%) and PlaceNear (0.0%) — i.e. its training distribution excludes the spatial-instruction parsing these tasks require — while Qwen-VLA-Instruct scores 33.3% and 39.6% respectively. The structural reading: Qwen-VLA's heavier VL co-training (8.5%) and spatial-grounding share (2.5%) carry spatial-language priors that translate to spatial-instruction generalization.
  • DOMINO zero-shot (Table 9). Qwen-VLA-Instruct's 26.6% SR / 39.5 MS beats every fine-tuned baseline on a benchmark where moving-object dynamics are not part of the Qwen-VLA training data. The strongest fine-tuned baseline (PUMA, 17.2% SR / 35.0 MS) uses DOMINO-specific fine-tuning and temporal motion inputs; Qwen-VLA-Instruct uses neither and surpasses it by 9.4 pp SR. This is the most defensible claim that the joint pretraining produces generalizable spatial-to-kinematic priors, not just multi-task amortization.

6.8 Contributions analysis — novel vs. recombination

The user-specific ask. In plain language:

Genuinely novel:

  1. The text-to-action (T2A) decompression warm-start. Freezing the VLM and pretraining the DiT on language + embodiment text alone — with images deliberately suppressed — is a new training-stage idea. The closest prior art is OpenVLA's initial fine-tune on FAST tokens or GR00T-N1's LAPA pseudo-action stage, but neither prohibits visual input. T2A's explicit thesis (the decoder should learn a structured language-indexed action prior before visual grounding) is original and the §5.2.1 ablations on Sigmoid-Normal vs. Beta timestep, full-sequence vs. chunk, and the rapid 2k-step convergence are unique contributions.
  2. Analytic flow-matching log-probability via ODE→SDE conversion for PPO. The standard trick of injecting per-step Gaussian noise during Euler integration exists in the literature (Song et al., 2021) but applying it to flow-matching action policies for PPO clipping — with a per-step random sub-sampling so only one extra DiT forward pass is needed during importance-ratio recomputation — is a clean RL contribution. This is a genuinely different solution to the "flow-matching policies don't have tractable log-probs" problem than ReinFlow's learnable-noise-injection or PI's RECAP (advantage-conditioned, avoids log-probs entirely).
  3. The RoboInf vision-synthesis pipeline (359,848 trajectories) + 7.2M-trajectory rendering-free language-only synthetic stage (the "over 8M synthetic" total is dominated by the language-only set, not RoboInf). Both pipelines are described in detail. The rendering-free language-only pipeline (six task templates × six robots × ~200k trajectories each at 50 Hz, via cuRobo motion planning, without physics simulation or scene rendering) is a deliberately cheap T2A-targeted data factory and a novel scale of language-action-only supervision.
  4. The 11-embodiment + human-MANO single-policy demonstration. No prior paper has shipped a single VLA evaluated head-to-head against 7+ specialist baselines on 4 distinct manipulation benchmarks plus navigation on R2R/RxR. GR00T N1.7 covers more humanoid embodiments but does not publish the corresponding sim-benchmark head-to-head; π0.6/π0.7 cover fewer platforms.

Recombination of known recipes:

  1. VL co-training to prevent VLM forgetting. This is the TRI LBM finding (Review-LBM-Cotraining) — robot-only training erodes MMBench/GQA; balanced VL co-training restores them. Qwen-VLA's §5.2.2 confirms VL+VLA matches VLA-only on LIBERO/Simpler-WidowX and beats it on RoboCasa-GR1 (+4.9 pp) and RoboTwin-2.0 (+4.6 pp). This is a known recipe applied; the contribution is the validation at Qwen-3.5 scale.
  2. Continuous flow-matching action expert. The architectural family is π0 / π0.6 / GR00T / LBM. The specific size (1.15B), block count (16), and AdaLN timestep injection are conventional choices.
  3. Embodiment-aware prompt conditioning. Specifying the platform via plain text is the GR00T series / π0.7 approach. The Qwen-VLA prompt template (robot tag + arm count + waist + mobile base + FPS + chunk size) is a particular instance, not a new idea.
  4. PPO + GAE + sparse binary reward for VLA RL. Standard recipe; the only novel piece is the flow-matching log-prob (see above).

Where Qwen-VLA aligns with PI's bets:

  • Flow matching as the action objective (not discrete diffusion).
  • Single-model cross-embodiment with text-based platform selection.
  • VL co-training as a side channel (low mixture weight, throughout training) rather than as a primary signal.
  • Distinct VL and action loss weights (Qwen-VLA's 0.1 / 1.0 echoes LBM's 0.02 / 1.0 in spirit — much smaller language weight to avoid VLM gradient domination).

Where Qwen-VLA aligns with TRI LBM's empirical findings:

  • No discrete action tokens. LBM found FAST / VQ-VAE / latent video tokens do not help and sometimes hurt at LBM scale; Qwen-VLA simply does not include them.
  • VL co-training helps and never interferes. LBM's largest single-modality gain finding is re-confirmed at smaller scale (8.5% combined VL share vs. LBM's much larger).
  • The general framing of "co-training as anti-forgetting" rather than purely as generalization gain. Qwen-VLA's §2.5 explicitly says "this objective stabilizes language grounding under heavy embodied co-training and prevents catastrophic forgetting".

Where Qwen-VLA disagrees with either lab:

  • vs. LBM on cross-embodiment phasing. LBM finds cross-embodiment data must be dropped in phase 3 to let the target platform specialize. Qwen-VLA keeps all 11 embodiments + human MANO through every stage including the SFT branches. The paper does not run the LBM-style "drop OXE in phase 3" ablation, so we cannot tell whether the LBM finding generalizes. But Qwen-VLA's Table 11 shows the final RL-refined model preserves performance on all benchmarks — so at minimum, keeping cross-embodiment in late stages does not collapse the target-task performance the way LBM warned it would on the dual-Franka platform.
  • vs. PI on backbone freezing. PI / KI never freezes the backbone — KI's gradient-routing trick lets the backbone train continuously with stable CE-only gradients. Qwen-VLA's T2A does freeze the backbone, then unfreezes at CPT. This is a third-way that the LBM/KI camps had not explicitly considered.
  • vs. GR00T on cross-attention vs. concatenation. GR00T's whole architectural identity is cross-attention from DiT into VLM hidden states with an evolving preprocessing pipeline (raw → vlln → vlln + vl_self_attention). Qwen-VLA picks concatenation + joint self-attention — no cross-attention, no preprocessing — and the §5.2.2 (b) ablation showing that a pretrained DiT outperforms a from-scratch one throughout SFT is the indirect evidence that the simpler routing suffices.

The Qwen-team-specific contribution:

The paper does not explicitly capitalize on Qwen3.5-specific features (gated linear attention, M-RoPE, QK-Norm) beyond inheriting them via the backbone. There is no ablation isolating "would this work as well with Qwen3-VL instead of Qwen3.5?" or "does the gated-linear-attention layer help vs. full GQA?". So the Qwen-developer-side advantage is latent in the backbone choice but is not the paper's headline contribution. The genuine team-specific contributions are (a) the data pipeline (RoboInf synthesis + language-only synthetic), (b) the 4-stage recipe with T2A warm-start, and (c) the analytic flow-matching log-prob for PPO.


7. Comprehensive benchmark results

7.1 Manipulation in simulation (Table 4)

Single generalist vs. specialists fine-tuned per-benchmark:

Method Type LIBERO RoboCasa-GR1 Simpler-WidowX RoboTwin-Easy RoboTwin-Hard
π0 Specialist 94.4 — — 65.9 58.4
StarVLA-OFT Specialist 96.6 48.8 64.6 50.4 —
GR00T N1.6 Specialist 97.2 49.9 63.2 47.6 —
π0.5 Specialist 97.6 37.0 46.9 82.7 76.8
ABot-M0 Specialist 98.6 58.3 — 86.0 85.0
Being-H0.5 Specialist 97.6 53.3 — — —
Qwen-VLA-Base Generalist 90.8 40.4 64.3 64.3 66.4
Qwen-VLA-Instruct Generalist 97.9 56.7 73.7 86.1 87.2

7.2 ALOHA real-world in-domain (Table 5)

Model Pick&Place Table Cleaning Bowl Stacking Bowl Pick&Place Towel Folding Fine-grained Avg.
GR00T N1.6 30.8 38.5 53.8 19.2 19.2 10.3 28.6
π0.5 73.1 84.6 88.5 69.2 80.8 33.3 71.6
Qwen-VLA-aloha (w/o pretrain) 30.8 53.8 61.5 64.1 50.0 30.8 48.5
Qwen-VLA-aloha (w/ pretrain) 96.2 92.3 98.7 87.2 65.4 61.5 83.6

7.3 ALOHA real-world OOD (Table 6)

Model Color Instance Position Background Instruction Avg.
GR00T N1.6 46.2 38.5 3.8 19.2 19.2 25.4
π0.5 57.7 61.5 19.2 26.9 42.3 41.5
Qwen-VLA-aloha (w/o pretrain) 42.3 30.8 34.6 30.8 42.3 36.2
Qwen-VLA-aloha (w/ pretrain) 88.5 76.9 53.8 80.8 84.6 76.9

The +35.4 pp delta vs. π0.5 and +40.7 pp delta vs. the from-scratch ablation are the cleanest evidence in the paper that the Qwen-VLA-Base pretraining contributes — not just the architecture.

7.4 Vision-and-language navigation (Table 7)

R2R Val-Unseen and RxR Val-Unseen:

Method R2R NE↓ R2R OS↑ R2R SR↑ R2R SPL↑ RxR NE↓ RxR SR↑ RxR SPL↑ RxR nDTW↑
NaVid 5.7 49.2 41.9 36.5 5.7 45.7 38.2 —
Uni-NaVid 5.6 53.3 47.0 42.7 6.2 48.7 40.9 —
NaVILA 5.2 62.5 54.0 49.0 6.8 49.3 44.0 58.8
StreamVLN 5.0 64.2 56.9 51.9 6.2 52.9 46.0 61.9
Qwen-VLA-Base 5.2 61.7 53.8 49.4 6.4 55.1 45.8 56.2
Qwen-VLA-Instruct 5.1 69.0 57.5 51.2 5.8 59.6 47.8 57.1

Qwen-VLA-Instruct leads on most metrics; SPL and nDTW on the two splits are within 1 pp of the StreamVLN best.

7.5 SimplerEnv-OOD (Table 8) — restricted-finetune positional generalization

All models fine-tuned only on the Bridge training split (simple pick-and-place with fixed object pairings); evaluation on unseen spatial-instruction tasks.

Method MoveAway MoveRight PlaceNear PlaceRight PutFront StackYellow Avg.
π0.5 26.1 0.0 0.0 32.1 13.0 4.2 12.6
Qwen-VLA-Base 31.3 31.6 16.7 47.1 6.3 18.8 25.3
Qwen-VLA-Instruct 43.8 33.3 39.6 47.9 4.2 22.9 32.0

7.6 DOMINO dynamic manipulation (Table 9)

Method SR↑ MS↑ Notes
OpenVLA (fine-tuned) 1.5 6.1 DOMINO fine-tune
RDT-1B (fine-tuned) 5.3 17.7 DOMINO fine-tune
π0 (fine-tuned) 8.2 24.0 DOMINO fine-tune
π0.5 (fine-tuned) 9.6 26.2 DOMINO fine-tune
InternVLA-M1 (fine-tuned) 5.4 27.6 DOMINO fine-tune
VLA-Adapter (fine-tuned) 4.4 24.3 DOMINO fine-tune
π0-FAST (fine-tuned) 3.5 20.9 DOMINO fine-tune
OpenVLA-OFT (fine-tuned) 9.1 24.1 DOMINO fine-tune
StarVLA-OFT (fine-tuned) 10.9 30.5 DOMINO fine-tune
PUMA (fine-tuned) 17.2 35.0 DOMINO fine-tune + temporal motion inputs
OpenVLA-OFT (zero-shot) 6.7 20.0 zero-shot
π0.5 (zero-shot) 7.5 20.4 zero-shot
LingBot-VLA w/ depth (zero-shot) 11.8 26.7 zero-shot
LingBot-VA (zero-shot) 24.1 36.1 zero-shot, WAM-style
Qwen-VLA-Base (zero-shot) 21.1 37.4 zero-shot, current frame only
Qwen-VLA-Instruct (zero-shot) 26.6 39.5 zero-shot, current frame only

Qwen-VLA-Instruct is the best zero-shot model and the best overall model — including beating every fine-tuned baseline. This is the headline empirical result of the paper.

7.7 Cumulative effect of post-training stages (Table 11)

Stage Simpler RoboCasa RoboTwin-E RoboTwin-H LIBERO Simpler-OOD DOMINO SR DOMINO MS
CPT (Base) 64.3 40.4 64.3 66.4 90.8 25.3 21.1 37.4
+ SFT 70.8 56.0 86.3 87.1 97.8 31.6 25.7 39.1
+ RL (Instruct) 73.7 56.7 86.1 87.2 97.9 32.0 26.6 39.5

RL gives a +2.9 pp boost on SimplerEnv (where the rollouts come from) and 0.1–0.9 pp positive transfer to every other benchmark including DOMINO — i.e., no catastrophic forgetting from RL, even on held-out task families. The paper attributes this to (a) CPT having already exposed the model to mixed sim+real distributions and (b) RL optimizing task-success which generalizes across visual domains.


8. Ablations

The paper publishes a substantial ablation set (§5.2.1–5.2.4). Summary:

8.1 T2A pretraining (§5.2.1)

Five design choices, all measured by SFT success rate on Simpler-WidowX:

Choice Best setting Effect of going off-best
T2A stage itself (baseline added in v2) With T2A Without T2A: 60.9% vs 71.1% (−10.2 pp)
Data composition 20% synthetic + 80% real Pure real: −20 pp; pure synthetic: −7 pp
Sequence prediction mode Full-sequence Chunk: −4.9 pp at 10% synthetic
Visual input during T2A None Adding images: −2.9 pp
Flow-matching timestep distribution Sigmoid-Normal at T2A, Beta at SFT All other combinations: −5.7 to −11.7 pp
T2A training duration 2,000 steps 40k steps: −10.7 pp (overfitting)

8.2 Vision-language co-training (§5.2.2)

Benchmark VLA-only VL + VLA Delta
LIBERO ≈ same ≈ same ~0
Simpler-WidowX ≈ same ≈ same ~0
RoboCasa-GR1 51.1 56.0 +4.9
RoboTwin-2.0 81.8 86.4 +4.6

VL co-training never hurts; helps on tasks requiring fine-grained object recognition and compositional instruction parsing. The pattern is consistent with TRI LBM's findings on object-recognition-heavy generalization.

8.3 Pretrained DiT vs from-scratch (§5.2.2 fig. 7b)

The pretrained DiT (after T2A) outperforms a from-scratch DiT throughout SFT training on RoboCasa-GR1 — converges faster in early stages, reaches a higher peak. This is the direct validation of the T2A stage's structural value.

8.4 Projection design for heterogeneous embodiments (§5.2.2 Table 10)

Design Bridge Robocasa
Single-embodiment baseline 62.8 / 53.4 —
Multi-MLP 63.3 52.1
Concatenation 63.0 52.8
Zero-Padding 63.0 53.2

All three within 1.2 pp; Zero-Padding chosen for parameter efficiency (2h × d_max vs. 2h × Σd for the others).

8.5 State conditioning (§5.2.4, Table 12)

Conditioning RoboTwin-Easy RoboTwin-Hard
No state 88.7 87.4
State in VLM prompt (256-bin discretized) 89.3 88.7
State in DiT (continuous-valued) 89.4 88.3

≤ +1.3 pp benefit from either injection; the multi-view visual observations + flow-matching predicting relative displacements together obviate explicit proprioception. Embodiment text prompt as the sole platform-specific interface is the final design choice. This is one of the cleaner "less is more" findings in 2026 VLA literature.

8.6 RL stage cumulative effect (§5.2.3, Table 11)

Already covered above. Key observation: RL transfers without catastrophic forgetting across all six in-distribution + OOD + dynamic benchmarks.


9. Limitations

9.1 Authors' stated limitations (§7)

  1. Embodied action data remains far smaller and less diverse than VL pretraining data. Robustness to long-tail objects, environments, embodiments, and contact-rich interactions is limited.
  2. Joint training introduces optimization trade-offs. Action-oriented training "can modestly regress some pure vision-language and navigation evaluations" — the paper acknowledges this is a balancing problem they did not fully solve.
  3. Evaluations are largely short-horizon and benchmark-driven. Long-duration, failure-prone real-world deployment is an open challenge.

9.2 Reviewer's concerns (not in the paper)

  1. No latency / inference-speed publication. π0.6 reports 63 ms / chunk on H100; GR00T N1 reports 63.9 ms / 16-action chunk on L40; LBM reports 0.146 s avg. Qwen-VLA gives no comparable number, which is significant given its concatenation-of-VLM-features-into-DiT-input pattern is a priori more expensive than prefix-KV or single-token-summary patterns.
  2. Hyperparameter reproducibility gap. Total pretraining steps, batch size, optimizer choice, total GPU-hours, peak LR for pretrain/SFT — none are stated as a single canonical number. The RL stage is well-specified; the supervised pretraining is under-documented.
  3. Weights and license status not addressed in the paper. The GitHub repo exists but the paper does not publicly commit to a license or weight release. Without weights, the public reproducibility of the 97.9% LIBERO, 73.7% Simpler-WidowX, and 26.6% DOMINO numbers depends on the team's future release decision.
  4. No head-to-head against π0.6/π0.7. The paper compares against π0.5 (April 2025) and GR00T N1.6 (Sep 2025), but not against the November 2025 π0.6 or April 2026 π0.7. Both are closed-weights so the omission is partly understandable, but the paper's claim that Qwen-VLA-Instruct is a leading 2026 VLA would land harder with at least the published π0.6 / π0.7 numbers.
  5. No discrete-diffusion VLA baseline. Discrete Diffusion VLA hits 96.3% on LIBERO; Qwen-VLA-Instruct hits 97.9%. The two architectural families (Cat B vs. Cat D) are not directly compared. A LIBERO head-to-head with a single shared eval protocol would settle a long-running debate.
  6. The "single generalist beats specialists" framing has caveats. Several "specialist" baselines (e.g., ABot-M0 on LIBERO 98.6 vs Qwen-VLA's 97.9) actually beat Qwen-VLA. The fairer claim is "matches specialists on the easiest benchmarks, beats them on the harder ones".
  7. VL co-training ablation is limited. §5.2.2 shows VL+VLA matches VLA-only on Libero/Simpler and helps on RoboCasa/RoboTwin, but does not run the LBM-style "what fraction of VL data?" sweep that would settle the 8.5% choice. It also does not measure VLM-benchmark preservation (MMBench/GQA) the way LBM does — so the "prevents catastrophic forgetting" claim in §2.5 is asserted, not measured.
  8. T2A ablations are run only on Simpler-WidowX. The 71.1% peak is a single-benchmark result. Whether the 20%-synthetic-80%-real, Sigmoid-Normal-then-Beta, full-sequence, 2k-step optimum is the optimum for all embodiments and benchmarks is not verified — it is presented as a global recipe.
  9. The 8.5% combined VL share is much lower than LBM's recipe. LBM finds VL co-training is the largest single-modality contribution. Qwen-VLA's 8.5% is consistent with π series posture, but the LBM finding suggests there may be unrealized gain available from increasing it. The paper does not run this sweep.
  10. No published failure-mode analysis. Real-world experiments are short-horizon (≤ Table 5's six categories). Long-horizon deployment with recovery, replanning, and observation-of-failure is acknowledged as future work, not measured.

10. Significance & positioning

10.1 Where Qwen-VLA sits in the 2026 landscape

The May 2026 release positions Qwen-VLA as:

  • First Qwen team VLA. Backbone-first VLM developers shipping their own VLA is a 2025–2026 trend (Google → Gemini Robotics; AllenAI → MolmoAct; Alibaba → Qwen-VLA). The technical contributions here are the four-stage recipe and the unified action representation; the strategic contribution is making Qwen a credible end-to-end VLA developer rather than just a backbone vendor.
  • The most-embodiment-diverse single-policy demonstration of 2026. 11 robot platforms + human MANO + navigation + autonomous driving VQA + spatial grounding all routed through one model. GR00T N1.7 covers more humanoid embodiments specifically; π0.7 covers fewer but with deeper post-training. No prior paper has demonstrated this breadth at this benchmark density.
  • The strongest "joint pretraining as transferable prior" claim of 2026. The 26.6% DOMINO zero-shot result — beating every fine-tuned baseline including PUMA — is the most defensible empirical result. It would have been even stronger had the paper published an ablation showing which pretraining sources matter for the DOMINO transfer.
  • Aligned with TRI LBM's empirical findings on three axes (discrete action tokens not needed; VL co-training helps and doesn't interfere; the action expert benefits from a structured warm-start).
  • A third-way between PI's KI and TRI LBM's frozen backbone for the VLM training policy: freeze for warm-start, unfreeze for joint training, stop-gradient for the RL value head.

10.2 Does it change the "which architecture should I pick?" calculus from Review-VLA-Architecture?

Not radically, but with three refinements:

  1. Category B (flow-matching expert) gets a new wiring pattern. Concatenation + joint self-attention is a viable third option alongside same-stack MoE (π) and adaLN-conditioning (LBM). Whether it has different latency / scaling properties is not yet measured publicly.
  2. The T2A warm-start is a new training-recipe primitive. It is generalizable in principle — any backbone-frozen, language-only DiT pretrain could be applied to any Cat B VLA. Whether it transfers to π0.6 / LBM-class systems remains to be seen.
  3. The ODE→SDE flow-matching log-prob is a new RL primitive. Applicable to any flow-matching VLA. This may be the most reusable contribution of the paper for the RL community.

Where Qwen-VLA does not change the calculus:

  • It does not displace π series on dexterous long-horizon real-world tasks (the laundry / box assembly tier — the paper's real-world tasks are shorter-horizon).
  • It does not displace GR00T on humanoid breadth (no Atlas, no Optimus, no Figure).
  • It does not address the discrete-diffusion-vs-flow-matching debate; it is firmly in the flow-matching camp.

10.3 Open questions Qwen-VLA leaves open

  • Does T2A scale beyond LBM/π/GR00T scale? The 2k-step convergence finding is on a small Simpler-WidowX eval. Whether the language-only prior helps at 100k-step pretraining is unverified.
  • Does the analytic flow-matching log-prob trick scale to long-horizon dexterous tasks beyond SimplerEnv? The RL stage was deliberately narrow ("a single simulation environment"); it has not yet been validated on a π*0.6-style laundry or box-assembly task.
  • Does the concatenation-routing scale to 50-step chunks (π-series scale) or beyond 16 DiT blocks?
  • Does increasing the VL co-training share beyond 8.5% — toward LBM's much larger share — help further?

11. Links


12. Related pages

← Back to Home

⚠️ **GitHub.com Fallback** ⚠️