Review Qwen Team VLA - Heungwoo/research GitHub Wiki

Cross-Paper Review — The Qwen Team's VLA Program: From VLM4VLA to the Qwen-Robot Suite

Scope: every VLA-relevant paper the Qwen team has authored or co-authored, read as one research program. Papers covered in depth: VLM4VLA (arXiv 2601.03309, ICLR 2026) · Qwen-VLA (arXiv 2605.30280) · Qwen-RobotManip (arXiv 2606.17846) · Qwen-RobotNav (arXiv 2606.18112) · Qwen-RobotWorld (arXiv 2606.17030) Caveat: the papers barely cite one another (the only visible connection is RobotWorld evaluating zero-shot on RobotManip's RoboTwin-IF benchmark). The "program" reconstructed here is an external reading from author lists, timing, and shared design choices — not the team's own stated roadmap.


1. TL;DR

  1. In six months (Jan–Jun 2026) the Qwen team went from diagnostic study to full embodied-AI stack. The arc: VLM4VLA (Jan, with Tsinghua) asks which VLM makes a good VLA and why; Qwen-VLA (May) ships the first flagship generalist; the Qwen-Robot Suite (Jun 16–17) productizes the stack into three specialized models — RobotManip (manipulation VLA), RobotNav (navigation VLN), RobotWorld (video world model).
  2. The program has a coherent core vision: a natively-multimodal Qwen backbone (though the family fragments: Qwen3.5-4B in the flagships, Qwen3-VL in RobotNav, frozen Qwen2.5-VL in RobotWorld) + flow-matching continuous action experts for manipulation + joint VL co-training against forgetting (λ = 0.1 in both flagships, 1.0 in RobotNav) + single cross-embodiment policy switched by text prompts + masked canonical action tensors + synthetic data as the scaling engine + deep skepticism of in-distribution benchmarks + language as the universal interface across policy, navigation, and world model.
  3. But the two flagship VLAs contradict each other on nearly every remaining axis: VLM↔expert wiring (concatenation vs cross-attention — and RobotManip's own ablation votes against Qwen-VLA's choice), action space (native conventions vs camera-frame delta EEF), proprioception (omitted vs first-class), recipe (4-stage + RL vs single-phase, no RL), data posture (in-house teleop vs open-data-only). Two teams, one org, five weeks apart, zero cross-citations.
  4. VLM4VLA reads as the program's empirical foundation — its two actionable findings (the vision encoder is the bottleneck and must receive action gradients; sequential embodied-VQA fine-tuning hurts, so supervision must be joint) are followed by both flagships. Its most inconvenient finding (VLM benchmark scores don't predict VLA performance — a 1.7B Kosmos-2 beat a 31B Qwen3VL MoE) did not stop the team from building both flagships on its own newest Qwen3.5-4B.
  5. The strategic break: closed weights. Despite Qwen's open-weights LLM tradition, no weights are released for any robot model, and the RobotManip/RobotNav README states there is no plan to. Qwen enters embodied AI as a full-stack closed provider — the same posture as Physical Intelligence and Gemini Robotics, opposite to its own LLM identity and to NVIDIA's Apache-2.0 GR00T.

2. The paper line

Date Paper arXiv What it is Venue / status
2026-01 VLM4VLA 2601.03309 Controlled study: 9 VLM backbones (24 variants) through a minimal <1%-new-params harness on Calvin / SimplerEnv / LIBERO-Long ICLR 2026 poster; Tsinghua-led, Qwen co-authors
2026-05-28 Qwen-VLA 2605.30280 Flagship generalist VLA: manipulation + navigation + AD-VQA, 11 embodiments + human MANO, 4-stage recipe (T2A → CPT → SFT → PPO RL) Technical report, 34 pp
2026-06-16/17 Qwen-RobotManip 2606.17846 Manipulation-specialist VLA: alignment-first thesis, camera-frame delta EEF, ~38,100 h open-data corpus, RoboChallenge #1 Technical report, 44 pp
2026-06-16/17 Qwen-RobotNav 2606.18112 Navigation specialist on Qwen3-VL (2B/4B/8B): parameterized observation interface (token budget · temporal decay · camera weights) with training-time randomization; 15.6M samples; SOTA on VLN-CE/ObjectNav/EVT-Bench/NAVSIM; agentic EQA SOTA with a Qwen3.6-Plus planner Technical report, in-depth review
2026-06-16/17 Qwen-RobotWorld 2606.17030 Language-actioned video world model: 20B 60-block double-stream MMDiT + frozen Qwen2.5-VL-7B action encoder + Wan-VAE; 8.6M-pair EWK corpus with five-layer action-language annotation; Scene2Robot H2R editing; 1st on EWMBench / DreamGen Bench, best open-source on WorldModelBench / PBench Technical report, in-depth review

The June trio was announced together as the Qwen-Robot Suite. Qwen-VLA (May) is conspicuously not part of the suite branding, and the suite's manipulation model is not an iteration of it — it is a different architecture from a different sub-team (§3).

Personnel threads (from author lists — verifiable)

  • VLM4VLA → Qwen-VLA: Qiuyue Wang and Mingsheng Li are VLM4VLA co-authors and become Qwen-VLA's first two (equal-contribution) authors; Shuai Bai and Junyang Lin (Qwen core) co-author VLM4VLA, and Shuai Bai is Qwen-VLA's corresponding author. VLM4VLA is led by Jianke Zhang / Jianyu Chen (Tsinghua). The diagnostic study directly seeded the first flagship.
  • Qwen-VLA ↔ Qwen-RobotManip: partial overlap (Zhixuan Liang, Pei Lin, Jie Zhang, Jinhui Ye, Sicheng Xie, Shuai Bai, Junyang Lin, Dayiheng Liu, Jingren Zhou) but disjoint leadership — Qwen-VLA answers to Shuai Bai (Qwen-VL lineage); RobotManip's corresponding/lead authors are Chenfei Wu and Xiong-Hui Chen, with a robotics/RL-heavy core-contributor list (Haoqi Yuan, Anzhe Chen, …) that does not appear on Qwen-VLA.
  • The most parsimonious organizational reading: two parallel embodied-AI efforts inside Qwen/Tongyi, one grown out of the VLM team, one staffed as a robotics team — with the June suite marking the robotics line as the productized track. This is interpretation, not documented fact.

3. The diagnostic: what VLM4VLA established

VLM4VLA is the program's empirical bedrock — a Tsinghua×Qwen controlled study whose findings define the design space the flagships later occupy:

  1. General VLM benchmark scores do not predict VLA performance (Pearson r = −0.36 on SimplerEnv, −0.19 on LIBERO-Long; only Calvin correlates at +0.84). Kosmos-2 at 1.7B ties π0 and beats Qwen3VL-30B-A3B (31B) on SimplerEnv.
  2. The vision encoder is the bottleneck, not language. Freezing the ViT during VLA training costs 21–42 points on SimplerEnv; freezing word embeddings costs ±0.2.
  3. Sequential fine-tuning on embodied auxiliary VQA (Robopoint, BridgeVQA, RoboBrain2, …) consistently hurts downstream VLA performance — all 7 auxiliary tasks tested produce negative Calvin deltas.
  4. Action-supervised vision-encoder fine-tuning is the one intervention that helps (+18.1 on SimplerEnv when the ViT receives action-token gradients).
  5. VLM pretraining is necessary (from-scratch collapses by 2.3–2.5 Calvin points) — just not sufficient, and not monotone in VLM quality.

Did the flagships follow their own diagnostic?

VLM4VLA finding Qwen-VLA (May) Qwen-RobotManip (Jun) Verdict
Vision encoder must receive action gradients Backbone (incl. ViT) frozen only during T2A warm-start, unfrozen from CPT onward Backbone fully unfrozen end-to-end; flow-matching gradients reach the ViT Followed — neither flagship ships a frozen vision encoder (contrast: TRI LBM freezes PaliGemma throughout)
Sequential embodied-VQA FT hurts No sequential VQA stage; VL supervision only as joint co-training (λ_vl = 0.1) Same — dual-stream joint co-training (9:1, λ = 0.1); +8.2 pp on RT-C2R Hard Followed, with a nuance: VLM4VLA tested sequential fine-tuning; both flagships use joint co-training, and RobotManip's ablation shows joint helps. The program implicitly claims the harm was in the sequencing, not the data
VQA score doesn't predict VLA quality — benchmark your candidates Adopts newest in-house Qwen3.5-4B without a backbone-selection ablation Same backbone, same absence of selection ablation Not followed. Neither flagship runs the study's own prescribed comparison; Qwen3.5 was not in VLM4VLA's test set. Pragmatics (in-house stack, early fusion) won over the diagnostic
Action-supervised vision FT is the missing recipe Realized implicitly via joint end-to-end training at scale Same, plus ECoT / 2D-trajectory VL data that inject action semantics into the VLM stream Followed in spirit — neither uses the paper's specific FAST-token vision-FT recipe, but both make action supervision reach the ViT
Minimal deterministic harness for clean comparison Flow-matching DiT (1.15B) Flow-matching DiT (small) N/A — the harness was diagnostic tooling, not a production recommendation

The scorecard: the two mechanistic findings shaped both flagships; the backbone-selection finding — the paper's headline — was overridden by the obvious institutional incentive to showcase Qwen3.5. Notably, VLM4VLA's closing open question ("what does a VLM trained for VLA look like?") is still unanswered by the program's own flagships: both consume a general-purpose Qwen3.5 rather than an action-aware backbone variant.


4. The shared vision — eight bets both flagships make

Where the two VLAs agree, they agree strikingly precisely. This is the closest thing to a documented "Qwen VLA doctrine":

  1. Qwen3.5-4B, natively multimodal, early fusion — for the flagships. Both manipulation flagships use the identical backbone — interleaved visual tokens from a dynamic-resolution ViT, no bolt-on encoder. The ~4B choice independently converges with PI (Gemma3-4B) and TRI (PaliGemma2-3B): three labs, three VLM lineages, one size class. But the suite fragments this: RobotNav runs on Qwen3-VL (2B/4B/8B) and RobotWorld uses a frozen Qwen2.5-VL-7B as action encoder — three backbone generations across one program, with no stated rationale for the split.
  2. Category B flow-matching continuous actions — no discrete action tokens anywhere. No FAST heads, no VQ-VAE, no latent-action vocabularies in either paper (nor in the suite's world model, which uses language as its action interface). The program never even ablates discrete tokens — aligning with LBM's negative result by omission.
  3. Joint VL co-training as the anti-forgetting mechanism — λ = 0.1 for manipulation, λ = 1.0 for navigation. Both flagships use joint VL supervision at exactly one-tenth weight of the action loss, and neither uses architectural insulation (KI) or freezing; the recurring 0.1 suggests a shared internal recipe. RobotNav keeps the mechanism (co-training against "collapse into reactive action-sequence mappers") but sets λ = 1.0 — a plausible domain difference (navigation's MSE waypoint loss is a weaker gradient source than flow matching), though no paper discusses it.
  4. One generalist policy, cross-embodiment by prompt + masked canonical tensor. Both encode the platform as text (Qwen-VLA's sentence template; RobotManip's structured fields) and both use zero-padded fixed action tensors with per-dimension loss masks. Neither uses per-embodiment heads (contrast GR00T's CategorySpecificMLP).
  5. Synthetic data as the scaling engine, not teleop fleets. Qwen-VLA: 7.2M rendering-free language-only trajectories + RoboInf vision synthesis. RobotManip: 24,808 h of human-to-robot re-rendering — 65% of its corpus. The program's data thesis is that synthesis + curation infrastructure substitutes for proprietary collection (RobotManip makes this explicit by using zero in-house teleop; Qwen-VLA still used >1,000 h in-house).
  6. Benchmark skepticism as identity. Qwen-VLA stakes its claim on zero-shot DOMINO transfer beating fine-tuned baselines; RobotManip escalates to a full manifesto (from-scratch matches pretrained in-distribution) and ships two new OOD benchmarks (RoboTwin-IF/XE). Both papers argue the field's standard evaluations reward memorization.
  7. Language as the universal interface of the whole stack. Embodiment prompts select the robot; ECoT reasoning is expressed in language; RobotWorld conditions video prediction on language as "a unified action interface"; RobotNav exposes a parameterized text interface for an agentic planner to switch task modes mid-episode. The suite's implicit System-1/System-2 story routes everything through natural language — the most Qwen-like possible bet.
  8. Closed weights. No model in the program has released weights; the suite README states no plan to. For a lab whose LLM brand is open weights, this is the clearest signal that embodied AI is being treated as a commercial moat, not an ecosystem play.

5. The internal contradictions — where the program disagrees with itself

Condensed from Review-Qwen-RobotManip §9.1 (full table there):

Axis Qwen-VLA Qwen-RobotManip Who's right?
VLM↔expert wiring Concatenation + joint self-attention (16-block, 1.15B DiT) Cross-attention into last-layer states, alternating vision/language by block parity (10-block, D=768 DiT) RobotManip's own Table 19 ablation: cross-attention 87.5 vs concatenation 87.0 on LIBERO-Plus — a direct, if narrow, vote against its sibling
Expert size 1.15B ~10× smaller Unresolved — no matched-scale comparison exists
Action space Per-dataset native conventions + quantile norm 80-dim canonical + camera-frame delta EEF (calibration required) RobotManip's scaling-law ablation is the strongest evidence in either paper: alignment creates the data scaling law
Proprioception Ablated to ≤+1.3 pp, omitted 80-dim state vector, first-class input Unreconciled — plausibly action-space-dependent (camera-frame deltas may need state anchoring)
Recipe 4 stages incl. T2A warm-start and PPO RL (ODE→SDE log-prob) Single-phase dual-stream, no warm-start, no RL Unresolved; RobotManip gets its OOD results without RL, questioning whether Qwen-VLA's stage-IV complexity earns its keep
Data posture In-house teleop + public + synthetic Open + synthesized only RobotManip's open-data-only result is the more disruptive claim if it holds
Scope Manip + nav + AD-VQA in one model Manip only; nav split into RobotNav The suite decomposition suggests the specialist view won internally

Two more observations:

  • The generalist-vs-suite fork now has a scoreboard. Qwen-VLA argues one model should do manipulation and navigation (it evaluates on R2R/RxR); three weeks later the suite ships navigation as a separate model — and RobotNav-8B beats Qwen-VLA-Instruct on the same VLN-CE Val-Unseen splits by +14.6 pp R2R SR (72.1 vs 57.5) and +16.9 pp RxR SR (76.5 vs 59.6) (with vastly more navigation training data, to be fair). On this axis the specialist decomposition is empirically vindicated, which reads as the Qwen-VLA generalist line being superseded.
  • Neither flagship evaluates the other, and neither compares against π0.6/π0.7. The program's internal contradictions are exactly the controlled experiments the field wants (concat vs cross-attn at matched scale; T2A+RL vs alignment-first; state vs no-state under both action spaces), and only this team can run them.

6. The suite as a stack — the productized vision

Read as a product architecture, the June suite decomposes embodied intelligence into three language-interfaced components:

flowchart LR
  subgraph AGENT["Agentic planner (Qwen3.6-Plus — demonstrated for navigation, not released)"]
    PLAN["Long-horizon goal decomposition<br/>task-mode + observation-config switching<br/>evidence-notebook memory"]
  end
  subgraph SUITE["Qwen-Robot Suite (Jun 2026)"]
    MANIP["Qwen-RobotManip<br/>manipulation VLA<br/>(flow-matching, camera-frame EEF)"]
    NAV["Qwen-RobotNav<br/>VLN engine<br/>(parameterized task interface)"]
    WORLD["Qwen-RobotWorld<br/>language-conditioned video WM<br/>(data synthesis · eval · planning signals)"]
  end
  PLAN -- "language sub-goals" --> MANIP
  PLAN -- "task mode + params" --> NAV
  WORLD -- "synthetic trajectories / rollout eval" --> MANIP
  WORLD -- "imagined futures" --> PLAN

  classDef m fill:#c8e6c9,stroke:#2e7d32,color:#000
  classDef n fill:#bbdefb,stroke:#1565c0,color:#000
  classDef w fill:#fff9c4,stroke:#f57f17,color:#000
  class MANIP m
  class NAV n
  class WORLD w
Loading
  • The decomposition maps cleanly onto the System 0/1/2 framing: RobotManip/RobotNav are System-1 executors, the planner is System-2, RobotWorld is the imagination module — with language as the bus between every layer.
  • The System-2 slot is demonstrated on the navigation side. RobotNav ships an agent-facing tool interface — nav_qwennav(L_i, τ_i, Φ_i) with runtime-switchable task modes and observation configs — and a working two-tier system (Qwen3.6-Plus planner + trajectory-evidence harness + persistent evidence notebook) that sets EQA SOTA (HM-EQA 76.7, EXPRESS-Bench 79.27) with 77% fewer navigation steps, validated on a real robot. No equivalent planner-over-RobotManip system exists yet.
  • RobotWorld's three stated application directions (synthetic data generation for policy training, virtual evaluation environments, language-guided planning signals) tie directly back to the two flagships' biggest gaps: RobotManip's synthesis-quality ceiling and the program's benchmark-skepticism (a world model as evaluator is the logical endpoint of "don't trust static benchmarks"). The one visible hand-off so far: RobotWorld evaluates itself zero-shot on RobotManip's RoboTwin-IF benchmark. But RobotWorld's action interface is language only — it cannot consume the continuous actions its sibling VLAs emit, so the policy-evaluation loop remains structurally open.
  • The suite's data engines are siblings: RobotManip's H2R synthesis (15 platforms, 24,808 h) and RobotWorld's Scene2Robot training pairs (MANO → 14 robot arms, MuJoCo IK + inpainting) are the same pipeline pattern re-tasked — one produces training trajectories, the other paired video-editing supervision.
  • Missing from the suite, notably: an RL layer (Qwen-VLA's stage IV has no successor here), a dexterous-hand story beyond the canonical vector's 12 hand dims, and any humanoid whole-body control — the suite is arms-and-wheels, not humanoids (contrast GR00T).

7. Positioning against the other 2026 programs

Axis Qwen program Physical Intelligence NVIDIA GR00T TRI LBM Google Gemini Robotics
Entry point VLM vendor moving down-stack Robotics-native startup Platform/ecosystem vendor Auto-industry research lab VLM vendor moving down-stack
Backbone Qwen3.5-4B (own) Gemma3-4B (external) Qwen3-VL-2B (external — Qwen's!) PaliGemma2-3B (external) Gemini (own)
Data thesis Open + synthesis (RobotManip: zero proprietary) In-house fleet-first Data pyramid: ego-video + sim + teleop Modest curated teleop + co-training science Undisclosed, fleet + web
RL Qwen-VLA only (PPO, ODE→SDE log-prob) RECAP (production) No canonical RL stage No Undisclosed
Evaluation posture OOD manifesto + new benchmarks Real-world task suites Sim + humanoid demos Controlled co-training studies Closed demos
Weights Closed (stated: no plan) Closed Open (Apache 2.0) Closed (papers open) Closed
Distinctive IP Camera-frame delta EEF + H2R synthesis + language-as-interface suite Same-stack MoE + KI + RECAP Cross-attn DiT + ego-data scaling Co-training evidence base Dual-system closed stack

The ironic detail: NVIDIA builds GR00T on Qwen's own open VLM while Qwen keeps its robot models closed — Qwen the LLM vendor powers a competitor's open robot stack that its own robot stack refuses to join.


8. Assessment — is there a coherent vision?

Yes, at the level of doctrine (§4's eight bets), and the doctrine is empirically grounded — VLM4VLA supplies the mechanistic rationale for end-to-end vision tuning and joint (not sequential) VL supervision, and both flagships implement it identically down to the λ = 0.1.

No, at the level of a single technical roadmap. The program currently maintains two incompatible flagship architectures with no public reconciliation, and the suite branding quietly sidelines the May flagship. The pattern resembles Qwen's LLM-era practice of parallel rapid iterations — but LLM iterations share an interface; VLAs with different action spaces, calibration requirements, and state conventions do not migrate as easily.

What would confirm the vision is working (watch-list):

  1. A Qwen-RobotManip 2 / unified successor that resolves the wiring, state, and RL contradictions — ideally with the matched-scale internal comparisons only Qwen can run.
  2. An action-aware Qwen backbone — VLM4VLA's own open question. If any lab ships "a VLM pretrained for VLA," it should be the one that owns both the VLM pretraining pipeline and the diagnosis.
  3. Weights or an API. Every headline number in the program is currently unreproducible; RoboChallenge (#1, externally verified) is the only independent validation. A closed suite with no deployment channel is a paper program.
  4. The agentic planner layer Partially delivered: RobotNav's two-tier agentic system (Qwen3.6-Plus planner) is demonstrated with EQA SOTA and a real-robot episode. Still open: the equivalent planner over RobotManip (long-horizon manipulation), and any release of the planner harness itself.
  5. Whether RobotWorld actually feeds RobotManip (synthetic data / evaluation loop) in a documented way — RobotWorld now cross-evaluates on RobotManip's RoboTwin-IF benchmark, but no policy has been trained on its rollouts or evaluated inside it; its language-only action interface cannot yet consume VLA action outputs, so the loop remains open.

9. Links


10. Related pages

← Back to Home

⚠️ **GitHub.com Fallback** ⚠️