Review Qwen Team VLA - Heungwoo/research GitHub Wiki
Scope: every VLA-relevant paper the Qwen team has authored or co-authored, read as one research program. Papers covered in depth: VLM4VLA (arXiv 2601.03309, ICLR 2026) · Qwen-VLA (arXiv 2605.30280) · Qwen-RobotManip (arXiv 2606.17846) · Qwen-RobotNav (arXiv 2606.18112) · Qwen-RobotWorld (arXiv 2606.17030) Caveat: the papers barely cite one another (the only visible connection is RobotWorld evaluating zero-shot on RobotManip's RoboTwin-IF benchmark). The "program" reconstructed here is an external reading from author lists, timing, and shared design choices — not the team's own stated roadmap.
- In six months (Jan–Jun 2026) the Qwen team went from diagnostic study to full embodied-AI stack. The arc: VLM4VLA (Jan, with Tsinghua) asks which VLM makes a good VLA and why; Qwen-VLA (May) ships the first flagship generalist; the Qwen-Robot Suite (Jun 16–17) productizes the stack into three specialized models — RobotManip (manipulation VLA), RobotNav (navigation VLN), RobotWorld (video world model).
- The program has a coherent core vision: a natively-multimodal Qwen backbone (though the family fragments: Qwen3.5-4B in the flagships, Qwen3-VL in RobotNav, frozen Qwen2.5-VL in RobotWorld) + flow-matching continuous action experts for manipulation + joint VL co-training against forgetting (λ = 0.1 in both flagships, 1.0 in RobotNav) + single cross-embodiment policy switched by text prompts + masked canonical action tensors + synthetic data as the scaling engine + deep skepticism of in-distribution benchmarks + language as the universal interface across policy, navigation, and world model.
- But the two flagship VLAs contradict each other on nearly every remaining axis: VLM↔expert wiring (concatenation vs cross-attention — and RobotManip's own ablation votes against Qwen-VLA's choice), action space (native conventions vs camera-frame delta EEF), proprioception (omitted vs first-class), recipe (4-stage + RL vs single-phase, no RL), data posture (in-house teleop vs open-data-only). Two teams, one org, five weeks apart, zero cross-citations.
- VLM4VLA reads as the program's empirical foundation — its two actionable findings (the vision encoder is the bottleneck and must receive action gradients; sequential embodied-VQA fine-tuning hurts, so supervision must be joint) are followed by both flagships. Its most inconvenient finding (VLM benchmark scores don't predict VLA performance — a 1.7B Kosmos-2 beat a 31B Qwen3VL MoE) did not stop the team from building both flagships on its own newest Qwen3.5-4B.
- The strategic break: closed weights. Despite Qwen's open-weights LLM tradition, no weights are released for any robot model, and the RobotManip/RobotNav README states there is no plan to. Qwen enters embodied AI as a full-stack closed provider — the same posture as Physical Intelligence and Gemini Robotics, opposite to its own LLM identity and to NVIDIA's Apache-2.0 GR00T.
| Date | Paper | arXiv | What it is | Venue / status |
|---|---|---|---|---|
| 2026-01 | VLM4VLA | 2601.03309 | Controlled study: 9 VLM backbones (24 variants) through a minimal <1%-new-params harness on Calvin / SimplerEnv / LIBERO-Long | ICLR 2026 poster; Tsinghua-led, Qwen co-authors |
| 2026-05-28 | Qwen-VLA | 2605.30280 | Flagship generalist VLA: manipulation + navigation + AD-VQA, 11 embodiments + human MANO, 4-stage recipe (T2A → CPT → SFT → PPO RL) | Technical report, 34 pp |
| 2026-06-16/17 | Qwen-RobotManip | 2606.17846 | Manipulation-specialist VLA: alignment-first thesis, camera-frame delta EEF, ~38,100 h open-data corpus, RoboChallenge #1 | Technical report, 44 pp |
| 2026-06-16/17 | Qwen-RobotNav | 2606.18112 | Navigation specialist on Qwen3-VL (2B/4B/8B): parameterized observation interface (token budget · temporal decay · camera weights) with training-time randomization; 15.6M samples; SOTA on VLN-CE/ObjectNav/EVT-Bench/NAVSIM; agentic EQA SOTA with a Qwen3.6-Plus planner | Technical report, in-depth review |
| 2026-06-16/17 | Qwen-RobotWorld | 2606.17030 | Language-actioned video world model: 20B 60-block double-stream MMDiT + frozen Qwen2.5-VL-7B action encoder + Wan-VAE; 8.6M-pair EWK corpus with five-layer action-language annotation; Scene2Robot H2R editing; 1st on EWMBench / DreamGen Bench, best open-source on WorldModelBench / PBench | Technical report, in-depth review |
The June trio was announced together as the Qwen-Robot Suite. Qwen-VLA (May) is conspicuously not part of the suite branding, and the suite's manipulation model is not an iteration of it — it is a different architecture from a different sub-team (§3).
- VLM4VLA → Qwen-VLA: Qiuyue Wang and Mingsheng Li are VLM4VLA co-authors and become Qwen-VLA's first two (equal-contribution) authors; Shuai Bai and Junyang Lin (Qwen core) co-author VLM4VLA, and Shuai Bai is Qwen-VLA's corresponding author. VLM4VLA is led by Jianke Zhang / Jianyu Chen (Tsinghua). The diagnostic study directly seeded the first flagship.
- Qwen-VLA ↔ Qwen-RobotManip: partial overlap (Zhixuan Liang, Pei Lin, Jie Zhang, Jinhui Ye, Sicheng Xie, Shuai Bai, Junyang Lin, Dayiheng Liu, Jingren Zhou) but disjoint leadership — Qwen-VLA answers to Shuai Bai (Qwen-VL lineage); RobotManip's corresponding/lead authors are Chenfei Wu and Xiong-Hui Chen, with a robotics/RL-heavy core-contributor list (Haoqi Yuan, Anzhe Chen, …) that does not appear on Qwen-VLA.
- The most parsimonious organizational reading: two parallel embodied-AI efforts inside Qwen/Tongyi, one grown out of the VLM team, one staffed as a robotics team — with the June suite marking the robotics line as the productized track. This is interpretation, not documented fact.
VLM4VLA is the program's empirical bedrock — a Tsinghua×Qwen controlled study whose findings define the design space the flagships later occupy:
- General VLM benchmark scores do not predict VLA performance (Pearson r = −0.36 on SimplerEnv, −0.19 on LIBERO-Long; only Calvin correlates at +0.84). Kosmos-2 at 1.7B ties π0 and beats Qwen3VL-30B-A3B (31B) on SimplerEnv.
- The vision encoder is the bottleneck, not language. Freezing the ViT during VLA training costs 21–42 points on SimplerEnv; freezing word embeddings costs ±0.2.
- Sequential fine-tuning on embodied auxiliary VQA (Robopoint, BridgeVQA, RoboBrain2, …) consistently hurts downstream VLA performance — all 7 auxiliary tasks tested produce negative Calvin deltas.
- Action-supervised vision-encoder fine-tuning is the one intervention that helps (+18.1 on SimplerEnv when the ViT receives action-token gradients).
- VLM pretraining is necessary (from-scratch collapses by 2.3–2.5 Calvin points) — just not sufficient, and not monotone in VLM quality.
| VLM4VLA finding | Qwen-VLA (May) | Qwen-RobotManip (Jun) | Verdict |
|---|---|---|---|
| Vision encoder must receive action gradients | Backbone (incl. ViT) frozen only during T2A warm-start, unfrozen from CPT onward | Backbone fully unfrozen end-to-end; flow-matching gradients reach the ViT | Followed — neither flagship ships a frozen vision encoder (contrast: TRI LBM freezes PaliGemma throughout) |
| Sequential embodied-VQA FT hurts | No sequential VQA stage; VL supervision only as joint co-training (λ_vl = 0.1) | Same — dual-stream joint co-training (9:1, λ = 0.1); +8.2 pp on RT-C2R Hard | Followed, with a nuance: VLM4VLA tested sequential fine-tuning; both flagships use joint co-training, and RobotManip's ablation shows joint helps. The program implicitly claims the harm was in the sequencing, not the data |
| VQA score doesn't predict VLA quality — benchmark your candidates | Adopts newest in-house Qwen3.5-4B without a backbone-selection ablation | Same backbone, same absence of selection ablation | Not followed. Neither flagship runs the study's own prescribed comparison; Qwen3.5 was not in VLM4VLA's test set. Pragmatics (in-house stack, early fusion) won over the diagnostic |
| Action-supervised vision FT is the missing recipe | Realized implicitly via joint end-to-end training at scale | Same, plus ECoT / 2D-trajectory VL data that inject action semantics into the VLM stream | Followed in spirit — neither uses the paper's specific FAST-token vision-FT recipe, but both make action supervision reach the ViT |
| Minimal deterministic harness for clean comparison | Flow-matching DiT (1.15B) | Flow-matching DiT (small) | N/A — the harness was diagnostic tooling, not a production recommendation |
The scorecard: the two mechanistic findings shaped both flagships; the backbone-selection finding — the paper's headline — was overridden by the obvious institutional incentive to showcase Qwen3.5. Notably, VLM4VLA's closing open question ("what does a VLM trained for VLA look like?") is still unanswered by the program's own flagships: both consume a general-purpose Qwen3.5 rather than an action-aware backbone variant.
Where the two VLAs agree, they agree strikingly precisely. This is the closest thing to a documented "Qwen VLA doctrine":
- Qwen3.5-4B, natively multimodal, early fusion — for the flagships. Both manipulation flagships use the identical backbone — interleaved visual tokens from a dynamic-resolution ViT, no bolt-on encoder. The ~4B choice independently converges with PI (Gemma3-4B) and TRI (PaliGemma2-3B): three labs, three VLM lineages, one size class. But the suite fragments this: RobotNav runs on Qwen3-VL (2B/4B/8B) and RobotWorld uses a frozen Qwen2.5-VL-7B as action encoder — three backbone generations across one program, with no stated rationale for the split.
- Category B flow-matching continuous actions — no discrete action tokens anywhere. No FAST heads, no VQ-VAE, no latent-action vocabularies in either paper (nor in the suite's world model, which uses language as its action interface). The program never even ablates discrete tokens — aligning with LBM's negative result by omission.
- Joint VL co-training as the anti-forgetting mechanism — λ = 0.1 for manipulation, λ = 1.0 for navigation. Both flagships use joint VL supervision at exactly one-tenth weight of the action loss, and neither uses architectural insulation (KI) or freezing; the recurring 0.1 suggests a shared internal recipe. RobotNav keeps the mechanism (co-training against "collapse into reactive action-sequence mappers") but sets λ = 1.0 — a plausible domain difference (navigation's MSE waypoint loss is a weaker gradient source than flow matching), though no paper discusses it.
-
One generalist policy, cross-embodiment by prompt + masked canonical tensor. Both encode the platform as text (Qwen-VLA's sentence template; RobotManip's structured fields) and both use zero-padded fixed action tensors with per-dimension loss masks. Neither uses per-embodiment heads (contrast GR00T's
CategorySpecificMLP). - Synthetic data as the scaling engine, not teleop fleets. Qwen-VLA: 7.2M rendering-free language-only trajectories + RoboInf vision synthesis. RobotManip: 24,808 h of human-to-robot re-rendering — 65% of its corpus. The program's data thesis is that synthesis + curation infrastructure substitutes for proprietary collection (RobotManip makes this explicit by using zero in-house teleop; Qwen-VLA still used >1,000 h in-house).
- Benchmark skepticism as identity. Qwen-VLA stakes its claim on zero-shot DOMINO transfer beating fine-tuned baselines; RobotManip escalates to a full manifesto (from-scratch matches pretrained in-distribution) and ships two new OOD benchmarks (RoboTwin-IF/XE). Both papers argue the field's standard evaluations reward memorization.
- Language as the universal interface of the whole stack. Embodiment prompts select the robot; ECoT reasoning is expressed in language; RobotWorld conditions video prediction on language as "a unified action interface"; RobotNav exposes a parameterized text interface for an agentic planner to switch task modes mid-episode. The suite's implicit System-1/System-2 story routes everything through natural language — the most Qwen-like possible bet.
- Closed weights. No model in the program has released weights; the suite README states no plan to. For a lab whose LLM brand is open weights, this is the clearest signal that embodied AI is being treated as a commercial moat, not an ecosystem play.
Condensed from Review-Qwen-RobotManip §9.1 (full table there):
| Axis | Qwen-VLA | Qwen-RobotManip | Who's right? |
|---|---|---|---|
| VLM↔expert wiring | Concatenation + joint self-attention (16-block, 1.15B DiT) | Cross-attention into last-layer states, alternating vision/language by block parity (10-block, D=768 DiT) | RobotManip's own Table 19 ablation: cross-attention 87.5 vs concatenation 87.0 on LIBERO-Plus — a direct, if narrow, vote against its sibling |
| Expert size | 1.15B | ~10× smaller | Unresolved — no matched-scale comparison exists |
| Action space | Per-dataset native conventions + quantile norm | 80-dim canonical + camera-frame delta EEF (calibration required) | RobotManip's scaling-law ablation is the strongest evidence in either paper: alignment creates the data scaling law |
| Proprioception | Ablated to ≤+1.3 pp, omitted | 80-dim state vector, first-class input | Unreconciled — plausibly action-space-dependent (camera-frame deltas may need state anchoring) |
| Recipe | 4 stages incl. T2A warm-start and PPO RL (ODE→SDE log-prob) | Single-phase dual-stream, no warm-start, no RL | Unresolved; RobotManip gets its OOD results without RL, questioning whether Qwen-VLA's stage-IV complexity earns its keep |
| Data posture | In-house teleop + public + synthetic | Open + synthesized only | RobotManip's open-data-only result is the more disruptive claim if it holds |
| Scope | Manip + nav + AD-VQA in one model | Manip only; nav split into RobotNav | The suite decomposition suggests the specialist view won internally |
Two more observations:
- The generalist-vs-suite fork now has a scoreboard. Qwen-VLA argues one model should do manipulation and navigation (it evaluates on R2R/RxR); three weeks later the suite ships navigation as a separate model — and RobotNav-8B beats Qwen-VLA-Instruct on the same VLN-CE Val-Unseen splits by +14.6 pp R2R SR (72.1 vs 57.5) and +16.9 pp RxR SR (76.5 vs 59.6) (with vastly more navigation training data, to be fair). On this axis the specialist decomposition is empirically vindicated, which reads as the Qwen-VLA generalist line being superseded.
- Neither flagship evaluates the other, and neither compares against π0.6/π0.7. The program's internal contradictions are exactly the controlled experiments the field wants (concat vs cross-attn at matched scale; T2A+RL vs alignment-first; state vs no-state under both action spaces), and only this team can run them.
Read as a product architecture, the June suite decomposes embodied intelligence into three language-interfaced components:
flowchart LR
subgraph AGENT["Agentic planner (Qwen3.6-Plus — demonstrated for navigation, not released)"]
PLAN["Long-horizon goal decomposition<br/>task-mode + observation-config switching<br/>evidence-notebook memory"]
end
subgraph SUITE["Qwen-Robot Suite (Jun 2026)"]
MANIP["Qwen-RobotManip<br/>manipulation VLA<br/>(flow-matching, camera-frame EEF)"]
NAV["Qwen-RobotNav<br/>VLN engine<br/>(parameterized task interface)"]
WORLD["Qwen-RobotWorld<br/>language-conditioned video WM<br/>(data synthesis · eval · planning signals)"]
end
PLAN -- "language sub-goals" --> MANIP
PLAN -- "task mode + params" --> NAV
WORLD -- "synthetic trajectories / rollout eval" --> MANIP
WORLD -- "imagined futures" --> PLAN
classDef m fill:#c8e6c9,stroke:#2e7d32,color:#000
classDef n fill:#bbdefb,stroke:#1565c0,color:#000
classDef w fill:#fff9c4,stroke:#f57f17,color:#000
class MANIP m
class NAV n
class WORLD w
- The decomposition maps cleanly onto the System 0/1/2 framing: RobotManip/RobotNav are System-1 executors, the planner is System-2, RobotWorld is the imagination module — with language as the bus between every layer.
-
The System-2 slot is demonstrated on the navigation side. RobotNav ships an agent-facing tool interface —
nav_qwennav(L_i, τ_i, Φ_i)with runtime-switchable task modes and observation configs — and a working two-tier system (Qwen3.6-Plus planner + trajectory-evidence harness + persistent evidence notebook) that sets EQA SOTA (HM-EQA 76.7, EXPRESS-Bench 79.27) with 77% fewer navigation steps, validated on a real robot. No equivalent planner-over-RobotManip system exists yet. - RobotWorld's three stated application directions (synthetic data generation for policy training, virtual evaluation environments, language-guided planning signals) tie directly back to the two flagships' biggest gaps: RobotManip's synthesis-quality ceiling and the program's benchmark-skepticism (a world model as evaluator is the logical endpoint of "don't trust static benchmarks"). The one visible hand-off so far: RobotWorld evaluates itself zero-shot on RobotManip's RoboTwin-IF benchmark. But RobotWorld's action interface is language only — it cannot consume the continuous actions its sibling VLAs emit, so the policy-evaluation loop remains structurally open.
- The suite's data engines are siblings: RobotManip's H2R synthesis (15 platforms, 24,808 h) and RobotWorld's Scene2Robot training pairs (MANO → 14 robot arms, MuJoCo IK + inpainting) are the same pipeline pattern re-tasked — one produces training trajectories, the other paired video-editing supervision.
- Missing from the suite, notably: an RL layer (Qwen-VLA's stage IV has no successor here), a dexterous-hand story beyond the canonical vector's 12 hand dims, and any humanoid whole-body control — the suite is arms-and-wheels, not humanoids (contrast GR00T).
| Axis | Qwen program | Physical Intelligence | NVIDIA GR00T | TRI LBM | Google Gemini Robotics |
|---|---|---|---|---|---|
| Entry point | VLM vendor moving down-stack | Robotics-native startup | Platform/ecosystem vendor | Auto-industry research lab | VLM vendor moving down-stack |
| Backbone | Qwen3.5-4B (own) | Gemma3-4B (external) | Qwen3-VL-2B (external — Qwen's!) | PaliGemma2-3B (external) | Gemini (own) |
| Data thesis | Open + synthesis (RobotManip: zero proprietary) | In-house fleet-first | Data pyramid: ego-video + sim + teleop | Modest curated teleop + co-training science | Undisclosed, fleet + web |
| RL | Qwen-VLA only (PPO, ODE→SDE log-prob) | RECAP (production) | No canonical RL stage | No | Undisclosed |
| Evaluation posture | OOD manifesto + new benchmarks | Real-world task suites | Sim + humanoid demos | Controlled co-training studies | Closed demos |
| Weights | Closed (stated: no plan) | Closed | Open (Apache 2.0) | Closed (papers open) | Closed |
| Distinctive IP | Camera-frame delta EEF + H2R synthesis + language-as-interface suite | Same-stack MoE + KI + RECAP | Cross-attn DiT + ego-data scaling | Co-training evidence base | Dual-system closed stack |
The ironic detail: NVIDIA builds GR00T on Qwen's own open VLM while Qwen keeps its robot models closed — Qwen the LLM vendor powers a competitor's open robot stack that its own robot stack refuses to join.
Yes, at the level of doctrine (§4's eight bets), and the doctrine is empirically grounded — VLM4VLA supplies the mechanistic rationale for end-to-end vision tuning and joint (not sequential) VL supervision, and both flagships implement it identically down to the λ = 0.1.
No, at the level of a single technical roadmap. The program currently maintains two incompatible flagship architectures with no public reconciliation, and the suite branding quietly sidelines the May flagship. The pattern resembles Qwen's LLM-era practice of parallel rapid iterations — but LLM iterations share an interface; VLAs with different action spaces, calibration requirements, and state conventions do not migrate as easily.
What would confirm the vision is working (watch-list):
- A Qwen-RobotManip 2 / unified successor that resolves the wiring, state, and RL contradictions — ideally with the matched-scale internal comparisons only Qwen can run.
- An action-aware Qwen backbone — VLM4VLA's own open question. If any lab ships "a VLM pretrained for VLA," it should be the one that owns both the VLM pretraining pipeline and the diagnosis.
- Weights or an API. Every headline number in the program is currently unreproducible; RoboChallenge (#1, externally verified) is the only independent validation. A closed suite with no deployment channel is a paper program.
-
The agentic planner layerPartially delivered: RobotNav's two-tier agentic system (Qwen3.6-Plus planner) is demonstrated with EQA SOTA and a real-robot episode. Still open: the equivalent planner over RobotManip (long-horizon manipulation), and any release of the planner harness itself. - Whether RobotWorld actually feeds RobotManip (synthetic data / evaluation loop) in a documented way — RobotWorld now cross-evaluates on RobotManip's RoboTwin-IF benchmark, but no policy has been trained on its rollouts or evaluated inside it; its language-only action interface cannot yet consume VLA action outputs, so the loop remains open.
- VLM4VLA: https://arxiv.org/abs/2601.03309 · OpenReview
- Qwen-VLA: https://arxiv.org/abs/2605.30280 · blog
- Qwen-RobotManip: https://arxiv.org/abs/2606.17846 · blog · GitHub
- Qwen-RobotNav: https://arxiv.org/abs/2606.18112 · blog · GitHub
- Qwen-RobotWorld: https://arxiv.org/abs/2606.17030 · blog
- Suite announcement: Alibaba Cloud blog
- VLM4VLA (in-depth) · Qwen-VLA (in-depth) · Qwen-RobotManip (in-depth) · Qwen-RobotNav (in-depth) · Qwen-RobotWorld (in-depth) — the five constituent deep-dives
- π series evolution · GR00T series — the competing lab-program reviews
- LBM Co-training — the co-training evidence base both flagships align with
- VLA Architectures · VLM↔Action Connection — where the wiring disagreement lives
- Cross-Embodiment · Action Space: EEF vs Joint — the axes RobotManip's alignment thesis extends
- System 0/1/2 — the framing the suite decomposition maps onto
- World Models — context for Qwen-RobotWorld
- Knowledge Insulation — the architectural alternative to the program's λ=0.1 co-training bet
← Back to Home