Review Qwen RobotManip - Heungwoo/research GitHub Wiki
In-Depth Review β Qwen-RobotManip: Alignment Unlocks Scale for Robotic Manipulation Foundation Models
Paper: Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models Authors: Qwen Team β core contributors Haoqi Yuan*, Zhixuan Liang*, Anzhe Chen*, Ye Wang*, Haoyang Li*, Pei Lin*, Yiyang Huang*, Zixing Lei*, Tong Zhang*, Jiazhao Zhang, Jie Zhang, Jingyang Fan, Gengze Zhou, Qihang Peng, Chenxu Lv, Xiaoyue Chen, An Yang, Fei Huang, Junyang Lin, Dayiheng Liu, Jingren Zhou, Chenfei Wuβ (corresponding), Xiong-Hui Chenβ‘β (project lead + corresponding), + 14 contributors Affiliation: Qwen Team, Alibaba Group arXiv: 2606.17846 Β· v1 Jun 16, 2026 Β· v2 Jun 17, 2026 (44 pages, cs.RO, CC BY 4.0) Code: github.com/QwenLM/Qwen-RobotManip Β· Blog: qwen.ai/blog?id=qwen-robotmanip Weights: the GitHub README states there is no plan to release model weights (for Qwen-RobotManip or Qwen-RobotNav) as of this review β see Β§11.2
This is the long-form companion to the per-paper summary. Companion reviews to read alongside: Qwen-VLA (the Qwen team's first VLA, May 2026 β the two papers are siblings and disagree architecturally) Β· Action Space: EEF vs Joint Β· Cross-Embodiment Β· LBM Co-training Β· VLA Architectures Β· VLMβAction Connection Β· LIBERO-Plus.
- The Qwen team's second VLA in two months β with a different thesis. Where Qwen-VLA (May 2026) was "one generalist, many benchmarks, 4-stage recipe + RL", Qwen-RobotManip (June 2026) is an alignment-first scaling manifesto: a canonical 80-dim state-action representation, a camera-frame delta EEF action space, and in-context policy adaptation are presented as prerequisites that make heterogeneous data scale at all. The ablation showing that unaligned representations produce no scaling law while aligned ones scale log-linearly from 1%β100% data is the paper's core claim.
- ~38,100 hours of pretraining data with zero proprietary collection. ~11,420 h open-source robot data (9 datasets) + 1,933 h egocentric human video (EgoDex/VITRA/EgoVerse) + 24,808 h human-to-robot synthesized demonstrations β each human clip re-rendered onto 15 bimanual robot platforms via a SAM3 β ProPainter β base-search + MuJoCo-IK β depth-composite pipeline. Two-thirds of the corpus is synthetic H2R data.
- An evaluation manifesto: standard benchmarks are broken. The paper shows from-scratch models match pretrained ones on LIBERO/RoboTwin (in-distribution), then argues OOD transfer is the only faithful measure β and introduces two new benchmarks: RoboTwin-IF (instruction following under held-out language templates) and RoboTwin-XE (zero-shot transfer to unseen robot morphologies).
- Headline OOD numbers vs Ο0.5: 91.4% LIBERO-Plus (vs 84.4), 69.4% RoboTwin-Clean2Rand Hard (vs 47.9), 45.6% EBench (vs 27.1), 35.9% RoboCasa365 (vs 16.9), 72.2% RoboTwin-IF (vs 49.6), 23.9% zero-shot cross-embodiment on RoboTwin-XE (3.2Γ Ο0.5's 7.5). Real-world CobotMagic ALOHA: 88.6% ID / 87.5% OOD vs Ο0.5's 42.9 / 37.5. Ranked 1st on RoboChallenge Table30-v1 generalist track (45% SR / 59.83 process score vs DM0's 37 / 48.43).
- Architecturally it contradicts Qwen-VLA. RobotManip uses a small 10-block cross-attention DiT (D=768) reading last-layer VLM hidden states β and its own ablation finds last-layer cross-attention beats the concatenation + joint self-attention design that Qwen-VLA shipped four weeks earlier. No RL stage, no T2A warm-start; a single dual-stream co-training pretraining phase instead.
By June 2026 the Qwen team has published two distinct manipulation foundation models:
| Qwen-VLA (arXiv 2605.30280, May 2026) | Qwen-RobotManip (arXiv 2606.17846, Jun 2026, this paper) | |
|---|---|---|
| Central thesis | Compression: languageβaction decompression prior (T2A) + 4-stage recipe ending in RL | Alignment unlocks scale: unified representation/motion/behavior alignment is the prerequisite for data scaling |
| Data posture | >1,000 h in-house teleop + public + synthetic | Open-source + synthesized only, zero proprietary collection (~38,100 h) |
| Scope | Manipulation + navigation + AD-VQA generalist | Manipulation only (a sibling Qwen-RobotNav is referenced in the repo) |
| Evaluation posture | Standard benchmark head-to-heads as a single generalist | OOD-first manifesto + two new benchmarks (RoboTwin-IF, RoboTwin-XE) |
Author overlap is partial (Zhixuan Liang, Pei Lin, Jie Zhang, Jinhui Ye, Sicheng Xie, Shuai Bai, Junyang Lin, Dayiheng Liu, Jingren Zhou appear on both), and the corresponding authors differ (Shuai Bai for Qwen-VLA; Chenfei Wu / Xiong-Hui Chen here). Neither paper cites the other. The most reasonable reading is two parallel efforts inside Tongyi/Qwen, and this one β bundled with Qwen-RobotNav under a "Qwen-Robot" suite umbrella β is the manipulation-specialized line.
- Alignment is a scaling prerequisite, not an engineering nicety. The Β§6.4 controlled data-scaling experiment (nested 1/5/10/25/50/100% subsets, held-out validation across 15 embodiment types Γ 154 tasks) is the first published demonstration that the same cross-embodiment corpus produces a clean log-linear MSE scaling law under an aligned action representation and an erratic, non-improving curve without it. This reframes debates about "does cross-embodiment data help?" (Cross-Embodiment) β the answer may depend on representation alignment more than on data volume.
- The data barrier may be lower than assumed. 38,100 hours with no proprietary teleop β two-thirds of it synthesized from human video β directly challenges the PI-style in-house-data-first posture. If the OOD results hold up, the moat shifts from data collection to synthesis + curation infrastructure.
- In-distribution benchmarks systematically overestimate. Figure 4's from-scratch-matches-pretrained result on LIBERO/RoboTwin (e.g., Qwen-RobotManip-scratch 98.2 LIBERO vs Ο0.5's 97.6) is an indictment of most 2025β2026 VLA evaluation practice, and the paper backs it with the observation that none of the action-representation variants show data-scaling gains on the in-distribution Easy settings while the OOD Hard settings separate them cleanly.
Qwen-RobotManip is the largest-scale commitment yet to camera-frame delta EEF actions (following Chen et al. 2025's embodiment-equivariance line). This connects to the Action Space: EEF vs Joint debate but goes a step further than "EEF beats joint": the claim is that which frame the EEF delta lives in decides whether cross-embodiment transfer works at all. RoboTwin-XE gives the cleanest evidence: joint-space zero-shot transfer to UR5 is 4.1%, camera-frame EEF is 22.8% (5.6Γ).
flowchart TB
subgraph OBS["Multi-view observations + prompt"]
V["Camera views (left / front / right / wrist)<br/>current + optional history frames"]
P["Structured embodiment prompt<br/>embodiment Β· instruction Β· speed Β· fps Β· camera view direction"]
ECOT["(co-training) ECoT / VL QA batches"]
end
subgraph CTX["In-context policy adaptation (Context variant)"]
C1["H history chunks (o_h, s_h, a_h)"]
C2["MLP_s / MLP_a encoders + temporal & slot embeddings"]
end
OBS --> VLM["Qwen3.5-4B backbone<br/>natively multimodal, early fusion<br/>ViT w/ dynamic-resolution spatial merging<br/>D_vlm = 2560"]
CTX -- "unified mode: context tokens appended to VLM input" --> VLM
VLM -- "last-layer hidden states (vision + language)" --> XATT
subgraph DIT["Flow-matching action expert β DiT, N = 10 blocks, D_act = 768, 12 heads"]
XATT["Cross-attention (alternating):<br/>even blocks β visual tokens<br/>odd blocks β language tokens"]
SA["Self-attention over state + noisy-action tokens<br/>CaPE (32/64 head dims) + RoPE (32/64)"]
FF["SwiGLU FFN Β· AdaLN(-Zero RMSNorm) conditioning:<br/>denoise timestep Β· EEF-type codebook Β· camera-params flag"]
end
STATE["Proprioceptive state (80-dim canonical)<br/>2-layer MLP, prepended"] --> DIT
NOISE["Noisy action chunk x_t"] --> DIT
DIT --> VEL["Velocity field v = a β Ξ΅"]
VEL --> EULER["4 Euler steps (10 for Context variant)"]
EULER --> ACT["Action chunk β 80-dim canonical vector<br/>camera-frame delta EEF (or abs joint)"]
classDef vlm fill:#bbdefb,stroke:#1565c0,color:#000
classDef act fill:#c8e6c9,stroke:#2e7d32,color:#000
classDef inp fill:#fff9c4,stroke:#f57f17,color:#000
class VLM vlm
class DIT,XATT,SA,FF,VEL,EULER,ACT act
class OBS,CTX,STATE,NOISE inp
- Qwen3.5-4B (Team, 2026) β same backbone family as Qwen-VLA. Natively multimodal with early vision-language fusion: ViT visual tokens with dynamic-resolution spatial merging interleaved directly into the text stream. Last-layer hidden dimension D_vlm = 2560.
- The backbone is trained end-to-end β flow-matching gradients flow into the VLM (no freezing, no Knowledge Insulation-style gradient routing). Forgetting is managed by data (dual-stream VL co-training, Β§5.1) rather than by architecture.
- N = 10 transformer blocks, D_act = 768, 12 attention heads β dramatically smaller than Qwen-VLA's 1.15B/16-block expert (parameter count is not stated, but at these dimensions it is roughly an order of magnitude smaller).
- Each block: self-attention over the concatenated state + noisy-action token sequence, then cross-attention to VLM last-layer hidden states, then a SwiGLU FFN. Cross-attention alternates by block parity: even-indexed blocks attend to visual tokens, odd-indexed blocks to language tokens β letting the expert ground actions in observations and instructions at separate processing stages.
- Proprioceptive state enters via a 2-layer MLP, prepended to the noisy action tokens.
- Conditioning via adaptive layer normalization (the figure labels it AdaZeroRMSNorm): denoising timestep + two additional embeddings (Β§3.5).
- Flow matching: t ~ Beta(1, 1.5), interpolant x_t = (1βt)Ξ΅ + tΒ·a, velocity target v = a β Ξ΅. Inference: 4 Euler steps (the Context variant needs 10 β see Β§8.2).
In the Review-VLA-Architecture taxonomy this is Category B (separate flow-matching expert) with GR00T-style cross-attention into VLM hidden states β not the concatenation + joint-self-attention wiring Qwen-VLA chose. Notably, Β§6.4's architecture ablation directly compares three wirings (layer-wise self-attention fusion; last-layer self-attention concatenation β i.e., the Qwen-VLA pattern; last-layer cross-attention with learned query tokens) and finds cross-attention wins on LIBERO-Plus (87.5 vs 87.0 vs 86.4) at the lowest compute cost. One Qwen paper's ablation quietly contradicts the other Qwen paper's design choice.
The cross-embodiment interface, and the first of the three "alignment" pillars:
| Segment | Dims | Content |
|---|---|---|
| Per-arm block Γ 2 | 29 each | Joint positions (7) + EEF pose (3 pos + 6D rotation = 9) + gripper (1) + dexterous-hand joints (12) |
| Reserved (shared) | 22 | Future DoF, e.g. mobile-base velocity |
| Total | 80 |
- States are absolute; actions are absolute for joints and relative deltas for EEF (rotation deltas as 3D rotation vectors; state rotations as 6D continuous representation).
- Each embodiment populates only its own slots; a per-dimension binary mask excludes zero-padded entries from the loss. The mask is AND-composed from three sources: the embodiment slot mask, a step-validity mask from the curation pipeline (once a step is invalid, all subsequent steps mask out to preserve causal consistency), and a per-hand validity mask for ego-human data that zeroes an arm slot from the moment the hand leaves the camera view.
- The DiT actually processes N_ee β {1,2} 40-dim per-end-effector tokens (29 active + 11 reserved) jointly via self-attention.
Same family as Qwen-VLA's zero-padded fixed tensor and GR00T's per-embodiment encoders, but with semantically fixed slot positions β and Β§6.4 shows the semantics matter: naive concatenate-and-pad ("w/o UnifiedSpace") destroys the scaling law.
The paper's most distinctive architectural contribution:
- Camera-frame delta pose (Eq. 5): the relative EEF rotation is conjugated into the reference camera frame, and the EEF displacement is projected into camera coordinates. Property: actions that look similar in the image are numerically proximate, regardless of robot base placement or morphology. Requires calibrated intrinsics + extrinsics at train and test time. The more compact full-camera-frame alternative (Eq. 6) was rejected as more sensitive to calibration error and long-tail translation coupling.
- Camera-aware positional encoding (CaPE): camera extrinsics are injected into the DiT's cross-attention as a rotational positional encoding occupying 32 of each 64-dim attention head (the other 32 are RoPE for temporal indexing). Following GTA/PRoPE practice, CaPE is applied to queries, keys, values, and attention outputs. Because it is rotational, the global world origin cancels in dot-product attention β only relative cameraβtoken poses survive. Intrinsics enter as a learned linear projection of normalized image-plane coordinates added per visual token.
- Reference-camera selection is randomized during training (any view for single-arm; shared head-camera vs per-wrist-camera strategies for dual-arm), so the policy learns to denoise in whatever frame the CaPE indicates.
- Fallback mode: an auxiliary binary flag embedding (camera parameters available or not) switches the action space between camera-frame delta and robot-base relative mode, so uncalibrated datasets remain usable.
embodiment: robot_aloha
instruction: Take the toy off the table and put it on the mat.
speed: 1000 # episode length, binned at 500 steps
fps: 30
camera view direction: arm side
Fields: platform tag, instruction, speed (episode-length bin), FPS, and camera-view direction (arm side / opposite side). The embodiment / speed / fps fields are randomly dropped with p = 0.15 during training for robustness to missing metadata. The Β§6.4 ablation ladder: soft prompt hurts (β1.5 avg vs no prompt), plain language tag helps slightly (+0.7), the structured prompt helps +2.2β3.2 points β i.e., the temporal metadata (FPS/speed), not the identity tag, carries most of the value.
The "-Context" variant conditions on a window of H recent execution chunks, each an (observation, state, action-chunk) triplet:
- Injection: historical frames prepend to the current frame through the VLM visual pathway (with an image-count annotation in the prompt); states and action chunks go through lightweight MLP encoders with temporal + slot embeddings, and the resulting context tokens are appended to the VLM input sequence ("unified mode" β chosen over injecting into the DiT, "dual mode", for richer cross-modal integration).
- Stochastic context sampling is the critical trick: always feeding the most recent chunks lets the model win training loss by copying the last action chunk (a recency shortcut that collapses the mechanism). Training instead samples the context window from a random position in the episode, forcing the model to extract the episode's behavioral profile (velocity patterns, grasping style, kinematic signature). At deployment, a rolling window of the most recent H chunks is used.
- The paper's interpretation: context works as an implicit embodiment identifier β intra-episode kinematics tell the policy what body it is driving, complementing the static prompt. Supporting evidence: the biggest per-dimension LIBERO-Plus gain from Context is Robot-state perturbation (75.5 β 83.9) and Camera (87.2 β 89.9).
- Cost: context makes the action distribution more complex; 4 denoising steps become insufficient (jittery motion, performance back at baseline) and 10 steps are required to unlock the benefit (Β§8.2). Also, at episode start the context is all zero-padding and the policy tends to hesitate before moving β the stated reason both Context and context-free variants exist.
| Data type | Embodiment | Sources | Setting | Hours |
|---|---|---|---|---|
| Robot β single-arm | Single-arm | OXE (Fractal/Bridge/BC-Z), RoboMIND, DROID, RH20T, β¦ | Tabletop | 3,808 |
| Robot β dual-arm | Dual-arm | AgiBotWorld-Beta, RoboCOIN, RDT, β¦ | Tabletop | 6,744 |
| Robot β mobile & humanoid | Mobile/humanoid | InternData-A1, Galaxea Open-World | Tabletop & indoor | 868 |
| Human | Human hands | EgoDex, VITRA, EgoVerse | Tabletop & in-the-wild | 1,933 |
| Human-to-Robot | 15 dual-arm platforms | Synthesized from the human data | Tabletop & in-the-wild | 24,808 |
Robot subtotal ~11,420 h across nine open-source datasets: OXE ~600 h (Fractal + Bridge + BC-Z only), AgiBotWorld-Beta ~2,400 h (gripper subset, ~200 task types), RoboMIND 1.0 + 2.0 ~1,400 h, Galaxea Open-World ~500 h, RoboCOIN ~430 h (10 embodiment types), DROID ~500 h (95k trajectories), RH20T ~1,100 h (contact-rich, 4 embodiments), RDT-1B 29 h, InternData-A1 >3,600 h (high-fidelity sim). Egocentric: EgoDex 732 h used (of 829), VITRA-processed Ego4D + EPIC-KITCHENS 247 h, EgoVerse industry portion 954 h; all hand poses unified to MANO parameters + 21 keypoints.
Each egocentric human clip is converted into robot demonstrations on 15 bimanual configurations (two identical arms from: Panda, UR5e, ARX-L5, xArm7, Sawyer, Kinova Gen3, IIWA, Jaco, FR3, UR10e, ViperX, WidowX, Piper, YAM, AgileX ALOHA), a 1,933 h β 24,808 h multiplication. Pipeline (inspired by Phantom/Masquerade):
- Action retargeting. Virtual finger k_vf = 0.7Β·k_index + 0.3Β·k_middle; EEF position = midpoint of thumb-tip and virtual finger; gripper width = their distance; gripper orientation built as a right-handed frame from the jaw-line axis and wrist-to-fingertip direction, with a handedness sign flip so both hands map to one gripper frame. SavitzkyβGolay filtering on positions/widths, Gaussian-weighted SLERP on orientations.
- Visual removal. SAM3 text-prompted arm masks β ProPainter optical-flow-guided inpainting β clean background plates.
- Base placement search. Egocentric trajectories are embodiment-free (no robot base exists), so base pose is optimized: grid search around the trajectory centroid maximizing the fraction of representative keyframes with feasible IK, independently per morphology.
- Rendering + compositing. MuJoCo (mink) IK tracks the smoothed trajectory; the rendered robot is composited onto the inpainted plate using an occlusion mask from Depth Anything v3 metric depth vs rendered robot depth.
- Speed alignment. Human motion is faster than teleop, so ego sources are frame-subsampled during training: EgoDex to 60% (~1.7Γ slower), EgoVerse to 45% (~2.2Γ), VITRA to 25% (~4Γ).
Applied to all pretraining data (and deliberately disabled for SFT, Β§5.2):
- Sudden-change detection β median + SavitzkyβGolay trend, flag on residual β§ (acceleration β¨ jerk) thresholds; per-dataset thresholds; frame-level removal up to full-episode discard (e.g., InternData-A1 collision episodes dropped entirely).
- State-action trend alignment β cross-correlation lag estimation + directional-agreement metric on lag-aligned first differences; DA < ~0.6β0.7 β episode excluded. This stage found 81% of RoboMIND UR-type episodes faulty β a remarkable public data-quality datapoint.
- Extreme-value filtering β protects the [q01, q99] β [β1,1] quantile normalization; gripper dims exempt (bimodal).
- FK consistency β Pinocchio FK vs logged EEF poses; primarily corrects (TCP offsets, shoulder-frame β world-frame) rather than filters; revealed same-robot-different-convention conflicts across datasets β the empirical motivation for the canonical representation.
- Base-frame / orientation alignment β per-dataset rotation corrections so +x is always robot-forward.
Cross-modal: instruction consistency (subtask segmentation β structured reasoning-guided VLM judgment β multi-VLM cross-model adjudication for flagged clips); video-state consistency (render URDF + joint states into the image, compare against a fine-tuned SAM3 robot mask by IoU; low overlap β fix camera parameters or drop); video quality (black/corrupt/blur/static removal, explicitly preserving gripper-closure keyframes).
Six categories: (1) general visual understanding, (2) spatial perception/reasoning (2D/3D grounding, pointing, counting, depth/distance/viewpoint, manipulation feasibility), (3) OCR/documents, (4) multimodal specialized knowledge, (5) instruction following/multilingual/pure text, and (6) embodied-centric data, which is the novel part:
- Embodied Chain-of-Thought (ECoT): three-part annotations (scene description β task-progress judgment with explicit "Task complete / not yet complete" β next atomic action from a 16-type taxonomy), synthesized by Qwen3.6-Plus in thinking mode using privileged context (episode memory summary, 6-frame future preview, coarse progress estimate) that is excluded from the training inputs β the model must produce the ECoT from current observation + instruction only.
- Egocentric video understanding: fine-grained hand/arm-motion and object-state-change descriptions over 4-frame clips (1.5β3 s), static clips filtered.
- 2D trajectory prediction: future EEF / hand trajectories as normalized image-coordinate sequences, low-motion samples filtered.
Sources include RoboPoint, RefSpatial, PixMo, CapsFusion, plus proprietary data.
No multi-stage curriculum (contrast: Qwen-VLA's four stages, LBM's three). One phase, two mutually exclusive batch streams:
- VLA stream: full manipulation corpus β masked flow-matching loss (Eq. 11), a per-sample average over valid mask entries so every sample contributes equally regardless of active-dimension count.
- VLM stream: the 28M VL mixture β standard next-token CE.
- Mixture ratio 9:1 (robot : VL); joint loss L = L_FM + λ·L_VLM with λ = 0.1; separate learning rates for backbone vs expert.
- K_repeat = 8: per VLA sample, the expert runs 8 independent (noise, timestep) draws against one VLM forward pass β amortizing the expensive backbone forward. A pragmatic efficiency trick not present in the Ο/GR00T/LBM recipes.
Gradients from the flow-matching loss do update the backbone (unlike LBM's frozen PaliGemma and unlike KI's insulated CE-only backbone gradients). The 9:1 co-training stream is the sole anti-forgetting mechanism β and Β§8.4 shows it carries real weight (β8.2 pp on RoboTwin-C2R Hard without VL data).
- Standard protocol: per target domain, all demonstrations pooled into one generalist SFT set (no per-task specialists); flow-matching loss only (no VL CE); curation filters disabled (keep every valid demo); color-jitter augmentation; fewer GPUs/steps than pretraining.
- The diagnosis: extended benchmark SFT causes what the paper names VLA-to-VA degradation β the policy stops conditioning on language and becomes a vision-action pattern matcher, because benchmark SFT data has concentrated layouts, shared train/test visual patterns, and weak language diversity. RoboTwin-IF exists precisely to measure this failure mode.
- Mixed post-training (optional, Β§6.5.1): co-train SFT with 10% VL data + auxiliary VLA data (pretraining trajectories filtered by distributional proximity, aligned in EEF space; 75% of the VLA portion). Result on RoboTwin-IF: plain SFT decays with training steps (overfitting), +VL mitigates, +VL+VLA eliminates the decay entirely β performance keeps rising through 90k steps to 75.8%. Crucially, mixed post-training without the UnifiedEEF representation collapses to 0.0% β alignment is the prerequisite for absorbing mixed data even at SFT time.
Remote-server inference over WiFi with Real-Time Chunking (RTC) (Black et al., 2026) generating the next chunk asynchronously during execution. No end-to-end latency numbers are published.
No RL stage (vs Qwen-VLA's PPO, PI's RECAP). No action-expert warm-start (vs Qwen-VLA's T2A). No discrete action tokens anywhere β consistent with LBM's finding. No stated optimizer/LR/batch/GPU-hours for pretraining β the same reproducibility gap as Qwen-VLA.
On LIBERO and RoboTwin (in-distribution), models without large-scale robot pretraining match or beat pretrained ones β StarVLA 98.0 / Qwen-RobotManip-scratch 98.2 on LIBERO vs Ο0.5's 97.6. On OOD versions the ordering inverts hard: StarVLA collapses 85.7 β 10.6 (RoboTwin Easy β Clean2Rand), scratch collapses 71.6 β 22.6, while pretrained Qwen-RobotManip retains 73.2 β 62.6 (~86% retention vs Ο0.5's ~66% and scratch's ~30%). The paper's practitioner argument: a deployer never has the benchmark's distribution β only OOD transfer efficiency matters.
Built on RoboTwin 2.0; models fine-tuned on RoboTwin-Clean only; evaluated with held-out unseen instruction templates + per-object description variants. Five suites: Pick-Diverse-Object (target grounding among 3 distractors from a 12-item pool), Place-Relative ("beside"/"on top of" spatial relations), Operate-Mic-Drawer (multi-step bimanual sequencing, optionally arm-specified), Operate-Stapler (verb discrimination β press vs move β with the placement pad always present as distractor), Operate-Tabletop (three-way verb-and-target discrimination in a multi-affordance scene).
Fine-tune on RoboTwin-Clean (AgileX only), evaluate under RoboTwin-Hard with the platform replaced by ARX-X5, UR5-WSG, or Franka Panda β initial EEF poses IK-aligned, camera extrinsics/scenes/seeds held identical, zero target-embodiment data. Tests both visual (arm appearance) and kinematic (DoF, joint arrangement, link lengths, workspace) generalization.
| Method | LIBERO | RoboTwin-Easy | RoboTwin-Hard |
|---|---|---|---|
| Ο0 | 94.4 | 65.9 | 58.4 |
| Ο0.5 | 97.6 | 82.7 | 76.8 |
| StarVLA | 98.0 | 85.7 | 87.3 |
| Abot-M0 | 98.6 | 86.1 | 85.1 |
| Being-H0.7 | 99.2 | 90.2 | 89.6 |
| Qwen-RobotManip-scratch | 98.2 | 88.7 | 88.4 |
| Qwen-RobotManip | 99.1 | 93.4 | 92.5 |
| Qwen-RobotManip-Context | 99.2 | 93.7 | 94.0 |
State-of-the-art or tied β but the paper itself immediately discounts these numbers as unable to distinguish generalization from memorization.
| Method | Camera | Robot | Language | Light | Background | Noise | Layout | Total |
|---|---|---|---|---|---|---|---|---|
| Ο0 | 13.8 | 6.0 | 58.8 | 85.0 | 81.4 | 79.0 | 68.9 | 53.6 |
| Ο0.5 | 78.4 | 73.6 | 80.8 | 96.2 | 94.1 | 89.0 | 84.5 | 84.4 |
| StarVLA | 52.5 | 49.8 | 88.5 | 95.7 | 95.7 | 73.0 | 76.9 | 74.1 |
| Abot-M0 | 60.4 | 67.9 | 86.4 | 96.2 | 91.6 | 86.4 | 82.6 | 80.5 |
| Cosmos-Policy | 75.8 | 63.3 | 81.7 | 96.5 | 88.9 | 92.7 | 82.2 | 82.2 |
| Being-H0.7 | 82.0 | 59.0 | 82.8 | 97.8 | 90.0 | 93.5 | 88.5 | 84.8 |
| Qwen-RobotManip-scratch | 70.4 | 44.9 | 88.1 | 95.8 | 95.5 | 84.4 | 79.1 | 78.3 |
| Qwen-RobotManip | 87.2 | 75.5 | 85.6 | 96.6 | 97.7 | 97.7 | 87.3 | 89.0 |
| Qwen-RobotManip-Context | 89.9 | 83.9 | 86.5 | 98.6 | 99.9 | 97.9 | 87.5 | 91.4 |
The structural reading the paper draws: Language/Light/Background robustness comes free from the VLM (scratch models match pretrained there); Camera and Robot-state robustness specifically require large-scale robot pretraining (+34.7 / +16.8 for pretrained-vs-scratch on Camera; scratch craters at 44.9 on Robot). Context adds an implicit kinematic prior on top (+8.4 Robot, +2.7 Camera).
| Method | Easy | Background | Light | Clutter | Height | Hard |
|---|---|---|---|---|---|---|
| StarVLA | 58.1 | 27.1 | 50.9 | 24.2 | 48.4 | 10.6 |
| GR00T-N1.7 | 43.6 | 40.4 | 41.9 | 27.1 | 39.0 | 20.7 |
| Ο0.5 | 73.1 | 67.0 | 69.2 | 57.9 | 67.6 | 47.9 |
| Abot-M0 | 70.7 | 56.5 | 68.8 | 46.0 | 56.3 | 36.0 |
| Qwen-RobotManip-scratch | 71.6 | 60.6 | 70.7 | 24.6 | 63.6 | 22.6 |
| Qwen-RobotManip (joint) | 73.2 | 74.6 | 68.4 | 61.3 | 71.0 | 62.6 |
| Qwen-RobotManip (eef) | 74.0 | 75.8 | 70.1 | 59.8 | 69.4 | 60.8 |
| Qwen-RobotManip-Context (joint) | 84.7 | 82.4 | 84.2 | 75.4 | 79.5 | 69.4 |
| Qwen-RobotManip-Context (eef) | 85.0 | 82.4 | 84.7 | 66.8 | 82.9 | 64.0 |
Notable details: Background randomization helps the pretrained model (74.6 > 73.2 Easy β diverse pretraining scenes make random backgrounds more in-distribution than sterile white); Clutter is the axis where scratch pretraining-free models die (71.6 β 24.6); Context adds +11.5 Easy / +6.8 Hard.
| Method | RoboCasa365 Atomic | Comp-Seen | Comp-Unseen | Total |
|---|---|---|---|---|
| Ο0.5 | 39.6 | 7.1 | 1.2 | 16.9 |
| GR00T-N1.5 | 50.7 | 14.8 | 2.7 | 23.9 |
| RLDX-1 | 63.0 | 27.5 | 5.4 | 33.2 |
| Qwen-RobotManip | 68.6 | 20.1 | 14.9 | 35.9 |
| Qwen-RobotManip-Context | 63.9 | 22.6 | 11.2 | 33.8 |
Composite-Unseen (long-horizon in OOD scenes) 14.9% nearly triples RLDX-1's 5.4%. EBench (Shanghai AI Lab, Isaac Sim, dual-arm mobile Lift2 + R5a, 26 task types / 794 instances): 45.6% SR / 60 score overall vs Ο0.5's 27.1 / 41, with Table-Top dexterous SR nearly 4Γ Ο0.5 (50.0 vs 12.9). Per-dimension: Qwen-RobotManip is nearly flat from Background (45.3) to Mix (46.8) while Ο0.5 declines 33% β perturbation stacking barely touches it.
| Method | Pick-Diverse | Place-Rel. | Op.-Mic-Drawer | Op.-Stapler | Op.-Tabletop | Avg. |
|---|---|---|---|---|---|---|
| StarVLA | 11 | 13 | 0 | 49 | 74 | 29.4 |
| GR00T-N1.7 | 20 | 17 | 0 | 14 | 32 | 16.6 |
| Ο0.5 | 44 | 20 | 15 | 92 | 66 | 49.6 |
| Qwen-RobotManip | 79 | 57 | 42 | 90 | 93 | 72.2 |
| Qwen-RobotManip-Context | 77 | 71 | 33 | 89 | 90 | 72.0 |
+22.6 pp average over Ο0.5; GR00T-N1.7's 16.6% is a striking demonstration of VLA-to-VA degradation on a model that scores respectably on visual-perturbation benchmarks.
| Method | ARX-X5 | UR5-WSG | Franka Panda | Total |
|---|---|---|---|---|
| Ο0.5 (joint) | 24.6 | 2.2 | 0.9 | 9.2 |
| Ο0.5 (eef) | 11.5 | 10.0 | 1.1 | 7.5 |
| Qwen-RobotManip (joint) | 37.6 | 4.1 | 1.8 | 14.5 |
| Qwen-RobotManip (eef) | 42.9 | 22.8 | 5.9 | 23.9 |
The performance gradient (ARX > UR5 > Franka) tracks visual/kinematic similarity to the AgileX training platform. Camera-frame EEF vs own-joint: 5.6Γ on UR5. This is the cleanest published validation of camera-frame action spaces for zero-shot embodiment transfer.
Fine-tuned on 22.9 h of teleop. In-domain, 7 tasks Γ 5 trials: 88.6% vs Ο0.5 42.9% / StarVLA 20.0% β 5/5 on five tasks including block-in-drawer-compartment and three-block-stacking where Ο0.5 scores 0/5; only yellow-disc-insertion (contact-rich precision insertion) stays hard (2/5). OOD, 4 tasks Γ 10 trials: 87.5% vs 37.5% / 0.0% β perfect 10/10 on cluttered-scene target grounding and left-right relational stacking (Ο0.5: 1/10 on the latter), 9/10 under disco-light illumination.
Few-shot (130 demos total across 5 tasks): best on 4/5 tasks; long-horizon Put Blocks 37.5% sub-step avg vs Ο0.5 25.0%; Unscrew Cap full completion 3/10 vs 1/10; Insert Screw defeats everyone (0/10 insertions). Cross-embodiment skill transfer β joint fine-tune on 6K CobotMagic + 130 ARX demos, evaluate on 4 ARX tasks with zero ARX demonstrations of those tasks: 55.0% vs 7.5% (w/o UnifiedSpace) and 12.5% (w/o UnifiedEEF) β the alignment stack is what makes skills compose across bodies.
Submitted as Lira_generalist; joint control; 30 tasks / 4 embodiments (ARX5, ALOHA, UR5, Franka), one policy per embodiment.
| Method | SR | Process score |
|---|---|---|
| Qwen-RobotManip | 45 | 59.83 |
| DM0_generalist | 37 | 48.43 |
| Ο0.5_generalist | 17.67 | 31.27 |
| GR00T-MULTI | 15.33 | 32.29 |
| Ο0_generalist | 9 | 20.22 |
Three analysis threads: bimanual coordination (40% avg over the 8 bimanual ALOHA tasks vs Ο0.5's 21.2% β attributed to the dual-arm-heavy corpus and the H2R pipeline's bimanual synthesis; only model non-zero on pour-fries-into-plate at 30%); pick-and-place across embodiments (63.3% over 12 tasks vs DM0's 48.3%); and emergent retry behavior β spontaneous re-attempts after failed grasps/placements, observed across picking, pouring, folding, wiping, sweeping, hypothesized to come from imperfect-then-corrected segments in the diverse pretraining data. On 6 hard long-horizon tasks, prior SOTA averages 5% vs Qwen-RobotManip's 36.7%.
Three action-space designs Γ nested data subsets (1/5/10/25/50/100%), validation MSE on a held-out OOD set (15 embodiment types, 154 unseen tasks):
| Variant | Representation | Scaling behavior |
|---|---|---|
| w/o UnifiedSpace | Raw per-embodiment fields, zero-padded to 80 dims, no semantic alignment | Unstable, high MSE, no consistent improvement with data |
| w/o UnifiedEEF | Canonical 80-dim slots; EEF deltas as axis-angle relative to initial pose | Log-linear MSE scaling β |
| Ours | + camera-frame delta EEF | Log-linear β with lowest EEF-prediction MSE |
Downstream (RoboTwin-C2R after fine-tuning each variant): on Hard, full alignment scales steadily 1% β 100% (reaching 50.2 joint / 56.6 eef) and beats both ablations at every data fraction; on Easy (in-distribution), no variant shows an upward data trend β the in-domain-benchmarks-can't-see-pretraining point, reproduced inside the ablation. Also: only the full model performs better in EEF mode than joint mode (72.5 vs 68.1 Easy); both ablations show the reverse.
| Config | Denoise steps | Easy | Hard | Avg |
|---|---|---|---|---|
| No prompt | 4 | 71.2 | 54.2 | 62.7 |
| Soft prompt | 4 | 70.2 | 52.1 | 61.2 |
| Language tag + FPS | 4 | 71.7 | 55.1 | 63.4 |
| Structured prompt | 4 | 73.4 | 58.3 | 65.9 |
| + Context | 4 | 72.1 | 54.4 | 63.3 |
| + Context | 10 | 80.1 | 61.6 | 70.9 |
| + Context | 20 | 79.8 | 62.1 | 71.0 |
The context mechanism's +5.0 over the structured prompt dwarfs every prompt-design delta β but only materializes with a 10-step denoising budget (at 4 steps the more complex action distribution produces jitter and the gain vanishes). 20 steps adds nothing.
Robot-only β +raw ego β +H2R at fixed 7:3 ratio: RoboTwin-C2R Hard 54.7 β 55.0 β 58.7; LIBERO-Plus total 87.1 β 88.4 β 89.0, with the largest per-dimension gain on Camera (72.8 β 80.0, +7.2) β ego viewpoint diversity converted into robot-frame robustness. Monotonic progression: raw ego helps via visual diversity; the H2R rendering unlocks the rest via action + visual alignment.
| LIBERO | LIBERO-Plus | RT-C2R Easy | RT-C2R Hard | RT-IF | |
|---|---|---|---|---|---|
| Full (VL in pretrain, none in post-train) | 99.1 | 90.1 | 73.2 | 62.6 | 71.6 |
| β VL in pretraining | 98.2 | 88.9 | 66.5 | 54.4 (β8.2) | 64.6 (β7.0) |
| + VL also in post-training | 98.6 | 91.4 | 74.0 | 62.5 | 73.1 |
VL co-training barely matters on in-distribution LIBERO (β0.9) but is worth 7β8 points on the hard OOD/instruction benchmarks β the co-training-as-generalization-infrastructure reading, consistent with LBM but measured on harder OOD axes. Adding VL to post-training specifically lifts language axes (LIBERO-Plus language perturbation 86.9 β 93.9; RT-IF Pick-Diverse 76 β 81).
| Variant | Total |
|---|---|
| Layer-wise self-attention (per-layer VLM feature fusion) | 86.4 |
| Last-layer self-attention (concatenation β the Qwen-VLA pattern) | 87.0 |
| Last-layer cross-attention (+ learned query tokens) | 87.5 |
Cross-attention wins and is the cheapest (no per-layer fusion, no full VLM token set stored in the DiT). Learned query tokens (register-style, outputs discarded) accompany the state/action tokens in cross-attending to VLM states.
| Axis | Qwen-RobotManip (Jun 2026) | Qwen-VLA (May 2026) |
|---|---|---|
| Backbone | Qwen3.5-4B (D=2560), fully unfrozen | Qwen3.5-4B, frozen at T2A then unfrozen |
| Action expert | DiT, 10 blocks, D=768, 12 heads (small) | DiT, 16 blocks, 1.15B (large) |
| VLMβexpert wiring | Cross-attention to last-layer states, alternating vision/language by block parity + learned queries | Concatenation + joint self-attention |
| Action space | 80-dim canonical; camera-frame delta EEF (CaPE-conditioned) or abs joint | HΓK fixed tensor; native per-dataset conventions + quantile norm |
| Cross-embodiment interface | Canonical slots + structured prompt + in-context history as implicit embodiment ID | Embodiment text prompt only |
| Training | Single-phase dual-stream co-training (9:1), K_repeat=8 | Four stages: T2A β CPT β SFT β PPO RL |
| RL | None | PPO with ODEβSDE flow log-prob |
| Data | ~38,100 h, zero proprietary, 65% H2R synthetic | >10,000 h public + >1,000 h in-house + language-only synthetic |
| Proprioceptive state | Explicit 80-dim state via MLP into DiT | Ablated to β€+1.3 pp; omitted |
| Camera calibration | Required for the main action mode (flag-switched fallback) | Not required |
| Evaluation identity | OOD-first; 2 new benchmarks; RoboChallenge #1 | Generalist-vs-specialist head-to-heads; DOMINO zero-shot |
| Scope | Manipulation only | Manipulation + navigation + AD-VQA |
Direct points of tension worth tracking:
- Wiring. RobotManip's Table 19 finds cross-attention > concatenation on LIBERO-Plus; Qwen-VLA shipped concatenation. The margin is small (0.5 pp) and the setting differs, but within one organization the two flagship VLAs disagree on the single most-debated design axis in Review-VLM-Action-Connection.
- State. Qwen-VLA found explicit proprioception nearly worthless; RobotManip makes an 80-dim state vector a first-class input. Not directly contradictory (different action spaces β camera-frame deltas may need state anchoring), but unreconciled.
- Recipe complexity. Qwen-VLA bets on stage curriculum + RL; RobotManip bets on representation alignment + data scale with the simplest possible one-phase recipe. RobotManip's LIBERO-Plus 91.4 vs the numbers Qwen-VLA reported on different benchmarks are not directly comparable β neither paper evaluates the other.
| Axis | Qwen-RobotManip | Ο0.5 | GR00T N1.x | TRI LBM |
|---|---|---|---|---|
| Expert wiring | Cross-attn DiT (small, D=768) | Same-stack expert w/ prefix attention | Cross-attn DiT | adaLN-conditioned 8-layer head |
| Action space | Camera-frame delta EEF (+ abs joint mode) | Robot-frame continuous | Embodiment-projected continuous | Robot-frame continuous |
| Camera geometry in the policy | CaPE in attention (q,k,v,out) + intrinsics tokens | No | No | No |
| Data thesis | Open + H2R synthesis (65% synthetic) | In-house-heavy + web co-training | Data pyramid incl. ~20k h ego (EgoScale) | 523 h target + OXE + ego |
| Anti-forgetting | 9:1 dual-stream co-training | Co-training + KI | Partial-layer tuning | Frozen backbone + VL stream |
| In-context adaptation | Yes (execution-history conditioning) | No | No | No |
| New benchmarks shipped | RoboTwin-IF, RoboTwin-XE | No | No | No |
Genuinely novel: (a) the camera-frame delta EEF + CaPE-in-the-action-expert combination at foundation scale, and the demonstration that it creates the data scaling law; (b) in-context policy adaptation with stochastic context sampling as an anti-shortcut mechanism, interpreted as implicit embodiment identification; (c) the 15-platform H2R synthesis pipeline at 24,808 h β an order of magnitude beyond Phantom/Masquerade-scale predecessors; (d) the RoboTwin-IF / RoboTwin-XE benchmarks and the "VLA-to-VA degradation" framing; (e) the curation pipeline's quantified findings (81% of RoboMIND UR data failing causality checks is a service to the field).
Recombination: the decoupled VLM + flow-matching DiT (Ο0/GR00T family), VL co-training as anti-forgetting (LBM, Ο0.5), masked canonical action vectors (Qwen-VLA, GR00T), ECoT supervision (Zawalski et al.), human-video-to-robot editing (Phantom, Masquerade), RTC deployment (Black et al.).
- The strongest open-data result in VLA to date. Every prior model at this performance tier (Ο0.5/0.6, GR00T, Qwen-VLA) leans on proprietary teleop. If alignment + synthesis really substitutes for in-house collection, the economics of manipulation foundation models change.
- Camera-frame actions graduate from technique to thesis. Prior appearances (Chen et al. 2025a; Zhang et al. 2026b) were single-robot studies; this is the first foundation-scale validation with a scaling-law argument attached. Expect the Action Space: EEF vs Joint debate to acquire a third axis: which frame.
- Evaluation-methodology influence may outlast the model. The from-scratch-matches-pretrained result plus RoboTwin-IF/XE give reviewers concrete tools to demand OOD evidence. RoboTwin-IF in particular measures a failure mode (language-conditioning collapse) no prior benchmark isolates β and GR00T-N1.7's 16.6% on it shows the failure is real in shipping models.
- A weaker claim than it appears in one respect: most OOD evaluation is still simulation (LIBERO-Plus, RoboTwin variants, RoboCasa365, EBench are all sim), which the authors themselves concede; the real-world OOD suite is 4 tasks on one platform.
- H2R synthesis quality bounds. Retargeting approximations and inpainting artifacts introduce distributional gaps that cap the effective quality of the 24,808 synthesized hours.
- OOD evaluation is still predominantly simulation-based; broader real-world deployment-condition coverage is needed.
- Fixed action-chunk length and inference latency constrain reactive sub-second control.
- No weights, contradicting the paper's own text. Β§6.4 says "we release both the context-conditioned and context-free variants," but the GitHub README (as of this review, Jul 2026) states there is no plan to release model weights for Qwen-RobotManip or Qwen-RobotNav. Every headline number is therefore currently unreproducible outside Alibaba, and the two new benchmarks' baselines cannot include the model that motivated them.
- Camera calibration as a hidden deployment tax. The flagship camera-frame mode requires calibrated intrinsics + extrinsics at inference. The fallback (base-relative mode via the flag embedding) exists, but no evaluation isolates how much performance survives without calibration β a key question for in-the-wild deployment, where DROID-style setups often have rough calibration at best.
- Reproducibility gap on training details. No optimizer, learning rates, batch sizes, step counts, GPU type/hours for pretraining. Same criticism applied to Qwen-VLA; unchanged here.
- No latency numbers despite a latency-sensitive design. Remote WiFi inference + RTC + (for the Context variant) 10 denoising steps and a longer VLM sequence; no ms-per-chunk is reported. The Context variant's start-of-episode hesitation is honestly disclosed, but its runtime cost is not.
- Context variant is not uniformly better. It loses to the base model on RoboCasa365 (33.8 vs 35.9), EBench (43.6 vs 45.6), and ties on RoboTwin-IF β the gains concentrate on perturbation-robustness benchmarks (LIBERO-Plus, RT-C2R). The paper does not analyze why history hurts long-horizon composite tasks; plausibly the stochastic-context training trades temporal-progress information away.
- No comparison against its own sibling. Qwen-VLA and Qwen-RobotManip share a backbone and a benchmark-capable evaluation stack, yet neither paper cites or evaluates the other. A controlled comparison (concatenation vs cross-attention at matched scale; T2A+RL vs alignment-first) is exactly the experiment the community needs and only this team can run.
- Ο0.6/Ο0.7 absent from baselines. Comparisons stop at Ο0.5 (Apr 2025) and GR00T-N1.5/1.6/1.7. Understandable for closed models, but the "substantially outperforms prior SOTA" claim is dated against a 14-month-old competitor.
- H2R data is 65% of the corpus but only ablated at small scale. Tables 16β17 show +1.9β4.0 pp from H2R at a reduced-scale 7:3 mixture. Whether 24,808 h of synthetic data provides value proportional to its two-thirds share of the corpus β versus, say, a 5,000 h subset β is not measured; no H2R-fraction sweep exists.
-
RoboChallenge is a leaderboard, not a paper protocol. The 1st-place claim (as
Lira_generalist) is externally verifiable, but per-task trial counts and evaluation dates are controlled by the challenge, and the "20% relative improvement" headline is over the next submission at a point in time. - The scaling-law evidence is MSE + one downstream benchmark. The log-linear claim rests on validation MSE (a proxy with a known weak correlation to task success in the literature) plus RoboTwin-C2R fine-tuning. Success-rate scaling curves on a second, non-RoboTwin domain would make the "alignment unlocks scale" thesis much harder to attack.
- arXiv: https://arxiv.org/abs/2606.17846 Β· PDF: https://arxiv.org/pdf/2606.17846
- Blog: https://qwen.ai/blog?id=qwen-robotmanip
- Code (docs only, no weights): https://github.com/QwenLM/Qwen-RobotManip
-
RoboChallenge leaderboard: https://robochallenge.cn/home (entry:
Lira_generalist) - HuggingFace paper page: https://huggingface.co/papers/2606.17846
- Qwen Team's VLA Program β cross-paper review placing this paper in the VLM4VLA β Qwen-VLA β Qwen-Robot Suite arc
- Qwen-RobotNav Β· Qwen-RobotWorld β the Qwen-Robot Suite siblings (RobotWorld evaluates zero-shot on this paper's RoboTwin-IF benchmark)
- Qwen-VLA β the sibling model; see Β§9.1 for the head-to-head
- Action Space: EEF vs Joint β the action-space debate this paper extends with the camera-frame axis
- Cross-Embodiment β one policy, many bodies
- LBM Co-training Study β the VL co-training evidence base
- VLA Architectures Β· VLMβAction Connection
- GR00T series β the other cross-attention DiT lineage
- Ο series evolution Β· Ο0.6 Β· Ο0.7
- Knowledge Insulation β the architectural alternative to dual-stream anti-forgetting
- LIBERO-Plus Β· RoboTwin 2.0 β the benchmarks this paper builds on
- X-VLA Β· Cosmos-Policy β baselines
- DM0 β RoboChallenge runner-up
- EgoDex Β· EgoScale β the egocentric-data scaling context
β Back to Home