Review Qwen RobotManip - Heungwoo/research GitHub Wiki

In-Depth Review β€” Qwen-RobotManip: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

Paper: Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models Authors: Qwen Team β€” core contributors Haoqi Yuan*, Zhixuan Liang*, Anzhe Chen*, Ye Wang*, Haoyang Li*, Pei Lin*, Yiyang Huang*, Zixing Lei*, Tong Zhang*, Jiazhao Zhang, Jie Zhang, Jingyang Fan, Gengze Zhou, Qihang Peng, Chenxu Lv, Xiaoyue Chen, An Yang, Fei Huang, Junyang Lin, Dayiheng Liu, Jingren Zhou, Chenfei Wu† (corresponding), Xiong-Hui Chen‑† (project lead + corresponding), + 14 contributors Affiliation: Qwen Team, Alibaba Group arXiv: 2606.17846 Β· v1 Jun 16, 2026 Β· v2 Jun 17, 2026 (44 pages, cs.RO, CC BY 4.0) Code: github.com/QwenLM/Qwen-RobotManip Β· Blog: qwen.ai/blog?id=qwen-robotmanip Weights: the GitHub README states there is no plan to release model weights (for Qwen-RobotManip or Qwen-RobotNav) as of this review β€” see Β§11.2

This is the long-form companion to the per-paper summary. Companion reviews to read alongside: Qwen-VLA (the Qwen team's first VLA, May 2026 β€” the two papers are siblings and disagree architecturally) Β· Action Space: EEF vs Joint Β· Cross-Embodiment Β· LBM Co-training Β· VLA Architectures Β· VLM↔Action Connection Β· LIBERO-Plus.


1. TL;DR

  1. The Qwen team's second VLA in two months β€” with a different thesis. Where Qwen-VLA (May 2026) was "one generalist, many benchmarks, 4-stage recipe + RL", Qwen-RobotManip (June 2026) is an alignment-first scaling manifesto: a canonical 80-dim state-action representation, a camera-frame delta EEF action space, and in-context policy adaptation are presented as prerequisites that make heterogeneous data scale at all. The ablation showing that unaligned representations produce no scaling law while aligned ones scale log-linearly from 1%β†’100% data is the paper's core claim.
  2. ~38,100 hours of pretraining data with zero proprietary collection. ~11,420 h open-source robot data (9 datasets) + 1,933 h egocentric human video (EgoDex/VITRA/EgoVerse) + 24,808 h human-to-robot synthesized demonstrations β€” each human clip re-rendered onto 15 bimanual robot platforms via a SAM3 β†’ ProPainter β†’ base-search + MuJoCo-IK β†’ depth-composite pipeline. Two-thirds of the corpus is synthetic H2R data.
  3. An evaluation manifesto: standard benchmarks are broken. The paper shows from-scratch models match pretrained ones on LIBERO/RoboTwin (in-distribution), then argues OOD transfer is the only faithful measure β€” and introduces two new benchmarks: RoboTwin-IF (instruction following under held-out language templates) and RoboTwin-XE (zero-shot transfer to unseen robot morphologies).
  4. Headline OOD numbers vs Ο€0.5: 91.4% LIBERO-Plus (vs 84.4), 69.4% RoboTwin-Clean2Rand Hard (vs 47.9), 45.6% EBench (vs 27.1), 35.9% RoboCasa365 (vs 16.9), 72.2% RoboTwin-IF (vs 49.6), 23.9% zero-shot cross-embodiment on RoboTwin-XE (3.2Γ— Ο€0.5's 7.5). Real-world CobotMagic ALOHA: 88.6% ID / 87.5% OOD vs Ο€0.5's 42.9 / 37.5. Ranked 1st on RoboChallenge Table30-v1 generalist track (45% SR / 59.83 process score vs DM0's 37 / 48.43).
  5. Architecturally it contradicts Qwen-VLA. RobotManip uses a small 10-block cross-attention DiT (D=768) reading last-layer VLM hidden states β€” and its own ablation finds last-layer cross-attention beats the concatenation + joint self-attention design that Qwen-VLA shipped four weeks earlier. No RL stage, no T2A warm-start; a single dual-stream co-training pretraining phase instead.

2. Why this paper matters in the 2026 landscape

2.1 One team, two VLAs, two theses

By June 2026 the Qwen team has published two distinct manipulation foundation models:

Qwen-VLA (arXiv 2605.30280, May 2026) Qwen-RobotManip (arXiv 2606.17846, Jun 2026, this paper)
Central thesis Compression: language→action decompression prior (T2A) + 4-stage recipe ending in RL Alignment unlocks scale: unified representation/motion/behavior alignment is the prerequisite for data scaling
Data posture >1,000 h in-house teleop + public + synthetic Open-source + synthesized only, zero proprietary collection (~38,100 h)
Scope Manipulation + navigation + AD-VQA generalist Manipulation only (a sibling Qwen-RobotNav is referenced in the repo)
Evaluation posture Standard benchmark head-to-heads as a single generalist OOD-first manifesto + two new benchmarks (RoboTwin-IF, RoboTwin-XE)

Author overlap is partial (Zhixuan Liang, Pei Lin, Jie Zhang, Jinhui Ye, Sicheng Xie, Shuai Bai, Junyang Lin, Dayiheng Liu, Jingren Zhou appear on both), and the corresponding authors differ (Shuai Bai for Qwen-VLA; Chenfei Wu / Xiong-Hui Chen here). Neither paper cites the other. The most reasonable reading is two parallel efforts inside Tongyi/Qwen, and this one β€” bundled with Qwen-RobotNav under a "Qwen-Robot" suite umbrella β€” is the manipulation-specialized line.

2.2 Three claims with field-wide implications

  1. Alignment is a scaling prerequisite, not an engineering nicety. The Β§6.4 controlled data-scaling experiment (nested 1/5/10/25/50/100% subsets, held-out validation across 15 embodiment types Γ— 154 tasks) is the first published demonstration that the same cross-embodiment corpus produces a clean log-linear MSE scaling law under an aligned action representation and an erratic, non-improving curve without it. This reframes debates about "does cross-embodiment data help?" (Cross-Embodiment) β€” the answer may depend on representation alignment more than on data volume.
  2. The data barrier may be lower than assumed. 38,100 hours with no proprietary teleop β€” two-thirds of it synthesized from human video β€” directly challenges the PI-style in-house-data-first posture. If the OOD results hold up, the moat shifts from data collection to synthesis + curation infrastructure.
  3. In-distribution benchmarks systematically overestimate. Figure 4's from-scratch-matches-pretrained result on LIBERO/RoboTwin (e.g., Qwen-RobotManip-scratch 98.2 LIBERO vs Ο€0.5's 97.6) is an indictment of most 2025–2026 VLA evaluation practice, and the paper backs it with the observation that none of the action-representation variants show data-scaling gains on the in-distribution Easy settings while the OOD Hard settings separate them cleanly.

2.3 The camera-frame action-space bet

Qwen-RobotManip is the largest-scale commitment yet to camera-frame delta EEF actions (following Chen et al. 2025's embodiment-equivariance line). This connects to the Action Space: EEF vs Joint debate but goes a step further than "EEF beats joint": the claim is that which frame the EEF delta lives in decides whether cross-embodiment transfer works at all. RoboTwin-XE gives the cleanest evidence: joint-space zero-shot transfer to UR5 is 4.1%, camera-frame EEF is 22.8% (5.6Γ—).


3. Architecture

flowchart TB
  subgraph OBS["Multi-view observations + prompt"]
    V["Camera views (left / front / right / wrist)<br/>current + optional history frames"]
    P["Structured embodiment prompt<br/>embodiment Β· instruction Β· speed Β· fps Β· camera view direction"]
    ECOT["(co-training) ECoT / VL QA batches"]
  end

  subgraph CTX["In-context policy adaptation (Context variant)"]
    C1["H history chunks (o_h, s_h, a_h)"]
    C2["MLP_s / MLP_a encoders + temporal & slot embeddings"]
  end

  OBS --> VLM["Qwen3.5-4B backbone<br/>natively multimodal, early fusion<br/>ViT w/ dynamic-resolution spatial merging<br/>D_vlm = 2560"]
  CTX -- "unified mode: context tokens appended to VLM input" --> VLM

  VLM -- "last-layer hidden states (vision + language)" --> XATT

  subgraph DIT["Flow-matching action expert β€” DiT, N = 10 blocks, D_act = 768, 12 heads"]
    XATT["Cross-attention (alternating):<br/>even blocks β†’ visual tokens<br/>odd blocks β†’ language tokens"]
    SA["Self-attention over state + noisy-action tokens<br/>CaPE (32/64 head dims) + RoPE (32/64)"]
    FF["SwiGLU FFN Β· AdaLN(-Zero RMSNorm) conditioning:<br/>denoise timestep Β· EEF-type codebook Β· camera-params flag"]
  end

  STATE["Proprioceptive state (80-dim canonical)<br/>2-layer MLP, prepended"] --> DIT
  NOISE["Noisy action chunk x_t"] --> DIT
  DIT --> VEL["Velocity field v = a βˆ’ Ξ΅"]
  VEL --> EULER["4 Euler steps (10 for Context variant)"]
  EULER --> ACT["Action chunk β€” 80-dim canonical vector<br/>camera-frame delta EEF (or abs joint)"]

  classDef vlm fill:#bbdefb,stroke:#1565c0,color:#000
  classDef act fill:#c8e6c9,stroke:#2e7d32,color:#000
  classDef inp fill:#fff9c4,stroke:#f57f17,color:#000
  class VLM vlm
  class DIT,XATT,SA,FF,VEL,EULER,ACT act
  class OBS,CTX,STATE,NOISE inp
Loading

3.1 Backbone

  • Qwen3.5-4B (Team, 2026) β€” same backbone family as Qwen-VLA. Natively multimodal with early vision-language fusion: ViT visual tokens with dynamic-resolution spatial merging interleaved directly into the text stream. Last-layer hidden dimension D_vlm = 2560.
  • The backbone is trained end-to-end β€” flow-matching gradients flow into the VLM (no freezing, no Knowledge Insulation-style gradient routing). Forgetting is managed by data (dual-stream VL co-training, Β§5.1) rather than by architecture.

3.2 Action expert β€” small cross-attention DiT

  • N = 10 transformer blocks, D_act = 768, 12 attention heads β€” dramatically smaller than Qwen-VLA's 1.15B/16-block expert (parameter count is not stated, but at these dimensions it is roughly an order of magnitude smaller).
  • Each block: self-attention over the concatenated state + noisy-action token sequence, then cross-attention to VLM last-layer hidden states, then a SwiGLU FFN. Cross-attention alternates by block parity: even-indexed blocks attend to visual tokens, odd-indexed blocks to language tokens β€” letting the expert ground actions in observations and instructions at separate processing stages.
  • Proprioceptive state enters via a 2-layer MLP, prepended to the noisy action tokens.
  • Conditioning via adaptive layer normalization (the figure labels it AdaZeroRMSNorm): denoising timestep + two additional embeddings (Β§3.5).
  • Flow matching: t ~ Beta(1, 1.5), interpolant x_t = (1βˆ’t)Ξ΅ + tΒ·a, velocity target v = a βˆ’ Ξ΅. Inference: 4 Euler steps (the Context variant needs 10 β€” see Β§8.2).

In the Review-VLA-Architecture taxonomy this is Category B (separate flow-matching expert) with GR00T-style cross-attention into VLM hidden states β€” not the concatenation + joint-self-attention wiring Qwen-VLA chose. Notably, Β§6.4's architecture ablation directly compares three wirings (layer-wise self-attention fusion; last-layer self-attention concatenation β€” i.e., the Qwen-VLA pattern; last-layer cross-attention with learned query tokens) and finds cross-attention wins on LIBERO-Plus (87.5 vs 87.0 vs 86.4) at the lowest compute cost. One Qwen paper's ablation quietly contradicts the other Qwen paper's design choice.

3.3 The 80-dimensional canonical state-action representation

The cross-embodiment interface, and the first of the three "alignment" pillars:

Segment Dims Content
Per-arm block Γ— 2 29 each Joint positions (7) + EEF pose (3 pos + 6D rotation = 9) + gripper (1) + dexterous-hand joints (12)
Reserved (shared) 22 Future DoF, e.g. mobile-base velocity
Total 80
  • States are absolute; actions are absolute for joints and relative deltas for EEF (rotation deltas as 3D rotation vectors; state rotations as 6D continuous representation).
  • Each embodiment populates only its own slots; a per-dimension binary mask excludes zero-padded entries from the loss. The mask is AND-composed from three sources: the embodiment slot mask, a step-validity mask from the curation pipeline (once a step is invalid, all subsequent steps mask out to preserve causal consistency), and a per-hand validity mask for ego-human data that zeroes an arm slot from the moment the hand leaves the camera view.
  • The DiT actually processes N_ee ∈ {1,2} 40-dim per-end-effector tokens (29 active + 11 reserved) jointly via self-attention.

Same family as Qwen-VLA's zero-padded fixed tensor and GR00T's per-embodiment encoders, but with semantically fixed slot positions β€” and Β§6.4 shows the semantics matter: naive concatenate-and-pad ("w/o UnifiedSpace") destroys the scaling law.

3.4 Camera-frame delta EEF + CaPE β€” the motion-alignment pillar

The paper's most distinctive architectural contribution:

  • Camera-frame delta pose (Eq. 5): the relative EEF rotation is conjugated into the reference camera frame, and the EEF displacement is projected into camera coordinates. Property: actions that look similar in the image are numerically proximate, regardless of robot base placement or morphology. Requires calibrated intrinsics + extrinsics at train and test time. The more compact full-camera-frame alternative (Eq. 6) was rejected as more sensitive to calibration error and long-tail translation coupling.
  • Camera-aware positional encoding (CaPE): camera extrinsics are injected into the DiT's cross-attention as a rotational positional encoding occupying 32 of each 64-dim attention head (the other 32 are RoPE for temporal indexing). Following GTA/PRoPE practice, CaPE is applied to queries, keys, values, and attention outputs. Because it is rotational, the global world origin cancels in dot-product attention β€” only relative camera↔token poses survive. Intrinsics enter as a learned linear projection of normalized image-plane coordinates added per visual token.
  • Reference-camera selection is randomized during training (any view for single-arm; shared head-camera vs per-wrist-camera strategies for dual-arm), so the policy learns to denoise in whatever frame the CaPE indicates.
  • Fallback mode: an auxiliary binary flag embedding (camera parameters available or not) switches the action space between camera-frame delta and robot-base relative mode, so uncalibrated datasets remain usable.

3.5 Structured embodiment prompt

embodiment: robot_aloha
instruction: Take the toy off the table and put it on the mat.
speed: 1000          # episode length, binned at 500 steps
fps: 30
camera view direction: arm side

Fields: platform tag, instruction, speed (episode-length bin), FPS, and camera-view direction (arm side / opposite side). The embodiment / speed / fps fields are randomly dropped with p = 0.15 during training for robustness to missing metadata. The Β§6.4 ablation ladder: soft prompt hurts (βˆ’1.5 avg vs no prompt), plain language tag helps slightly (+0.7), the structured prompt helps +2.2–3.2 points β€” i.e., the temporal metadata (FPS/speed), not the identity tag, carries most of the value.

3.6 In-context policy adaptation β€” the behavior-alignment pillar

The "-Context" variant conditions on a window of H recent execution chunks, each an (observation, state, action-chunk) triplet:

  • Injection: historical frames prepend to the current frame through the VLM visual pathway (with an image-count annotation in the prompt); states and action chunks go through lightweight MLP encoders with temporal + slot embeddings, and the resulting context tokens are appended to the VLM input sequence ("unified mode" β€” chosen over injecting into the DiT, "dual mode", for richer cross-modal integration).
  • Stochastic context sampling is the critical trick: always feeding the most recent chunks lets the model win training loss by copying the last action chunk (a recency shortcut that collapses the mechanism). Training instead samples the context window from a random position in the episode, forcing the model to extract the episode's behavioral profile (velocity patterns, grasping style, kinematic signature). At deployment, a rolling window of the most recent H chunks is used.
  • The paper's interpretation: context works as an implicit embodiment identifier β€” intra-episode kinematics tell the policy what body it is driving, complementing the static prompt. Supporting evidence: the biggest per-dimension LIBERO-Plus gain from Context is Robot-state perturbation (75.5 β†’ 83.9) and Camera (87.2 β†’ 89.9).
  • Cost: context makes the action distribution more complex; 4 denoising steps become insufficient (jittery motion, performance back at baseline) and 10 steps are required to unlock the benefit (Β§8.2). Also, at episode start the context is all zero-padding and the policy tends to hesitate before moving β€” the stated reason both Context and context-free variants exist.

4. Data: the ~38,100-hour corpus

4.1 Composition (Table 1 of the paper)

Data type Embodiment Sources Setting Hours
Robot β€” single-arm Single-arm OXE (Fractal/Bridge/BC-Z), RoboMIND, DROID, RH20T, … Tabletop 3,808
Robot β€” dual-arm Dual-arm AgiBotWorld-Beta, RoboCOIN, RDT, … Tabletop 6,744
Robot β€” mobile & humanoid Mobile/humanoid InternData-A1, Galaxea Open-World Tabletop & indoor 868
Human Human hands EgoDex, VITRA, EgoVerse Tabletop & in-the-wild 1,933
Human-to-Robot 15 dual-arm platforms Synthesized from the human data Tabletop & in-the-wild 24,808

Robot subtotal ~11,420 h across nine open-source datasets: OXE ~600 h (Fractal + Bridge + BC-Z only), AgiBotWorld-Beta ~2,400 h (gripper subset, ~200 task types), RoboMIND 1.0 + 2.0 ~1,400 h, Galaxea Open-World ~500 h, RoboCOIN ~430 h (10 embodiment types), DROID ~500 h (95k trajectories), RH20T ~1,100 h (contact-rich, 4 embodiments), RDT-1B 29 h, InternData-A1 >3,600 h (high-fidelity sim). Egocentric: EgoDex 732 h used (of 829), VITRA-processed Ego4D + EPIC-KITCHENS 247 h, EgoVerse industry portion 954 h; all hand poses unified to MANO parameters + 21 keypoints.

4.2 Human-to-robot (H2R) synthesis β€” the scaling engine

Each egocentric human clip is converted into robot demonstrations on 15 bimanual configurations (two identical arms from: Panda, UR5e, ARX-L5, xArm7, Sawyer, Kinova Gen3, IIWA, Jaco, FR3, UR10e, ViperX, WidowX, Piper, YAM, AgileX ALOHA), a 1,933 h β†’ 24,808 h multiplication. Pipeline (inspired by Phantom/Masquerade):

  1. Action retargeting. Virtual finger k_vf = 0.7Β·k_index + 0.3Β·k_middle; EEF position = midpoint of thumb-tip and virtual finger; gripper width = their distance; gripper orientation built as a right-handed frame from the jaw-line axis and wrist-to-fingertip direction, with a handedness sign flip so both hands map to one gripper frame. Savitzky–Golay filtering on positions/widths, Gaussian-weighted SLERP on orientations.
  2. Visual removal. SAM3 text-prompted arm masks β†’ ProPainter optical-flow-guided inpainting β†’ clean background plates.
  3. Base placement search. Egocentric trajectories are embodiment-free (no robot base exists), so base pose is optimized: grid search around the trajectory centroid maximizing the fraction of representative keyframes with feasible IK, independently per morphology.
  4. Rendering + compositing. MuJoCo (mink) IK tracks the smoothed trajectory; the rendered robot is composited onto the inpainted plate using an occlusion mask from Depth Anything v3 metric depth vs rendered robot depth.
  5. Speed alignment. Human motion is faster than teleop, so ego sources are frame-subsampled during training: EgoDex to 60% (~1.7Γ— slower), EgoVerse to 45% (~2.2Γ—), VITRA to 25% (~4Γ—).

4.3 Curation β€” five signal-quality stages + three cross-modal checks

Applied to all pretraining data (and deliberately disabled for SFT, Β§5.2):

  1. Sudden-change detection β€” median + Savitzky–Golay trend, flag on residual ∧ (acceleration ∨ jerk) thresholds; per-dataset thresholds; frame-level removal up to full-episode discard (e.g., InternData-A1 collision episodes dropped entirely).
  2. State-action trend alignment β€” cross-correlation lag estimation + directional-agreement metric on lag-aligned first differences; DA < ~0.6–0.7 β†’ episode excluded. This stage found 81% of RoboMIND UR-type episodes faulty β€” a remarkable public data-quality datapoint.
  3. Extreme-value filtering β€” protects the [q01, q99] β†’ [βˆ’1,1] quantile normalization; gripper dims exempt (bimodal).
  4. FK consistency β€” Pinocchio FK vs logged EEF poses; primarily corrects (TCP offsets, shoulder-frame β†’ world-frame) rather than filters; revealed same-robot-different-convention conflicts across datasets β€” the empirical motivation for the canonical representation.
  5. Base-frame / orientation alignment β€” per-dataset rotation corrections so +x is always robot-forward.

Cross-modal: instruction consistency (subtask segmentation β†’ structured reasoning-guided VLM judgment β†’ multi-VLM cross-model adjudication for flagged clips); video-state consistency (render URDF + joint states into the image, compare against a fine-tuned SAM3 robot mask by IoU; low overlap β†’ fix camera parameters or drop); video quality (black/corrupt/blur/static removal, explicitly preserving gripper-closure keyframes).

4.4 VL co-training mixture (~28M samples)

Six categories: (1) general visual understanding, (2) spatial perception/reasoning (2D/3D grounding, pointing, counting, depth/distance/viewpoint, manipulation feasibility), (3) OCR/documents, (4) multimodal specialized knowledge, (5) instruction following/multilingual/pure text, and (6) embodied-centric data, which is the novel part:

  • Embodied Chain-of-Thought (ECoT): three-part annotations (scene description β†’ task-progress judgment with explicit "Task complete / not yet complete" β†’ next atomic action from a 16-type taxonomy), synthesized by Qwen3.6-Plus in thinking mode using privileged context (episode memory summary, 6-frame future preview, coarse progress estimate) that is excluded from the training inputs β€” the model must produce the ECoT from current observation + instruction only.
  • Egocentric video understanding: fine-grained hand/arm-motion and object-state-change descriptions over 4-frame clips (1.5–3 s), static clips filtered.
  • 2D trajectory prediction: future EEF / hand trajectories as normalized image-coordinate sequences, low-motion samples filtered.

Sources include RoboPoint, RefSpatial, PixMo, CapsFusion, plus proprietary data.


5. Training methodology

5.1 Pretraining β€” single-phase dual-stream co-training

No multi-stage curriculum (contrast: Qwen-VLA's four stages, LBM's three). One phase, two mutually exclusive batch streams:

  • VLA stream: full manipulation corpus β†’ masked flow-matching loss (Eq. 11), a per-sample average over valid mask entries so every sample contributes equally regardless of active-dimension count.
  • VLM stream: the 28M VL mixture β†’ standard next-token CE.
  • Mixture ratio 9:1 (robot : VL); joint loss L = L_FM + λ·L_VLM with Ξ» = 0.1; separate learning rates for backbone vs expert.
  • K_repeat = 8: per VLA sample, the expert runs 8 independent (noise, timestep) draws against one VLM forward pass β€” amortizing the expensive backbone forward. A pragmatic efficiency trick not present in the Ο€/GR00T/LBM recipes.

Gradients from the flow-matching loss do update the backbone (unlike LBM's frozen PaliGemma and unlike KI's insulated CE-only backbone gradients). The 9:1 co-training stream is the sole anti-forgetting mechanism β€” and Β§8.4 shows it carries real weight (βˆ’8.2 pp on RoboTwin-C2R Hard without VL data).

5.2 Post-training β€” generalist SFT, and the "VLA-to-VA degradation" diagnosis

  • Standard protocol: per target domain, all demonstrations pooled into one generalist SFT set (no per-task specialists); flow-matching loss only (no VL CE); curation filters disabled (keep every valid demo); color-jitter augmentation; fewer GPUs/steps than pretraining.
  • The diagnosis: extended benchmark SFT causes what the paper names VLA-to-VA degradation β€” the policy stops conditioning on language and becomes a vision-action pattern matcher, because benchmark SFT data has concentrated layouts, shared train/test visual patterns, and weak language diversity. RoboTwin-IF exists precisely to measure this failure mode.
  • Mixed post-training (optional, Β§6.5.1): co-train SFT with 10% VL data + auxiliary VLA data (pretraining trajectories filtered by distributional proximity, aligned in EEF space; 75% of the VLA portion). Result on RoboTwin-IF: plain SFT decays with training steps (overfitting), +VL mitigates, +VL+VLA eliminates the decay entirely β€” performance keeps rising through 90k steps to 75.8%. Crucially, mixed post-training without the UnifiedEEF representation collapses to 0.0% β€” alignment is the prerequisite for absorbing mixed data even at SFT time.

5.3 Deployment

Remote-server inference over WiFi with Real-Time Chunking (RTC) (Black et al., 2026) generating the next chunk asynchronously during execution. No end-to-end latency numbers are published.

5.4 What is not in the recipe

No RL stage (vs Qwen-VLA's PPO, PI's RECAP). No action-expert warm-start (vs Qwen-VLA's T2A). No discrete action tokens anywhere β€” consistent with LBM's finding. No stated optimizer/LR/batch/GPU-hours for pretraining β€” the same reproducibility gap as Qwen-VLA.


6. The evaluation manifesto and the two new benchmarks

6.1 Standard benchmarks fail to measure pretraining (Β§6.1)

On LIBERO and RoboTwin (in-distribution), models without large-scale robot pretraining match or beat pretrained ones β€” StarVLA 98.0 / Qwen-RobotManip-scratch 98.2 on LIBERO vs Ο€0.5's 97.6. On OOD versions the ordering inverts hard: StarVLA collapses 85.7 β†’ 10.6 (RoboTwin Easy β†’ Clean2Rand), scratch collapses 71.6 β†’ 22.6, while pretrained Qwen-RobotManip retains 73.2 β†’ 62.6 (~86% retention vs Ο€0.5's ~66% and scratch's ~30%). The paper's practitioner argument: a deployer never has the benchmark's distribution β€” only OOD transfer efficiency matters.

6.2 RoboTwin-IF (new β€” instruction following)

Built on RoboTwin 2.0; models fine-tuned on RoboTwin-Clean only; evaluated with held-out unseen instruction templates + per-object description variants. Five suites: Pick-Diverse-Object (target grounding among 3 distractors from a 12-item pool), Place-Relative ("beside"/"on top of" spatial relations), Operate-Mic-Drawer (multi-step bimanual sequencing, optionally arm-specified), Operate-Stapler (verb discrimination β€” press vs move β€” with the placement pad always present as distractor), Operate-Tabletop (three-way verb-and-target discrimination in a multi-affordance scene).

6.3 RoboTwin-XE (new β€” zero-shot cross-embodiment)

Fine-tune on RoboTwin-Clean (AgileX only), evaluate under RoboTwin-Hard with the platform replaced by ARX-X5, UR5-WSG, or Franka Panda β€” initial EEF poses IK-aligned, camera extrinsics/scenes/seeds held identical, zero target-embodiment data. Tests both visual (arm appearance) and kinematic (DoF, joint arrangement, link lengths, workspace) generalization.


7. Comprehensive benchmark results

7.1 In-distribution (Table 3)

Method LIBERO RoboTwin-Easy RoboTwin-Hard
Ο€0 94.4 65.9 58.4
Ο€0.5 97.6 82.7 76.8
StarVLA 98.0 85.7 87.3
Abot-M0 98.6 86.1 85.1
Being-H0.7 99.2 90.2 89.6
Qwen-RobotManip-scratch 98.2 88.7 88.4
Qwen-RobotManip 99.1 93.4 92.5
Qwen-RobotManip-Context 99.2 93.7 94.0

State-of-the-art or tied β€” but the paper itself immediately discounts these numbers as unable to distinguish generalization from memorization.

7.2 LIBERO-Plus (Table 4) β€” seven perturbation dimensions

Method Camera Robot Language Light Background Noise Layout Total
Ο€0 13.8 6.0 58.8 85.0 81.4 79.0 68.9 53.6
Ο€0.5 78.4 73.6 80.8 96.2 94.1 89.0 84.5 84.4
StarVLA 52.5 49.8 88.5 95.7 95.7 73.0 76.9 74.1
Abot-M0 60.4 67.9 86.4 96.2 91.6 86.4 82.6 80.5
Cosmos-Policy 75.8 63.3 81.7 96.5 88.9 92.7 82.2 82.2
Being-H0.7 82.0 59.0 82.8 97.8 90.0 93.5 88.5 84.8
Qwen-RobotManip-scratch 70.4 44.9 88.1 95.8 95.5 84.4 79.1 78.3
Qwen-RobotManip 87.2 75.5 85.6 96.6 97.7 97.7 87.3 89.0
Qwen-RobotManip-Context 89.9 83.9 86.5 98.6 99.9 97.9 87.5 91.4

The structural reading the paper draws: Language/Light/Background robustness comes free from the VLM (scratch models match pretrained there); Camera and Robot-state robustness specifically require large-scale robot pretraining (+34.7 / +16.8 for pretrained-vs-scratch on Camera; scratch craters at 44.9 on Robot). Context adds an implicit kinematic prior on top (+8.4 Robot, +2.7 Camera).

7.3 RoboTwin-Clean2Rand (Table 5) β€” fine-tune on Clean, evaluate randomized

Method Easy Background Light Clutter Height Hard
StarVLA 58.1 27.1 50.9 24.2 48.4 10.6
GR00T-N1.7 43.6 40.4 41.9 27.1 39.0 20.7
Ο€0.5 73.1 67.0 69.2 57.9 67.6 47.9
Abot-M0 70.7 56.5 68.8 46.0 56.3 36.0
Qwen-RobotManip-scratch 71.6 60.6 70.7 24.6 63.6 22.6
Qwen-RobotManip (joint) 73.2 74.6 68.4 61.3 71.0 62.6
Qwen-RobotManip (eef) 74.0 75.8 70.1 59.8 69.4 60.8
Qwen-RobotManip-Context (joint) 84.7 82.4 84.2 75.4 79.5 69.4
Qwen-RobotManip-Context (eef) 85.0 82.4 84.7 66.8 82.9 64.0

Notable details: Background randomization helps the pretrained model (74.6 > 73.2 Easy β€” diverse pretraining scenes make random backgrounds more in-distribution than sterile white); Clutter is the axis where scratch pretraining-free models die (71.6 β†’ 24.6); Context adds +11.5 Easy / +6.8 Hard.

7.4 RoboCasa365 (Table 6) and EBench (Table 7)

Method RoboCasa365 Atomic Comp-Seen Comp-Unseen Total
Ο€0.5 39.6 7.1 1.2 16.9
GR00T-N1.5 50.7 14.8 2.7 23.9
RLDX-1 63.0 27.5 5.4 33.2
Qwen-RobotManip 68.6 20.1 14.9 35.9
Qwen-RobotManip-Context 63.9 22.6 11.2 33.8

Composite-Unseen (long-horizon in OOD scenes) 14.9% nearly triples RLDX-1's 5.4%. EBench (Shanghai AI Lab, Isaac Sim, dual-arm mobile Lift2 + R5a, 26 task types / 794 instances): 45.6% SR / 60 score overall vs Ο€0.5's 27.1 / 41, with Table-Top dexterous SR nearly 4Γ— Ο€0.5 (50.0 vs 12.9). Per-dimension: Qwen-RobotManip is nearly flat from Background (45.3) to Mix (46.8) while Ο€0.5 declines 33% β€” perturbation stacking barely touches it.

7.5 RoboTwin-IF (Table 8)

Method Pick-Diverse Place-Rel. Op.-Mic-Drawer Op.-Stapler Op.-Tabletop Avg.
StarVLA 11 13 0 49 74 29.4
GR00T-N1.7 20 17 0 14 32 16.6
Ο€0.5 44 20 15 92 66 49.6
Qwen-RobotManip 79 57 42 90 93 72.2
Qwen-RobotManip-Context 77 71 33 89 90 72.0

+22.6 pp average over Ο€0.5; GR00T-N1.7's 16.6% is a striking demonstration of VLA-to-VA degradation on a model that scores respectably on visual-perturbation benchmarks.

7.6 RoboTwin-XE (Table 9) β€” zero-shot to unseen morphologies

Method ARX-X5 UR5-WSG Franka Panda Total
Ο€0.5 (joint) 24.6 2.2 0.9 9.2
Ο€0.5 (eef) 11.5 10.0 1.1 7.5
Qwen-RobotManip (joint) 37.6 4.1 1.8 14.5
Qwen-RobotManip (eef) 42.9 22.8 5.9 23.9

The performance gradient (ARX > UR5 > Franka) tracks visual/kinematic similarity to the AgileX training platform. Camera-frame EEF vs own-joint: 5.6Γ— on UR5. This is the cleanest published validation of camera-frame action spaces for zero-shot embodiment transfer.

7.7 Real-world β€” CobotMagic ALOHA (Tables 10–11)

Fine-tuned on 22.9 h of teleop. In-domain, 7 tasks Γ— 5 trials: 88.6% vs Ο€0.5 42.9% / StarVLA 20.0% β€” 5/5 on five tasks including block-in-drawer-compartment and three-block-stacking where Ο€0.5 scores 0/5; only yellow-disc-insertion (contact-rich precision insertion) stays hard (2/5). OOD, 4 tasks Γ— 10 trials: 87.5% vs 37.5% / 0.0% β€” perfect 10/10 on cluttered-scene target grounding and left-right relational stacking (Ο€0.5: 1/10 on the latter), 9/10 under disco-light illumination.

7.8 Real-world β€” ARX ALOHA few-shot & skill transfer (Tables 12–13)

Few-shot (130 demos total across 5 tasks): best on 4/5 tasks; long-horizon Put Blocks 37.5% sub-step avg vs Ο€0.5 25.0%; Unscrew Cap full completion 3/10 vs 1/10; Insert Screw defeats everyone (0/10 insertions). Cross-embodiment skill transfer β€” joint fine-tune on 6K CobotMagic + 130 ARX demos, evaluate on 4 ARX tasks with zero ARX demonstrations of those tasks: 55.0% vs 7.5% (w/o UnifiedSpace) and 12.5% (w/o UnifiedEEF) β€” the alignment stack is what makes skills compose across bodies.

7.9 RoboChallenge Table30-v1 β€” 1st place, generalist track

Submitted as Lira_generalist; joint control; 30 tasks / 4 embodiments (ARX5, ALOHA, UR5, Franka), one policy per embodiment.

Method SR Process score
Qwen-RobotManip 45 59.83
DM0_generalist 37 48.43
Ο€0.5_generalist 17.67 31.27
GR00T-MULTI 15.33 32.29
Ο€0_generalist 9 20.22

Three analysis threads: bimanual coordination (40% avg over the 8 bimanual ALOHA tasks vs Ο€0.5's 21.2% β€” attributed to the dual-arm-heavy corpus and the H2R pipeline's bimanual synthesis; only model non-zero on pour-fries-into-plate at 30%); pick-and-place across embodiments (63.3% over 12 tasks vs DM0's 48.3%); and emergent retry behavior β€” spontaneous re-attempts after failed grasps/placements, observed across picking, pouring, folding, wiping, sweeping, hypothesized to come from imperfect-then-corrected segments in the diverse pretraining data. On 6 hard long-horizon tasks, prior SOTA averages 5% vs Qwen-RobotManip's 36.7%.


8. Ablations

8.1 The scaling-law ablation (Β§6.4, Figs. 18–19) β€” the paper's thesis in one experiment

Three action-space designs Γ— nested data subsets (1/5/10/25/50/100%), validation MSE on a held-out OOD set (15 embodiment types, 154 unseen tasks):

Variant Representation Scaling behavior
w/o UnifiedSpace Raw per-embodiment fields, zero-padded to 80 dims, no semantic alignment Unstable, high MSE, no consistent improvement with data
w/o UnifiedEEF Canonical 80-dim slots; EEF deltas as axis-angle relative to initial pose Log-linear MSE scaling βœ”
Ours + camera-frame delta EEF Log-linear βœ” with lowest EEF-prediction MSE

Downstream (RoboTwin-C2R after fine-tuning each variant): on Hard, full alignment scales steadily 1% β†’ 100% (reaching 50.2 joint / 56.6 eef) and beats both ablations at every data fraction; on Easy (in-distribution), no variant shows an upward data trend β€” the in-domain-benchmarks-can't-see-pretraining point, reproduced inside the ablation. Also: only the full model performs better in EEF mode than joint mode (72.5 vs 68.1 Easy); both ablations show the reverse.

8.2 Embodiment prompt & context (Table 15, RoboTwin-C2R joint, early checkpoint)

Config Denoise steps Easy Hard Avg
No prompt 4 71.2 54.2 62.7
Soft prompt 4 70.2 52.1 61.2
Language tag + FPS 4 71.7 55.1 63.4
Structured prompt 4 73.4 58.3 65.9
+ Context 4 72.1 54.4 63.3
+ Context 10 80.1 61.6 70.9
+ Context 20 79.8 62.1 71.0

The context mechanism's +5.0 over the structured prompt dwarfs every prompt-design delta β€” but only materializes with a 10-step denoising budget (at 4 steps the more complex action distribution produces jitter and the gain vanishes). 20 steps adds nothing.

8.3 Human-to-robot data (Tables 16–17)

Robot-only β†’ +raw ego β†’ +H2R at fixed 7:3 ratio: RoboTwin-C2R Hard 54.7 β†’ 55.0 β†’ 58.7; LIBERO-Plus total 87.1 β†’ 88.4 β†’ 89.0, with the largest per-dimension gain on Camera (72.8 β†’ 80.0, +7.2) β€” ego viewpoint diversity converted into robot-frame robustness. Monotonic progression: raw ego helps via visual diversity; the H2R rendering unlocks the rest via action + visual alignment.

8.4 VL co-training (Table 18)

LIBERO LIBERO-Plus RT-C2R Easy RT-C2R Hard RT-IF
Full (VL in pretrain, none in post-train) 99.1 90.1 73.2 62.6 71.6
βˆ’ VL in pretraining 98.2 88.9 66.5 54.4 (βˆ’8.2) 64.6 (βˆ’7.0)
+ VL also in post-training 98.6 91.4 74.0 62.5 73.1

VL co-training barely matters on in-distribution LIBERO (βˆ’0.9) but is worth 7–8 points on the hard OOD/instruction benchmarks β€” the co-training-as-generalization-infrastructure reading, consistent with LBM but measured on harder OOD axes. Adding VL to post-training specifically lifts language axes (LIBERO-Plus language perturbation 86.9 β†’ 93.9; RT-IF Pick-Diverse 76 β†’ 81).

8.5 Architecture wiring (Table 19, LIBERO-Plus)

Variant Total
Layer-wise self-attention (per-layer VLM feature fusion) 86.4
Last-layer self-attention (concatenation β€” the Qwen-VLA pattern) 87.0
Last-layer cross-attention (+ learned query tokens) 87.5

Cross-attention wins and is the cheapest (no per-layer fusion, no full VLM token set stored in the DiT). Learned query tokens (register-style, outputs discarded) accompany the state/action tokens in cross-attending to VLM states.


9. Comparative analysis

9.1 Qwen-RobotManip vs Qwen-VLA β€” the intra-Qwen comparison

Axis Qwen-RobotManip (Jun 2026) Qwen-VLA (May 2026)
Backbone Qwen3.5-4B (D=2560), fully unfrozen Qwen3.5-4B, frozen at T2A then unfrozen
Action expert DiT, 10 blocks, D=768, 12 heads (small) DiT, 16 blocks, 1.15B (large)
VLM↔expert wiring Cross-attention to last-layer states, alternating vision/language by block parity + learned queries Concatenation + joint self-attention
Action space 80-dim canonical; camera-frame delta EEF (CaPE-conditioned) or abs joint HΓ—K fixed tensor; native per-dataset conventions + quantile norm
Cross-embodiment interface Canonical slots + structured prompt + in-context history as implicit embodiment ID Embodiment text prompt only
Training Single-phase dual-stream co-training (9:1), K_repeat=8 Four stages: T2A β†’ CPT β†’ SFT β†’ PPO RL
RL None PPO with ODE→SDE flow log-prob
Data ~38,100 h, zero proprietary, 65% H2R synthetic >10,000 h public + >1,000 h in-house + language-only synthetic
Proprioceptive state Explicit 80-dim state via MLP into DiT Ablated to ≀+1.3 pp; omitted
Camera calibration Required for the main action mode (flag-switched fallback) Not required
Evaluation identity OOD-first; 2 new benchmarks; RoboChallenge #1 Generalist-vs-specialist head-to-heads; DOMINO zero-shot
Scope Manipulation only Manipulation + navigation + AD-VQA

Direct points of tension worth tracking:

  1. Wiring. RobotManip's Table 19 finds cross-attention > concatenation on LIBERO-Plus; Qwen-VLA shipped concatenation. The margin is small (0.5 pp) and the setting differs, but within one organization the two flagship VLAs disagree on the single most-debated design axis in Review-VLM-Action-Connection.
  2. State. Qwen-VLA found explicit proprioception nearly worthless; RobotManip makes an 80-dim state vector a first-class input. Not directly contradictory (different action spaces β€” camera-frame deltas may need state anchoring), but unreconciled.
  3. Recipe complexity. Qwen-VLA bets on stage curriculum + RL; RobotManip bets on representation alignment + data scale with the simplest possible one-phase recipe. RobotManip's LIBERO-Plus 91.4 vs the numbers Qwen-VLA reported on different benchmarks are not directly comparable β€” neither paper evaluates the other.

9.2 Against the field

Axis Qwen-RobotManip Ο€0.5 GR00T N1.x TRI LBM
Expert wiring Cross-attn DiT (small, D=768) Same-stack expert w/ prefix attention Cross-attn DiT adaLN-conditioned 8-layer head
Action space Camera-frame delta EEF (+ abs joint mode) Robot-frame continuous Embodiment-projected continuous Robot-frame continuous
Camera geometry in the policy CaPE in attention (q,k,v,out) + intrinsics tokens No No No
Data thesis Open + H2R synthesis (65% synthetic) In-house-heavy + web co-training Data pyramid incl. ~20k h ego (EgoScale) 523 h target + OXE + ego
Anti-forgetting 9:1 dual-stream co-training Co-training + KI Partial-layer tuning Frozen backbone + VL stream
In-context adaptation Yes (execution-history conditioning) No No No
New benchmarks shipped RoboTwin-IF, RoboTwin-XE No No No

Genuinely novel: (a) the camera-frame delta EEF + CaPE-in-the-action-expert combination at foundation scale, and the demonstration that it creates the data scaling law; (b) in-context policy adaptation with stochastic context sampling as an anti-shortcut mechanism, interpreted as implicit embodiment identification; (c) the 15-platform H2R synthesis pipeline at 24,808 h β€” an order of magnitude beyond Phantom/Masquerade-scale predecessors; (d) the RoboTwin-IF / RoboTwin-XE benchmarks and the "VLA-to-VA degradation" framing; (e) the curation pipeline's quantified findings (81% of RoboMIND UR data failing causality checks is a service to the field).

Recombination: the decoupled VLM + flow-matching DiT (Ο€0/GR00T family), VL co-training as anti-forgetting (LBM, Ο€0.5), masked canonical action vectors (Qwen-VLA, GR00T), ECoT supervision (Zawalski et al.), human-video-to-robot editing (Phantom, Masquerade), RTC deployment (Black et al.).


10. Significance & positioning

  • The strongest open-data result in VLA to date. Every prior model at this performance tier (Ο€0.5/0.6, GR00T, Qwen-VLA) leans on proprietary teleop. If alignment + synthesis really substitutes for in-house collection, the economics of manipulation foundation models change.
  • Camera-frame actions graduate from technique to thesis. Prior appearances (Chen et al. 2025a; Zhang et al. 2026b) were single-robot studies; this is the first foundation-scale validation with a scaling-law argument attached. Expect the Action Space: EEF vs Joint debate to acquire a third axis: which frame.
  • Evaluation-methodology influence may outlast the model. The from-scratch-matches-pretrained result plus RoboTwin-IF/XE give reviewers concrete tools to demand OOD evidence. RoboTwin-IF in particular measures a failure mode (language-conditioning collapse) no prior benchmark isolates β€” and GR00T-N1.7's 16.6% on it shows the failure is real in shipping models.
  • A weaker claim than it appears in one respect: most OOD evaluation is still simulation (LIBERO-Plus, RoboTwin variants, RoboCasa365, EBench are all sim), which the authors themselves concede; the real-world OOD suite is 4 tasks on one platform.

11. Limitations

11.1 Authors' stated limitations (Β§7)

  1. H2R synthesis quality bounds. Retargeting approximations and inpainting artifacts introduce distributional gaps that cap the effective quality of the 24,808 synthesized hours.
  2. OOD evaluation is still predominantly simulation-based; broader real-world deployment-condition coverage is needed.
  3. Fixed action-chunk length and inference latency constrain reactive sub-second control.

11.2 Reviewer's concerns (not in the paper)

  1. No weights, contradicting the paper's own text. Β§6.4 says "we release both the context-conditioned and context-free variants," but the GitHub README (as of this review, Jul 2026) states there is no plan to release model weights for Qwen-RobotManip or Qwen-RobotNav. Every headline number is therefore currently unreproducible outside Alibaba, and the two new benchmarks' baselines cannot include the model that motivated them.
  2. Camera calibration as a hidden deployment tax. The flagship camera-frame mode requires calibrated intrinsics + extrinsics at inference. The fallback (base-relative mode via the flag embedding) exists, but no evaluation isolates how much performance survives without calibration β€” a key question for in-the-wild deployment, where DROID-style setups often have rough calibration at best.
  3. Reproducibility gap on training details. No optimizer, learning rates, batch sizes, step counts, GPU type/hours for pretraining. Same criticism applied to Qwen-VLA; unchanged here.
  4. No latency numbers despite a latency-sensitive design. Remote WiFi inference + RTC + (for the Context variant) 10 denoising steps and a longer VLM sequence; no ms-per-chunk is reported. The Context variant's start-of-episode hesitation is honestly disclosed, but its runtime cost is not.
  5. Context variant is not uniformly better. It loses to the base model on RoboCasa365 (33.8 vs 35.9), EBench (43.6 vs 45.6), and ties on RoboTwin-IF β€” the gains concentrate on perturbation-robustness benchmarks (LIBERO-Plus, RT-C2R). The paper does not analyze why history hurts long-horizon composite tasks; plausibly the stochastic-context training trades temporal-progress information away.
  6. No comparison against its own sibling. Qwen-VLA and Qwen-RobotManip share a backbone and a benchmark-capable evaluation stack, yet neither paper cites or evaluates the other. A controlled comparison (concatenation vs cross-attention at matched scale; T2A+RL vs alignment-first) is exactly the experiment the community needs and only this team can run.
  7. Ο€0.6/Ο€0.7 absent from baselines. Comparisons stop at Ο€0.5 (Apr 2025) and GR00T-N1.5/1.6/1.7. Understandable for closed models, but the "substantially outperforms prior SOTA" claim is dated against a 14-month-old competitor.
  8. H2R data is 65% of the corpus but only ablated at small scale. Tables 16–17 show +1.9–4.0 pp from H2R at a reduced-scale 7:3 mixture. Whether 24,808 h of synthetic data provides value proportional to its two-thirds share of the corpus β€” versus, say, a 5,000 h subset β€” is not measured; no H2R-fraction sweep exists.
  9. RoboChallenge is a leaderboard, not a paper protocol. The 1st-place claim (as Lira_generalist) is externally verifiable, but per-task trial counts and evaluation dates are controlled by the challenge, and the "20% relative improvement" headline is over the next submission at a point in time.
  10. The scaling-law evidence is MSE + one downstream benchmark. The log-linear claim rests on validation MSE (a proxy with a known weak correlation to task success in the literature) plus RoboTwin-C2R fine-tuning. Success-rate scaling curves on a second, non-RoboTwin domain would make the "alignment unlocks scale" thesis much harder to attack.

12. Links


13. Related pages

← Back to Home

⚠️ **GitHub.com Fallback** ⚠️