Review WAM vs VLA Robustness - Heungwoo/research GitHub Wiki

In-Depth Review β€” Do World Action Models Generalize Better than VLAs? A Robustness Study

Paper: Do World Action Models Generalize Better than VLAs? A Robustness Study Authors: Zhanguang ZhangΒΉ*, Zhiyuan LiΒΉΒ·Β² (Huawei Canada internship), Behnam RahmatiΒΉ, Rui Heng YangΒΉ, Yintao MaΒΉ, Amir RasouliΒΉ, Sajjad PakdamansavojiΒΉ, Yangzheng WuΒΉ, Lingfeng ZhangΒΉ, Tongtong CaoΒΉ, Feng WenΒΉ, Xinyu WangΒΉ, Xingyue QuanΒΉ, Yingxue ZhangΒΉ* (corresponding: [email protected], [email protected]) Affiliations: ΒΉHuawei Technologies Β· Β²University of Toronto arXiv: 2603.22078 Β· v3, April 30, 2026 Code / benchmark: RoboTwin 2.0-Plus benchmark introduced in the paper (not yet publicly released as of the v3 PDF)

This page sits in the VLA architectures review as the controlled empirical comparison between Category-E world-action-model architectures and the rest of the VLA tree β€” the WAM-side counterpart to TRI's data-side study LBM Co-training and to the VLA-only robustness sibling RobustVLA.


1. TL;DR

Huawei's team builds a side-by-side benchmark of seven VLAs, one VLA+world-model hybrid, and four WAMs across two single-arm and bimanual perturbation suites (LIBERO-Plus + the paper's new RoboTwin 2.0-Plus). 13,000+ rollout episodes give a controlled answer to whether the "video-pretrain wins" thesis advocated by Cosmos-Policy, LingBot-VA, GE-Act and others actually shows up under perturbation.

Five verdicts:

  1. WAMs win on visual perturbations. LingBot-VA (74.2% on RoboTwin 2.0-Plus) and Cosmos-Policy (82.2% on LIBERO-Plus) lead their respective benchmarks under noise, lighting, layout, and background perturbations. The spatiotemporal prior from video-generation pretraining transfers directly to these axes.
  2. WAMs lose on geometric perturbations. Camera viewpoint and robot initial-state perturbations remain the WAM weak spot on both benchmarks. Video pretraining does not appear to teach the policy what its own kinematics look like from a shifted viewpoint.
  3. Data-diverse VLAs can match or beat WAMs β€” but only when the VLA is trained at PI's scale. Ο€0.5 achieves 85.7% overall on LIBERO-Plus (best in the table) by leaning on web data + diverse robot data + a JAX flow-matching expert. Smaller / less-diverse VLAs (Ο€0, Ο€0-FAST, OpenVLA-OFT, UniVLA, RIPT-VLA) collapse on the same benchmark.
  4. Inference latency is the unsolved WAM problem. Cosmos-Policy is 6.2Γ— slower than Ο€0.5; LingBot-VA at RoboTwin denoising settings is 83Γ— slower (5.2 s per chunk). Even the fastest WAM (Fast-WAM at 190 ms) is 3Γ— slower than Ο€0.5 (63 ms).
  5. The thesis is paradigmatic, not universal. Hybrid approaches (MOTUS, VLA-JEPA) sit between pure VLAs and pure WAMs, and Fast-WAM's LIBERO collapse (97.6% β†’ 51.5%) shows that the WAM video prior is necessary but not sufficient β€” task-specific training-data diversity is still a hard requirement.

The conclusion is mixed-with-asymmetry: WAMs win on perturbation-side generalization given video priors, but Ο€0.5-class VLAs match them via data diversity, and the latency gap means WAMs are not yet shippable in the deployment regimes where Ο€0.5 already runs.


2. Why this paper matters in the 2026 landscape

By early 2026 the WAM thread (defined here as video-generation-model-as-backbone for robot control) had three loud claims out:

  • NVIDIA Cosmos-Policy (ICLR 2026): "fine-tune Cosmos-Predict2 with light modifications; latent-frame action encoding gives you a planning-capable policy with future-state prediction for free."
  • LingBot-VA / DreamZero / GigaWorld-Policy: "autoregressive interleaved video + action generation produces state-of-the-art robot policies once you accept the latency cost."
  • GE-Act (Genie Envisioner) and Fast-WAM: "you can cut WAM latency by limiting denoising steps or skipping video generation at test time."

Each individual WAM paper claimed strong generalization or robustness wins on its own evaluation setup β€” but no paper had compared more than two WAMs to more than one VLA family on a perturbation-controlled axis. This study is the first to do so:

Claim (2025–2026) Verdict in this paper
WAMs generalize better than VLAs under visual perturbation βœ… confirmed (Cosmos-Policy, LingBot-VA top noise/light/layout)
WAMs generalize better under all perturbations ❌ camera + robot-state perturbations remain a weak spot
Hybrid VLA + WM (auxiliary objective) is "free" ⚠️ helps, but falls short of dedicated WAMs (VLA-JEPA 77.9% vs Cosmos-Policy 82.2% on LIBERO-Plus)
Removing test-time video generation closes the latency gap (Fast-WAM) ⚠️ 190 ms is closer but still 3Γ— Ο€0.5
WAMs don't need embodied pre-training (Fast-WAM thesis) ⚠️ true on RoboTwin (with domain-randomized training data); false on LIBERO (with clean-only data)
VLAs need lots of diverse data to match WAMs βœ… Ο€0.5 confirms (web + mobile + cross-embodiment + post-train), smaller VLAs collapse

This is the controlled study the architecture review Β§5 Category E section had flagged as missing β€” the architectural-family-level empirical reference that the VAM (E4) and auxiliary-world-model (E1) sub-patterns can now both be evaluated against.


3. WAM definition used in the paper

The paper uses a strict definition of WAM that distinguishes it from the broader "world model in a VLA" thread:

A World Action Model (WAM) is built on a pretrained video-generation backbone (Cosmos-Predict2, LTX-Video, Wan2.1, Wan2.2, Stable Video Diffusion, etc.), with minimal or no architecture modifications, and is fine-tuned to produce robot actions. Robot joint state encodings are added as lightweight modifications to the video backbone.

This is stricter than the Category-E definition in Review-VLA-Architecture:

  • E1 (auxiliary world-model loss on a VLM) is not a WAM in this paper's sense β€” those models (DreamVLA, WorldVLA, Cosmos-Policy as auxiliary loss) are categorized as "VLA + WM" hybrids.
  • E2 (DreamGen / GR00T-Dreams as data factory) is not a WAM β€” those are data-augmentation pipelines, not policies.
  • E3 (Ctrl-World / WMPO world-model-as-environment) is not a WAM β€” those are RL-on-WM, not video-backbone policies.
  • E4 (mimic-video / GE-Act / Cosmos-Policy / LingBot-VA / DreamZero / GigaWorld-Policy / Fast-WAM) is a WAM.

So in the paper's vocabulary: WAM = video-generation-model-as-VLA-backbone = Category E4 in the architecture review.

Table 1 in the paper enumerates the WAM candidates considered, with their characteristic flags:

Model Params Video Backbone MOT Pretrain-Free Causal-Pred AR-Gen
VPP 1.5B Stable Video Diffusion βœ— βœ— βœ“ βœ—
GE-Act 2.2B LTX-Video-2B βœ“ βœ— βœ— βœ“
Cosmos-Policy 2B Cosmos-Predict2-2B βœ— βœ“ βœ— βœ—
LingBot-VA 5.3B Wan2.2-5B βœ“? βœ— βœ“ βœ“
DreamZero 14B Wan2.1-14B βœ— βœ— βœ— βœ“
GigaWorld-Policy >5B Wan2.2-5B βœ— βœ— βœ“ βœ—
Fast-WAM 6B Wan2.2-5B βœ“ βœ“ βœ— βœ—
  • MOT = mixture-of-transformers (separate backbones for video and action with cross-modal attention).
  • Pretrain-Free = no task-agnostic embodied pre-training required.
  • Causal-Pred = action conditioned on predicted future state (IDM-style) or vice-versa.
  • AR-Gen = autoregressive generation.

The paper also notes that MOTUS is excluded from the WAM list because, although it uses a pretrained video backbone, action generation is routed through a separate VLM expert β€” making it a hybrid, not a video-backbone policy.

DreamZero (14B Wan2.1 backbone) and GigaWorld-Policy are surveyed but excluded from quantitative evaluation β€” DreamZero because re-training is prohibitively expensive and warm-up exceeds 15 minutes; GigaWorld-Policy because no checkpoint is yet public.


4. Comparison setup

4.1 Architectural families evaluated

flowchart TB
  subgraph V[VLAs - pure language/action models]
    V1[Ο€0 β€” RSS 2025 baseline]
    V2[Ο€0-FAST β€” FAST tokenized AR]
    V3[Ο€0.5 β€” CoRL 2025 Oral, JAX]
    V4[OpenVLA-OFT_m]
    V5[UniVLA β€” DCT tokens]
    V6[RIPT-VLA]
    V7[X-VLA β€” soft-prompt cross-embodiment]
    V8[HoloBrain0-GD, ABot-M0 β€” only on LIBERO-Plus]
  end
  subgraph H["VLA + WM (hybrid)"]
    H1[VLA-JEPA β€” future-state aux on Qwen3-VL-2B]
    H2[MOTUS β€” Wan2.2-5B video + separate VLM expert]
  end
  subgraph W[WAMs - video-generation-as-backbone]
    W1[GE-Act β€” LTX-Video-2B + flow-matching decoder]
    W2[Cosmos-Policy β€” Cosmos-Predict2-2B latent-frame action]
    W3[LingBot-VA β€” Wan2.2-5B AR interleaved]
    W4[Fast-WAM β€” Wan2.2-5B joint denoising, skip-video at test]
  end

  classDef vla fill:#bbdefb,stroke:#1565c0,color:#000
  classDef hybrid fill:#fff9c4,stroke:#f57f17,color:#000
  classDef wam fill:#c8e6c9,stroke:#2e7d32,color:#000
  class V1,V2,V3,V4,V5,V6,V7,V8 vla
  class H1,H2 hybrid
  class W1,W2,W3,W4 wam
Loading

4.2 Benchmarks

Aspect LIBERO-Plus RoboTwin 2.0-Plus
Simulator MuJoCo (robosuite) SAPIEN (ManiSkill3)
Robot Franka Panda (7-DoF) Aloha-AgileX (14-DoF bimanual)
Arms Single Dual
Cameras 2 (third-person + wrist) 3 (head + 2 wrist)
Image resolution 256Γ—256 320Γ—240
Native action space 7-dim delta EEF 14-dim joint positions
Control mode OSC (delta EE pose) Joint position control
Control frequency 10 Hz 25–30 Hz
Base tasks 40 (4 suites Γ— 10) 50 collaborative tasks
Training demos 50 / task 50 clean + 500 domain-randomized
Total trajectories 22,400 27,500
Distractor objects 416 731 (147 categories)

LIBERO-Plus is the existing open benchmark from Fei et al. (2025) β€” single-arm robustness.

RoboTwin 2.0-Plus is introduced by this paper β€” built on top of RoboTwin 2.0 (Chen et al., 2025a), applies LIBERO-Plus-style perturbations to dual-arm bimanual tasks. The two benchmarks differ in embodiment, sensing configuration, action space, and base-task style, which the authors argue makes them complementary rather than redundant.

4.3 Perturbation taxonomy

Both benchmarks use 7 perturbation dimensions Γ— 21 sub-dimensions with one perturbation enabled per evaluation branch (8 configs total: 1 clean + 7 perturbed).

flowchart LR
  subgraph PT[7 perturbation dimensions]
    N["Sensor Noise (N1–N5)<br/>motion blur Β· Gaussian Β· zoom blur Β· fog Β· glass"]
    L["Lighting (L1–L4)<br/>diffuse Β· direction Β· specular Β· shadows"]
    C["Camera (C1–C3)<br/>distance Β· spherical pos Β· orientation"]
    R[Robot Init State<br/>joint Gaussian + extreme gripper]
    B["Background (B1+B2)<br/>scene theme + surface appearance"]
    O["Layout (O1+O2)<br/>distractor count + target pose"]
    I["Language (R1+R2+R3)<br/>distraction Β· reword Β· reasoning chain"]
  end
  classDef pt fill:#fff9c4,stroke:#f57f17,color:#000
  class N,L,C,R,B,O,I pt
Loading

C2 (spherical-position camera perturbation) is disabled by default in the RoboTwin 2.0-Plus evaluation config because the authors found it caused too much off-distribution instability to be informative.

4.4 Evaluation protocol

  • Per task per branch: 50 rollout episodes
  • Total LIBERO-Plus episodes (per model): ~16,000 (40 tasks Γ— 8 configs Γ— 50 ep)
  • Total RoboTwin 2.0-Plus episodes (per model): ~20,000 (50 tasks Γ— 8 configs Γ— 50 ep)
  • Ο€0.5 is finetuned in JAX from the pretrained Ο€0.5 checkpoint on full RoboTwin 2.0 (27.5k training data, 60k steps, AdamW, cosine 2.5e-5 β†’ 2.5e-6, BS=64), per the openpi recommended config β€” done in-paper because no JAX Ο€0.5 checkpoint existed for RoboTwin 2.0 at evaluation time.
  • All other models use their publicly released checkpoints (X-VLA, LingBot-VA, MOTUS on RoboTwin 2.0; Ο€0 / Ο€0-FAST / OpenVLA-OFT / UniVLA / RIPT-VLA / VLA-JEPA / GE-Act / Cosmos-Policy / Fast-WAM on LIBERO).
  • The Fast-WAM LIBERO checkpoint was trained on clean demonstrations only; the Fast-WAM RoboTwin checkpoint was trained on clean + domain-randomized demonstrations β€” this asymmetry becomes a natural experiment in RQ 2.

5. Headline results

5.1 RoboTwin 2.0-Plus (bimanual)

Table 3 from the paper:

Model Original Camera Robot Lang. Light BG Noise Layout Total
VLAs
Ο€0.5 78.4 45.6 27.6 74.4 49.6 71.7 64.9 56.8 58.6
X-VLA 65.6 23.2 65.2 64.4 63.1 58.6 49.7 34.8 53.1
VLA + WM
MOTUS 87.0 21.6 85.0 83.2 84.6 84.4 43.1 82.8 71.5
WAMs
LingBot-VA 92.1 28.9 36.2 87.3 89.0 91.3 80.9 87.9 74.2
Fast-WAM 91.2 30.4 53.2 86.7 88.8 90.0 76.4 83.2 72.7

LingBot-VA wins 5/7 perturbation categories and the overall ranking. The two it loses are exactly the two where video priors do not apply: camera viewpoint (won by Ο€0.5, 45.6%) and robot init state (won by MOTUS, 85.0%).

5.2 LIBERO-Plus (single-arm)

Table 4:

Model Original Camera Robot Lang. Light BG Noise Layout Total
VLAs
Ο€0 94.2 13.8 6.0 58.8 85.0 81.4 79.0 68.9 53.6
Ο€0 (rerun, JAX) 91.3 61.0 40.8 63.5 89.3 84.1 80.1 76.4 69.4
Ο€0-FAST 85.5 65.1 21.6 61.0 73.2 73.2 74.4 68.8 61.6
Ο€0.5 96.9 75.4 77.5 85.6 96.9 94.6 89.7 85.7 85.7
OpenVLA-OFT_m 97.6 55.6 21.7 81.0 92.7 91.0 78.6 68.7 67.9
UniVLA 95.2 1.8 46.2 69.6 69.0 81.0 21.2 31.9 42.9
RIPT-VLA 97.5 55.2 31.2 77.6 88.4 91.6 73.5 74.2 68.4
X-VLA 98.1 23.4 89.7 75.7 88.2 96.0 62.7 71.8 71.4
HoloBrain0-GD 96.7 65.5 58.2 78.7 88.1 90.3 66.9 79.5 74.0
ABot-M0 98.6 60.4 67.9 86.4 96.2 91.6 86.4 82.6 80.5
VLA + WM
VLA-JEPA 97.2 64.2 67.7 88.1 91.8 93.4 65.8 83.9 77.9
WAMs
GE-Act 94.4 60.7 77.0 77.4 95.8 86.0 90.9 80.2 80.3
Cosmos-Policy 98.5 75.8 63.3 81.7 96.5 88.9 92.7 82.2 82.2
Fast-WAM 97.6 16.4 44.5 68.9 78.2 53.7 37.7 60.7 51.5

The picture is sharper here than on RoboTwin. On LIBERO-Plus, Ο€0.5 (85.7%) is the single best model overall, outscoring every WAM. Cosmos-Policy (82.2%) takes the WAM crown and matches Ο€0.5 on Camera and Noise. The gap is driven by Ο€0.5's superiority on Robot Init State (77.5% vs Cosmos-Policy 63.3%) and BG (94.6% vs 88.9%).

Fast-WAM collapses on LIBERO-Plus (51.5%) β€” see RQ 2 below.

5.3 Inference latency (Table 5)

Model Chunk size Inference (wall-clock) Relative to Ο€0.5
Ο€0.5 50 63 ms 1.0Γ—
X-VLA 30 195 ms 3.1Γ—
Fast-WAM* 16 190 ms 3.0Γ—
GE-Act 36 300 ms 4.8Γ—
Cosmos-Policy 16 390 ms 6.2Γ—
LingBot-VA (RW config) 32 480 ms 7.6Γ—
MOTUS 16 1175 ms 18.6Γ—
LingBot-VA (RT config) 32 5230 ms 83.0Γ—

*Fast-WAM latency reported by Yuan et al. on their hardware, not re-measured by the authors.

Why the variance: number of denoising steps for state and action dominates the runtime, with state denoising typically dominating action denoising. LingBot-VA uses 3 state + 5 action steps in real-world deployment vs 25 + 50 in RoboTwin β€” the ~11Γ— wall-clock gap (480 ms β†’ 5230 ms) between those two configurations on the same model is striking. GE-Act uses 1 state + 10 action denoising steps (the lightest); Cosmos-Policy uses 5 + 5; MOTUS uses 10 + 10.

By omitting visual state generation entirely at test time, Fast-WAM achieves the lowest WAM latency in the comparison (190 ms) β€” but it is still 3Γ— slower than Ο€0.5 (63 ms).


6. Side-by-side tradeoff table

This is the OBJECTIVE balanced summary requested. Numbers are all from this paper unless noted otherwise.

Axis VLA (Ο€0.5-class, data-diverse) VLA (other, e.g. OpenVLA-OFT / UniVLA / Ο€0-FAST) VLA + WM hybrid (MOTUS / VLA-JEPA) WAM (E4 β€” Cosmos-Policy / GE-Act / LingBot-VA) WAM (Fast-WAM β€” joint-denoising, skip-video at test)
In-distribution success (clean) High (Ο€0.5: 96.9% LIBERO, 78.4% RT) Mixed (Ο€0-FAST 85.5%, UniVLA 95.2%, OpenVLA-OFT 97.6% LIBERO; collapses on RoboTwin if not trained on RT 2.0) High (MOTUS 87.0% RT, VLA-JEPA 97.2% LIB) Very high (LingBot-VA 92.1% RT, Cosmos-Policy 98.5% LIB, GE-Act 94.4% LIB) High (Fast-WAM 91.2% RT, 97.6% LIB)
OOD generalization (avg under perturbations) Ο€0.5 best in study (85.7% LIB Β· 58.6% RT). Smaller VLAs collapse badly. Bottom-tier under perturbation (UniVLA 42.9%, Ο€0-FAST 61.6%, OpenVLA-OFT 67.9% on LIB) Strong (MOTUS 71.5% RT, VLA-JEPA 77.9% LIB) β€” between VLA and WAM Best on visual perturbations (Cosmos-Policy 82.2% LIB, LingBot-VA 74.2% RT) Highly dependent on training data: 72.7% RT (with domain randomization) vs 51.5% LIB (clean only)
Robustness to action perturbations Not tested directly (this paper). See RobustVLA for action-noise eval β€” Ο€0 wins there. Not tested. Not tested. Not tested. Not tested.
Robustness to observation perturbations (noise / lighting) Ο€0.5 strong (96.9% light, 89.7% noise on LIB) Variable; UniVLA collapses on noise (21.2%) MOTUS weak on noise (43.1% RT); VLA-JEPA mid (65.8% noise LIB) Strongest (Cosmos-Policy 92.7%, GE-Act 90.9% on LIB noise; LingBot-VA 80.9% on RT noise) Strong with diverse data (76.4% RT noise) / weak without (37.7% LIB noise)
Robustness to language perturbations Ο€0.5 best on LIB (85.6%); decent on RT (74.4%) Mid (60–80%) VLA-JEPA strong on LIB (88.1%), MOTUS strong on RT (83.2%) Mid (LingBot-VA 87.3% RT, Cosmos-Policy 81.7% LIB) Variable (Fast-WAM 86.7% RT, 68.9% LIB)
Robustness to environment perturbations (BG / layout) Ο€0.5 strong (94.6% BG, 85.7% layout LIB; 71.7%, 56.8% on RT) OpenVLA-OFT weak on layout (68.7% LIB) MOTUS strong on RT (84.4% BG, 82.8% layout) Strongest on layout (LingBot-VA 87.9%, Cosmos-Policy 82.2%, GE-Act 80.2%) Variable (53.7% LIB BG when clean-only training)
Robustness to camera viewpoint Ο€0.5 best in study (75.4% LIB, 45.6% RT) Mostly collapse (UniVLA 1.8%, X-VLA 23.4–24.5% on LIB) Strong (VLA-JEPA 64.2% LIB) Weak: Cosmos-Policy 75.8% on LIB (only this one comparable to Ο€0.5), GE-Act 60.7%, LingBot-VA 28.9% on RT Very weak (16.4% LIB)
Robustness to robot init state Ο€0.5 strong on LIB (77.5%), weak on RT (27.6%) UniVLA 46.2%, Ο€0 6.0%, OpenVLA-OFT 21.7% on LIB MOTUS strongest on RT (85.0%), VLA-JEPA strong on LIB (67.7%) Weak: Cosmos-Policy 63.3%, GE-Act 77.0%, LingBot-VA 36.2% (RT) Weak (44.5% LIB, 53.2% RT)
Sample efficiency (embodied pre-training requirement) High requirement (Ο€0.5: web data + 400h mobile manip + cross-embodiment + post-train) Moderate (e.g. OpenVLA: 970k cross-embodiment) Moderate (VLA-JEPA: 220k human ego + 76k single-emb; MOTUS: 231k human ego + 781k cross-emb + 1k task-agnostic) Moderate-to-Low (Cosmos-Policy: only 185 task-specific trajectories; GE-Act: 3k h single-embodiment + 1 h task-specific; LingBot-VA: 16k h cross-embodiment + 50 trajectories) Lowest β€” Fast-WAM requires no embodied pre-training, only 60 h task-specific demos
Inference latency Fastest (Ο€0.5 63 ms) X-VLA 195 ms MOTUS 1175 ms (slowest hybrid); VLA-JEPA not measured Slowest (GE-Act 300 ms, Cosmos-Policy 390 ms, LingBot-VA 480 ms RW / 5230 ms RT) Lightest WAM (190 ms = 3Γ— Ο€0.5)
Backbone parameters Ο€0.5 ~4B class OpenVLA 7B; smaller models <1B 2B (VLA-JEPA Qwen3-VL-2B) – 5B (MOTUS Wan2.2-5B + extra VLM expert) 2B–14B (Cosmos-Predict2-2B = 2B; Wan2.2-5B = 5.3B; Wan2.1-14B = 14B) ~6B (Wan2.2-5B + decoder)
Cross-embodiment readiness Ο€0.5 designed for it; cross-embodiment in pre-training X-VLA designed for it (soft prompts) MOTUS uses optical-flow latent action space β€” cross-embodiment friendly Latent-frame / absolute-EEF / quaternion+joint representations are not natively cross-embodiment Unknown β€” small literature
Causal coupling style Pure VLA mapping h_t β†’ a_t Pure VLA VLA + WM (VLA-JEPA: future state alignment loss; MOTUS: separate video + VLM expert via MoT) Variable: LingBot-VA (state-then-action causal, IDM-style), Cosmos-Policy (jointly denoise), GigaWorld-Policy (action-then-state) Joint state+action denoising; skip state at test
Failure mode under noise Misalignment cascades on first grasp (Fig. 2c) Action-token decoder fails outright in extreme cases MOTUS noise-weak (43.1% RT noise) Robust noise denoising in video predictions (Fig. 3) Robust if training data was domain-randomized; collapses otherwise
Failure mode under camera perturbation Ο€0.5 holds up (45.6% RT, 75.4% LIB) Catastrophic collapse for most VLA-JEPA holds up (64.2% LIB) Severe collapse (LingBot-VA 28.9% RT; Fast-WAM 16.4% LIB) β€” video priors do not transfer geometry Severe collapse
Failure mode under background Ο€0.5 strong (94.6% LIB) Mostly strong (X-VLA 96.0%, OpenVLA-OFT 91.0%) Strong (93.4% LIB) Cosmos-Policy 88.9% LIB; predicted future image distorts in spatial layout under unusual backgrounds (paper Fig. 3) β€” degrades action accuracy Fast-WAM 53.7% LIB (clean-trained only)

The biggest takeaway from this table: no architectural family is uniformly best. WAMs lead on noise/light/layout but lose on camera/robot-state; Ο€0.5 (data-diverse VLA) is the most balanced single model in the study but at the cost of the largest training pipeline.


7. Detailed per-research-question findings

The paper organizes its analysis around four research questions; each is summarized below.

7.1 RQ 1 β€” Are WAMs robust to perturbations?

Verdict: Yes, WAMs demonstrate strong robustness on both benchmarks, but not uniformly.

  • LingBot-VA: 92.1% on original RoboTwin 2.0 β†’ 74.2% averaged across the 7 perturbations (drop ~18 pp). Best in 5 of 7 categories.
  • Cosmos-Policy: 98.5% on original LIBERO β†’ 82.2% averaged. Best of WAMs on LIBERO.
  • Ο€0.5 on LIBERO-Plus (85.7% total) is higher than any WAM on the same benchmark β€” establishing the upper bound of what data-diverse VLAs can do.
  • Fast-WAM (no embodied pre-training, RT-only domain-randomized data) ties LingBot-VA on RoboTwin (72.7% vs 74.2%) β€” showing video backbone + diverse task-data is sufficient without embodied pre-training.
  • MOTUS (hybrid) is third best on RT (71.5%), VLA-JEPA (hybrid) is mid-tier on LIBERO (77.9%). Partial integration of video priors yields intermediate robustness.

Three representative failure cases for Ο€0.5 are documented (Fig. 2):

  • Beat-block-with-hammer (noise N3): Ο€0.5 collides with the hammer; LingBot-VA succeeds.
  • Handover block (layout): Ο€0.5 collides with the red block during approach; LingBot-VA succeeds.
  • Rank RGB blocks (lighting): Ο€0.5 fails to grasp the first red block due to misalignment and does not recover; LingBot-VA succeeds.

7.2 RQ 2 β€” Is the advantage consistent across perturbation types?

Verdict: No β€” geometric perturbations (camera, robot init state) are a WAM weakness, while visual perturbations (noise, light, layout, background) are a WAM strength.

Fast-WAM's two checkpoints provide the cleanest natural experiment in the paper:

Setting Original Avg under perturbations Drop
Fast-WAM on RT 2.0-Plus (clean + domain-randomized training data) 91.2% 72.7% -18.5
Fast-WAM on LIBERO-Plus (clean training data only) 97.6% 51.5% -46.1

The architecture is identical; only the training data diversity differs. The verdict: the WAM video prior is necessary but not sufficient for robustness. Task-specific training-data diversity remains a critical lever even when the backbone is video-pretrained.

The paper also speculates that joint-denoising WAMs (Fast-WAM) may rely on training-data diversity more than IDM-style WAMs (LingBot-VA) that explicitly condition action on a predicted future state β€” the explicit causal coupling acts as an additional architectural prior for robustness.

Cosmos-Policy's future-image predictions (paper Fig. 3) show:

  • βœ… Predictions remain accurate under noise (model actively denoises the moving arm β€” attributed to web-video pretraining on dynamic objects).
  • βœ… Predictions remain accurate under light variation.
  • ❌ Predictions exhibit severe spatial distortions and inconsistent color schemes under certain background perturbations β€” likely because the perturbed backgrounds are far from the training distribution. This degradation propagates to inaccurate action generation.

7.3 RQ 3 β€” Why is there a difference between VLAs and WAMs?

Verdict: Backbone pretraining objective + embodied-data requirement.

  • VLA backbones (VLMs) are pretrained on static image-text data with next-token-prediction objective β†’ lack fine-grained dynamic priors β†’ require diverse robot + web + cross-embodiment data to compensate.
  • WAM backbones (video-generation models like Cosmos-Predict2, Wan2.2, LTX-Video) are pretrained on web-scale video β†’ already capture fine-grained spatiotemporal transitions β†’ embodied pre-training can focus on learning action prediction conditioned on visual state (a relatively easier learning problem given priors).
  • Hybrid (MOTUS, VLA-JEPA) integrates video / human-video data into VLA training pipelines β€” partial benefit because the backbone is still VLM-class.

The paper frames this as a data-efficiency advantage of the WAM paradigm: Cosmos-Policy was finetuned on only 185 task-specific trajectories (Table 2) and still achieved 82.2% on LIBERO-Plus. Ο€0.5 needed web data + mobile manipulation 400h + cross-embodiment + multi-env tabletop + post-training data to reach its 85.7% number.

The decomposition the paper appeals to is:

$$p_\phi(h_{t+1}, a_t \mid h_t) \quad \text{(joint, WAM)} \qquad \text{vs.} \qquad p_\theta(a_t \mid h_t) \quad \text{(VLA)}$$

or for IDM-style WAMs:

$$p_\phi(h_{t+1} \mid h_t) \cdot g_\psi(a_t \mid h_t, h_{t+1})$$

The video pretraining objective trains the $p_\phi(h_{t+1} \mid h_t)$ part directly. Embodied pre-training is then mostly about training $g_\psi$, which is easier given $h_{t+1}$ as an additional conditioning signal. The authors cite Richens et al. (2025) β€” "General agents contain world models" (arXiv:2506.01622) β€” for the claim that agents capable of multi-step goal-directed generalization must effectively learn predictive structure in the environment. (The $p_\phi \cdot g_\psi$ vs. joint-denoising distinction mirrors the paper's own contrast β€” LingBot-VA conditions action on a predicted visual state, whereas Cosmos-Policy and DreamZero jointly denoise both modalities β€” but the equation form is the reviewer's formalization, not verbatim from the paper.)

7.4 RQ 4 β€” How does WAM runtime compare?

Verdict: WAMs are 4.8Γ— to 83Γ— slower than Ο€0.5; even the lightest WAM (Fast-WAM, 190 ms) is 3Γ— slower.

Key drivers (per the paper's Β§4.4):

  • Number of denoising steps for state and action dominates runtime β€” the dominant cost is the future-state diffusion process.
  • State denoising typically dominates over action denoising.
  • Backbone size (2B β†’ 5.3B) is a secondary driver.
  • Generation strategy (joint vs separate action decoding, AR vs single-shot) matters.

The two WAM acceleration strategies in the literature:

  1. Reduce denoising steps (Cosmos-Policy: 5+5; LingBot-VA real-world: 3+5; GE-Act: 1+10; MOTUS: 10+10). Workable for inference, but training-time still expensive.
  2. Skip future-state generation at test time (Fast-WAM: jointly trains video + action but omits video at deployment; GigaWorld-Policy: conditions video on action so video can be skipped). Best WAM latency to date (190 ms β€” Fast-WAM), but still well behind Ο€0.5 (63 ms).

LingBot-VA at the RoboTwin denoising setting (25 state + 50 action steps) is 5.2 seconds per chunk β€” well beyond any plausible deployment frequency.


8. Authors' own conclusions on the WAM-vs-VLA question

Directly from the paper's Β§4 conclusion (paraphrased and consolidated):

  1. WAMs win on visual perturbations (noise, lighting, layout) β€” a pattern that holds across both single-arm and bimanual settings.
  2. WAMs lose on geometric perturbations (camera viewpoint, robot initial state) β€” video priors do not transfer when scene geometry changes.
  3. VLAs can match WAMs if trained with enough diverse data β€” but this is expensive. Ο€0.5's recipe is the empirical existence proof.
  4. Hybrid approaches (MOTUS, VLA-JEPA) sit between β€” the manner in which video priors are integrated matters as much as their presence.
  5. Inference speed remains a major WAM challenge. β‰₯4.8Γ— slower than Ο€0.5 across the board.
  6. The Fast-WAM thesis (no embodied pre-training needed) is conditional β€” it works on RoboTwin with diverse task data but collapses on LIBERO with clean-only data. WAM video priors are necessary but not sufficient.
  7. Open challenge: more efficient exploitation of dynamic priors + improved training/inference efficiency.

The paper is admirably explicit that this is not a one-architecture-wins conclusion β€” it is a paradigmatic comparison with axis-by-axis tradeoffs.


9. Limitations of the study itself

The paper does not have a dedicated limitations section, but several caveats emerge from the methodology:

9.1 WAM coverage gaps

  • DreamZero (14B Wan2.1) is excluded because re-training was prohibitive and warm-up exceeded 15 min β€” but DreamZero is the largest WAM in the field. If scaling effects matter, the study under-samples them.
  • GigaWorld-Policy is excluded because no public checkpoint was available at evaluation time.
  • mimic-video (VAM position piece) is referenced but not in the quantitative comparison.

9.2 VLA coverage gaps

  • The strongest data-diverse VLA in the comparison is Ο€0.5 β€” but Ο€0.6, Ο€0.7 (the actual 2026 production baselines from PI) are not evaluated. Both have improvements (Knowledge Insulation, MEM history, subgoal-image conditioning) that could change the comparison.
  • Newer reasoning-augmented VLAs (ECoT-Lite, dVLA, Embodied-R1, MolmoAct, ABot-M0) are mostly absent from the RoboTwin evaluation.
  • HoloBrain0-GD and ABot-M0 are reported on LIBERO-Plus only, not RoboTwin 2.0-Plus.

9.3 Benchmark coverage gaps

  • Both benchmarks are simulation-only (MuJoCo + SAPIEN). No real-robot rollouts are reported.
  • Both are tabletop manipulation. Mobile manipulation, humanoid loco-manipulation, dexterous, and contact-rich tasks are not tested.
  • The robustness suite (LIBERO-Plus + RoboTwin 2.0-Plus) tests observation-side perturbations primarily. Action-side robustness (action noise, delayed actions, dropped frames) is not tested β€” see RobustVLA for that axis.

9.4 Training-data confounds

  • The Fast-WAM clean-only vs domain-randomized comparison is the only clean A/B in the paper. For all other comparisons, the models come with different training-data mixtures, different backbones, and different objectives β€” disentangling video-backbone effect from data-mixture effect is hard.
  • Ο€0.5's leading LIBERO-Plus score is conflated with its training pipeline (web + mobile + cross-embodiment + post-training). The study cannot isolate which ingredient matters most.

9.5 Evaluation distribution

  • 50 rollouts per task per branch is reasonable but introduces meaningful variance β€” single-digit differences between models may not be statistically meaningful.
  • C2 (spherical camera perturbation) is disabled by default β€” meaning camera perturbations are under-stressed relative to what they could be.
  • Language perturbations use only ~2,500 pre-generated variants (50 tasks Γ— 50 variants on RoboTwin), which may not stress true OOD verb generalization.

9.6 Latency confounds

  • Fast-WAM latency (190 ms) is reported from Yuan et al.'s hardware, not re-measured on the same device as the other models. So the 3Γ— ratio to Ο€0.5 may not be apples-to-apples.

These caveats do not invalidate the central claims but should be kept in mind when citing single numbers.


10. Cross-paper positioning

flowchart TB
  Q[Question: VLA vs WAM at architectural-family level]
  Q --> THIS[This paper Mar 2026<br/>Controlled benchmark on LIBERO-Plus + RoboTwin 2.0-Plus]

  THIS -. extends .-> ROBUST[RobustVLA - ICLR 2026<br/>VLA-only multi-modal robustness]
  THIS -. data-side companion .-> LBM[LBM Co-training - Feb 2026<br/>data modality x strategy ablation]
  THIS -. architecture taxonomy .-> ARCH[Review-VLA-Architecture<br/>Β§5 Category E sub-patterns E1-E5]
  THIS -. Cat E4 strict definition .-> ARCH
  THIS -. WAM evaluation reference for .-> GR00T[GR00T N1.x series review<br/>NVIDIA WAM-influenced VLA]
  THIS -. evaluation reference for .-> COSMOS[Cosmos-Policy ICLR 2026 per-paper]
  THIS -. evaluation reference for .-> GENIE[Genie-Envisioner / GE-Act ICLR 2026 per-paper]

  classDef self fill:#bbdefb,stroke:#1565c0,color:#000
  classDef rel fill:#fff9c4,stroke:#f57f17,color:#000
  class THIS self
  class ROBUST,LBM,ARCH,GR00T,COSMOS,GENIE rel
Loading

vs. GR00T N1 β†’ N1.7 (NVIDIA's evolving WAM-influenced VLA)

GR00T is not classified as a WAM in this paper's vocabulary β€” GR00T retains an Eagle-class VLM backbone with a separate DiT action expert (Category C/F in the architecture review). However, NVIDIA's broader Cosmos / DreamGen / GR00T-Dreams ecosystem is the place where the WAM thread is most actively being integrated into a VLA. This paper provides the cleanest empirical answer to "should N1.x flip to a video-generation backbone?": the WAM perturbation advantage is real on noise/light/layout, but lost on camera/robot-state, and latency is currently a deal-breaker for humanoid control rates. GR00T N1.7's current trajectory (Qwen3-VL backbone + EgoScale data + 32-layer DiT) is more consistent with the data-diverse VLA strategy that paid off for Ο€0.5 in this study than with the pure WAM strategy of Cosmos-Policy.

vs. Review-VLA-Architecture Β§5 Category E (world-model sub-patterns)

The architecture review's E1–E5 sub-pattern taxonomy is complementary to this paper. The review distinguishes how world models are used (auxiliary loss vs data factory vs environment vs backbone vs geometry-first); this paper provides the first controlled benchmark for E4 (VAM/video-backbone) against B (flow-matching) and other VLA families. Where the architecture review is structural, this paper is empirical.

A revised mental model after reading this paper:

  • E1 (auxiliary world-model loss on VLM) β‰ˆ MOTUS / VLA-JEPA in the paper's taxonomy β†’ intermediate robustness.
  • E4 (VAM β€” video-backbone) = the WAMs in the paper (Cosmos-Policy, GE-Act, LingBot-VA, Fast-WAM) β†’ strong visual robustness, weak geometric robustness, slow inference.

vs. LBM Co-training (the data-side empirical companion)

TRI's LBM paper is the data-side controlled study (which data modalities help during co-training); this paper is the architecture-side controlled study (which architectural family is more robust under perturbation). They are complementary β€” TRI keeps the architecture fixed (PaliGemma2-3B + flow-matching) and varies data; Huawei keeps perturbation conditions fixed and varies architecture.

Both papers reach a similar high-level conclusion: data diversity matters at least as much as architectural cleverness. TRI's "VL + cross-embodiment dominates" finding and this paper's "Fast-WAM collapses on clean-only LIBERO data" finding are two faces of the same coin.

vs. RobustVLA (VLA-only robustness study)

RobustVLA (Beihang + PKU + CUHK + Tsinghua, ICLR 2026) evaluates VLAs only under a 17-perturbation suite spanning action / observation / environment / instruction modalities. This paper extends the spotlight to include WAMs and the VLA+WM hybrids. The two papers are best read together:

  • RobustVLA's three findings:
    1. Action perturbations are the most fragile axis.
    2. Visual-robust methods do not transfer across modalities.
    3. Ο€0 > Ο€0-FAST > OpenVLA in robustness (within the VLA family).
  • This paper's complementary findings:
    1. Camera + robot-state perturbations are the WAM-specific fragile axis.
    2. WAM video priors transfer to visual perturbations but not geometric ones.
    3. Ο€0.5 > Ο€0 > Ο€0-FAST > OpenVLA in robustness (and Ο€0.5 matches WAMs on LIBERO-Plus).

A reader picking a 2026 production VLA should consult both β€” RobustVLA for action-side stability, this paper for architecture-family-level visual generalization.

ICRA 2026 developments

The ICRA 2026 cohort (ICRA 2026 Survey) approaches robustness from a deployment-centric angle rather than the architecture-family lens used here, but three threads bear directly on this review's thesis.

Domain-randomized generalization as the practicality axis. Rethinking VLA Practicality introduces CEBench, a single-arm / bimanual / mobile-bimanual benchmark (14.4k sim trajectories over 36 tasks + 1.6k real over 8 tasks) built with explicit domain randomization — clutter, lighting, texture, table height — so that seen vs domain-randomized success can be read off directly. Its 0.5B LLaVA-VLA baseline scores 40.3% seen → 28.6% DR on RoboTwin and 44.2% → 30.7% on real bimanual, quantifying a seen→DR gap on a small data-lean model. This is the cheap-deployment mirror of this paper's central lever: where Huawei shows the WAM video prior is necessary-but-not-sufficient and Fast-WAM collapses on clean-only LIBERO data (97.6%→51.5%), CEBench shows the same data-diversity dependence survives all the way down to a 0.5B backbone trained on a single 4090. Both findings reinforce that domain-randomized / diverse training data — not parameter count or backbone family alone — is the dominant robustness knob.

Adversarial and failure-side robustness. The ICRA "generalization, robustness & security" cluster (ICRA-2026-Topic-VLA Β§9) adds axes this study does not touch: Exploiting Vulnerabilities demonstrates universal adversarial attacks on VLAs (physical-world VLAs are attackable), and SVP (Dual Stochastic Visual Prompting) diagnoses "distracted attention" as shortcut learning in OpenVLA-class models and corrects it without architecture changes. These are an adversarial complement to the natural-perturbation suites (LIBERO-Plus / RoboTwin 2.0-Plus) benchmarked here β€” worst-case rather than distributional robustness.

Clutter as a perturbation dimension. CEBench's clutter randomization and the cohort's clutter-based evaluation protocols extend the Layout/distractor axis of this paper's 7-dimension taxonomy toward denser, more deployment-realistic scenes β€” a natural next stress test for the WAMs that already lead this study on layout (LingBot-VA 87.9% RT, Cosmos-Policy 82.2% LIB).


10b. NVIDIA's thought about WAM (the "imagine β†’ act" blog + Cosmos 3)

NVIDIA's developer blog "Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models" (developer.nvidia.com, June 2026) is NVIDIA's own editorial framing of the WAM thesis β€” and it explicitly cites this very paper (Zhang et al. 2603.22078) as evidence that WAMs can reach strong robustness "without the broader training-data mixture used by VLA baselines." Read alongside Cosmos 3 (launched June 1, 2026), it shows where NVIDIA thinks the field is going.

In brief (full, figure-rich treatment moved to the dedicated page NVIDIA WAM thesis + Cosmos 3):

  • The blog's thesis β€” a WAM "starts from a pretrained world-model or video backbone and adapts it to … emit corresponding actions"; three paradigms (inverse-dynamics Β· joint-prediction Β· representation-only) matching Β§4.1 here. NVIDIA's datapoint: DreamZero (2602.15922) scores RoboArena 1750 vs Ο€0.5's 1622, trained only on DROID. It concedes the costs this study measures β€” ~590–800 ms vs ~190 ms (3–4Γ—) latency, ~10Γ— longer sequences, open grounding gap β€” and bets the "next generation will be WAM+VLA hybrids" (an editorial prediction, not a roadmap).
  • Cosmos 3 (Jun 1 2026) β€” a single Mixture-of-Transformers omnimodel pairing a reasoning transformer with an expert generation transformer ("think before it acts"), with native action output (joint angles / gripper / waypoints), three tiers (Super / Nano / Edge, OpenMDW 1.1), and three developer roles (VLM Β· world model Β· WAM backbone). Scale and "ranks first" benchmarks are vendor claims with no absolute scores.

C. How this reframes this review's verdict

This study's core finding is asymmetric: WAMs win visual robustness (noise/light/layout) but lose camera-viewpoint, robot-state, and language grounding, and pay a 3–4Γ— latency tax (Β§7). NVIDIA's blog concedes every one of those costs and does not claim WAMs dominate β€” it argues the thesis is defensible and the future is hybrid. Cosmos 3's design reads as a direct response to exactly the legs WAMs were losing here:

  • The grounding/semantic leg (where pure video WAMs were weakest, and VLAs strongest) β†’ Cosmos 3 adds a reasoning transformer in front of generation, re-introducing the VLM-style semantic grounding a raw video backbone lacks.
  • The "one model, both strengths" bet β†’ Cosmos 3's MoT is the WAM+VLA hybrid the blog predicts, collapsed into a single pretraining rather than a bolted-on VLA+WM (contrast MOTUS / VLA-JEPA's intermediate ~70–78% in Β§5–§6).

Caveats (keep the skeptic's hat from Β§9). Cosmos 3's claims β€” the data scale, the "ranks first" on RoboArena/RoboLab/Physics-IQ, the "fraction of a second" Nano latency β€” are launch-PR / vendor statements with no absolute numbers and no peer review, and there is no independent LIBERO-Plus / RoboTwin-2.0-Plus robustness eval of Cosmos 3 yet β€” precisely the controlled apples-to-apples test this paper argues the field still needs. In particular, the Β§7 finding that WAMs lose camera-viewpoint and robot-state robustness and pay a 3–4Γ— latency tax is not addressed by any Cosmos 3 number on the table; the "ranks first on RoboArena" claim uses a different protocol from this study's perturbation suites, and the blog itself flags benchmark-gaming risk (it calls for RoboLab). The strongest measured NVIDIA WAM result remains DreamZero's RoboArena 1750. So the honest read: NVIDIA's "thought about WAM" matches this paper's evidence β€” WAM video priors are real but necessary-not-sufficient β€” and Cosmos 3 is its biggest bet that a reasoning-augmented MoT (plus a low-latency Nano tier) can buy WAM robustness and VLA grounding at once. Whether it actually closes the viewpoint / state / latency gaps this review identified is, as of June 2026, unverified.

Sources: NVIDIA developer blog "Pretrained to Imagine, Fine-Tuned to Act" (Jun 2026) Β· Cosmos 3 newsroom Β· Cosmos 3 blog Β· Cosmos 2501.03575 Β· Cosmos Policy 2601.16163 Β· DreamZero 2602.15922.


11. The "is the answer WAM or VLA?" question, plainly

If forced to pick a single sentence answer: mixed-with-asymmetry.

  • WAMs win when:

    • You need visual perturbation robustness (noise, lighting, layout, background) out of the box.
    • You have minimal task-specific data (Cosmos-Policy: 185 trajectories).
    • You have moderate compute and inference latency is not deployment-critical.
    • You have access to a strong web-video-pretrained backbone (Cosmos-Predict2, Wan2.2, LTX-Video).
  • VLAs win when:

    • You need camera-viewpoint or robot-state robustness.
    • You can afford the diverse-data training pipeline (web + cross-embodiment + post-train).
    • Inference latency is a hard requirement (≀100 ms control rate).
    • You need verb / language generalization that the VLM's NL grounding provides.
  • Hybrid (VLA + WM) wins when:

    • You want a portion of the WAM robustness without giving up the VLA pipeline.
    • You can engineer a working integration (MOTUS's MoT or VLA-JEPA's predictive-encoder alignment).
    • You're willing to live with intermediate robustness (~70–78%) for moderate cost.

The 2026 architectural meta-question β€” will video-generation backbones replace VLM backbones for embodied AI? β€” is not yet answered by this paper. The paper does show the WAM thesis is empirically defensible on visual robustness, but it does not show WAMs dominate when latency and geometric generalization are required. The next round of comparisons (Ο€0.7 vs LingBot-VA on real bimanual, GR00T N1.7 vs Cosmos-Policy on humanoid loco-manip, dexterous tasks) will likely decide it.

πŸ—“ IROS 2026 update β€” the "versus" is dissolving into "WAM-as-scaffold" πŸ†•

The IROS 2026 survey full-coverage set shows the debate framing shifting. Across IROS 2026's representative WAM papers, the world model is pushed out of the inference loop and into the training/representation scaffold β€” none deploys the WM as the policy:

  • AtomVLA β€” WM as an offline-RL critic (scores action chunks, absent at deployment).
  • DreamMimic β€” WM (RSSM) as distillation supervision + representation; the deployed student is a plain vision policy.
  • Scaling Cross-Embodiment WM β€” WM as a shared cross-embodiment interface feeding classical MPC, not an end-to-end VLA.

So the emerging answer to "WAM or VLA?" is neither: the WAM is being absorbed as scaffolding (critic Β· distillation teacher Β· shared representation Β· MPC model) around a reactive VLA/policy β€” which sidesteps this study's latency verdict by construction (the generative WM never runs at control rate). The in-path generative-WAM backbone (DreamZero / Cosmos-Policy) is now the minority position. Full treatment: VLA Hybrid Architectures Β§4.2b (Axis 2 gained a "critic-only" mode) and IROS 2026 survey Β§5.1.


12. Links


13. Pointers to related pages

  • VLA Architectures review β€” Β§5 Category E (world-model / VAM sub-patterns E1–E5) β€” this paper is the empirical reference for E4.
  • GR00T N1 β†’ N1.7 β€” NVIDIA's evolving humanoid VLA; uses Cosmos backbone + DreamGen data but keeps VLM-style architecture rather than going full WAM.
  • RobustVLA β€” VLA-only robustness sibling.
  • Rethinking VLA Practicality β€” CEBench seenβ†’domain-randomized benchmark + 0.5B LLaVA-VLA baseline; the deployment-centric companion to this study's data-diversity lever.
  • ICRA 2026 Survey Β· ICRA 2026 VLA topic analysis β€” adversarial-attack, SVP attention-shortcut, and clutter-based robustness threads (Β§9 / Β§11).
  • LBM Co-training Study β€” data-side empirical companion (TRI).
  • Cosmos-Policy β€” the strongest WAM on LIBERO-Plus in this paper.
  • Genie-Envisioner / GE-Act β€” the LTX-Video-2B-backed WAM (80.3% LIBERO-Plus).
  • Review-VLA-Attention β€” for the attention-architecture details of LingBot-VA's released unified-transformer variant.
  • CVPR-2026-LIBERO-Plus β€” the LIBERO-Plus benchmark paper (Fei et al. 2025) that this paper extends to RoboTwin 2.0.

← Back to Home Β· ICLR-2026 Β· Review-VLA-Architecture

⚠️ **GitHub.com Fallback** ⚠️