Review NVIDIA WAM Cosmos3 - Heungwoo/research GitHub Wiki
A self-contained bundle of (a) NVIDIA's developer blog "Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models" (developer.nvidia.com, Jun 2026) and (b) a detailed analysis of Cosmos 3 (NVIDIA, launched Jun 1 2026). Companion to the empirical study in WAM vs VLA Robustness (the blog cites that very paper) and to World Models.
Figure note: the diagrams below are original schematics recreated to convey the blog's and Cosmos 3's structure — NVIDIA's own copyrighted figures are linked, not embedded. Data tables reproduce only factual numbers/claims from the sources.
NVIDIA's position, in one line: "pretrained to imagine, fine-tuned to act." Start from a web-video world model that already knows how scenes change, then adapt it to emit actions — because predicting future visual change is strongly correlated with the motor commands that cause it, so a video prior offloads what scarce robot data can't teach. Cosmos 3 is NVIDIA's frontier realization: a single Mixture-of-Transformers (MoT) omnimodel that reasons about a scene before generating both video and action trajectories — i.e. NVIDIA collapsing the "WAM + VLA hybrid" it predicts as the field's future into one pretraining.
Honest framing up front: the blog is a technical-editorial survey (it discusses non-NVIDIA models too) and concedes WAMs' costs (3–4× latency, open grounding gap); Cosmos 3's headline numbers are launch-PR / vendor claims with no absolute scores or peer review.
"a policy that starts from a pretrained world-model or video backbone and adapts it to represent or predict how the scene changes over time and emit corresponding actions."
The contrast with a VLM-based VLA: a VLA grounds language → action directly; a WAM inserts an imagine stage — predict the visual future, then act.
flowchart LR
L[language instruction] --> WM[Pretrained video / world model<br/>imagine how the scene changes]
O[current observation] --> WM
WM --> ACT[emit corresponding actions]
2.2 The three WAM paradigms (= the family axis in Review-WAM-vs-VLA-Robustness §4.1)
flowchart TB
Q{How are actions produced?}
Q -- generate future frames then infer --> P1[Inverse dynamics<br/>video model + IDM head · LingBot-VA]
Q -- denoise video and actions together --> P2[Joint prediction<br/>one monolithic DiT · DreamZero]
Q -- use features only · skip generation --> P3[Representation-only<br/>Fast-WAM]
| Paradigm | Mechanism | Example | Trade |
|---|---|---|---|
| Inverse dynamics | video model predicts future frames; a separate IDM head reads actions off the transition | LingBot-VA | modular; two stages |
| Joint prediction | one model denoises future video + action tokens in a single pass | DreamZero | unified; heaviest sequence |
| Representation-only | use the video backbone for features; skip generation at inference | Fast-WAM | fast; weaker world prior |
| Method | Idea | Example |
|---|---|---|
| Action-as-token | actions are a new modality alongside video tokens | many WAMs |
| Action-as-image | encode action / proprioception / value as synthetic latent frames | Cosmos Policy |
| Latent plans / actions | compress behaviour into learned latent abstractions | latent-action WAMs |
DreamZero (NVIDIA; adapts Wan 2.1-I2V-14B, 2602.15922; joint-prediction monolithic DiT) is the blog's "signal for WAMs" — it beats strong VLAs while trained only on DROID (~50k demos), with no large cross-embodiment stage:
| Model | RoboArena (Apr 2026) |
|---|---|
| DreamZero (WAM) | 1750 |
| π0.5 (VLA) | 1622 |
| π-FAST (VLA) | 1592 |
| π0 (VLA) | 1475 |
(GR-1 on CALVIN ABC→D: 3.06/5 avg sequence length, cited separately.)
| Cost | Figure |
|---|---|
| Sequence length | video token sequence ~10× longer than a VLA's |
| Inference latency | WAM ~590–800 ms / action chunk vs ~190 ms for π0.5 → 3–4× slower |
| Training compute | DreamZero action-tuning alone ~9 ZFLOP (+ ~51 ZFLOP for video pretrain) |
| Grounding gap | language→physical-action grounding "remains open" even for modern VLAs |
The blog's forward-looking conclusion — "the likely next generation of robot foundation models will be WAM+VLA hybrids", canonically a three-tower Mixture-of-Transformers (understanding/VLM + video-generation + action expert, exchanging via shared self-attention) — naming, as modular/bolted-on: π0.7's BAGEL subgoal images, Cortex 2.0 (Sereact, foresight trajectory scoring); as unified: MOTUS, BagelVLA, Being-H0.7; plus robotics-first foundation models and latent world models (V-JEPA 2) for cheaper rollouts. See VLA Architectures §4.2b for the full hybrid taxonomy; the reactive corner (DYNA-2, ω-0) drops the video tower at inference.
Attribution caveat. This "future is hybrid" line is the developer blog's editorial prediction (NVIDIA authors synthesizing the field — most hybrid exemplars it cites are non-NVIDIA), not a binding corporate-roadmap statement. Notably, NVIDIA's own product move (Cosmos 3, below) is a single-model unification, not the bolted-on VLA+WM the cited examples illustrate.
| Stage | Model | Role |
|---|---|---|
| Imagine | Cosmos / Cosmos-Predict (2501.03575) | web-scale video world model — pure future-frame prediction |
| Fine-tune to act | Cosmos Policy (Cosmos-Predict2-2B, 2601.16163) | action-as-latent-frame WAM; 185-trajectory data efficiency; a top performer in Review-WAM-vs-VLA-Robustness |
| Unify reason + imagine + act | Cosmos 3 (Jun 1 2026) | single MoT omnimodel — §3.2 onward |
NVIDIA describes Cosmos 3 as a Mixture-of-Transformers pairing a reasoning transformer with an expert generation transformer: reasoning "first interprets what is happening in a scene" (object interactions, motion, spatial-temporal relations); generation then "uses that context to create physically grounded outputs" — video and action trajectories. This is the blog's monolithic joint-prediction WAM, but with a semantic-reasoning stage in front — NVIDIA's literal "think before it acts," and the architectural opposite of a pure video backbone that denoises frames with no semantic stage.
flowchart LR
IN[multimodal input<br/>text · image · video · sound] --> R[Reasoning transformer<br/>interpret scene · object interaction · motion · space]
R --> G[Expert generation transformer<br/>physically grounded outputs]
G --> OV[video / synthetic data]
G --> OA[action trajectories<br/>joint angles · gripper · waypoints]
| Tier | Purpose | Status |
|---|---|---|
| Cosmos 3 Super | post-training robotics / AV models needing the highest physics accuracy & generation quality | available now |
| Cosmos 3 Nano | high-quality video + action reasoning "in a fraction of a second" — the latency-conscious tier | available now |
| Cosmos 3 Edge | real-time inference at the edge | coming soon |
Released under OpenMDW 1.1 (Linux Foundation); available on Hugging Face / GitHub / build.nvidia.com / as NVIDIA NIM microservices.
flowchart TB
C[Cosmos 3<br/>one pretraining] --> RoleA[Role 1 · VLM<br/>reason across modalities]
C --> RoleB[Role 2 · World / video model<br/>simulate + predict future states · train & eval data]
C --> RoleC[Role 3 · WAM backbone<br/>train robots for specific tasks]
That a single pretraining serves all three roles — the grounding leg (VLAs' strength in the robustness study), the synthetic-data / evaluation leg (cf. DreamGen, GR00T-Dreams), and the WAM leg — is precisely the WAM+VLA convergence the blog predicts, here as a unified model rather than a bolted-on hybrid.
Unlike a pure video WAM that needs a separate inverse-dynamics head, Cosmos 3 claims native numeric action output — "joint angles, gripper positions and trajectory points." Cited early robotics uses: Agile Robots ("generate action-conditioned robot data for policy development") and the NVIDIA GEAR team (video action models for embodied agents).
| Item | Claim | Provenance |
|---|---|---|
| Training data | "billions of samples across text, image, video, sound and action trajectories" | NVIDIA newsroom |
| (third-party) | ~20T tokens / ~1B images / 400M videos | press coverage — not on NVIDIA's page |
| World generation | "ranks first" — Physics-IQ, PAI-Bench, R-Bench, Artificial Analysis | NVIDIA (no absolute scores) |
| Action policy | "ranks first" — RoboLab, RoboArena | NVIDIA (no absolute scores) |
| Vision understanding | "ranks first" — VANTAGE-Bench, TAR | NVIDIA (no absolute scores) |
| Cosmos Coalition | Agile Robots, Black Forest Labs, Generalist, LTX, Runway, Skild AI | NVIDIA |
The empirical study (Review-WAM-vs-VLA-Robustness, Zhang et al. 2603.22078) — which the blog cites — finds an asymmetric result: WAMs win visual robustness (noise/light/layout) but lose camera-viewpoint, robot-state, and language grounding, at a 3–4× latency cost. Cosmos 3 reads as a direct answer to the legs WAMs were losing:
- Grounding/semantic leg → a reasoning transformer in front of generation re-introduces the VLM-style grounding a raw video backbone lacks.
- Latency leg → the Nano tier explicitly targets sub-second action reasoning.
- "Both strengths in one model" → the MoT is the predicted WAM+VLA hybrid, collapsed into one pretraining (contrast the bolted-on MOTUS / VLA-JEPA at intermediate ~70–78% robustness).
- Cosmos 3's data scale, "ranks first" benchmarks, and Nano's "fraction of a second" latency are launch-PR / vendor statements with no absolute numbers and no peer review.
- There is no independent LIBERO-Plus / RoboTwin-2.0-Plus robustness eval of Cosmos 3 — exactly the controlled, apples-to-apples test the cited study argues the field still needs. The §7 viewpoint / state / latency gaps are not addressed by any Cosmos 3 figure on the table.
- The strongest measured NVIDIA WAM result remains DreamZero's RoboArena 1750, on a different protocol from the perturbation suites.
- The "future is hybrid" line is an editorial prediction, not a roadmap (see §2.6).
- Net read: NVIDIA's WAM thesis matches the study's evidence — video priors are real but necessary-not-sufficient — and Cosmos 3 is its biggest bet that a reasoning-augmented MoT (+ low-latency Nano) can buy WAM robustness and VLA grounding at once. Whether it closes the identified gaps is, as of June 2026, unverified.
- NVIDIA developer blog: "Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models" — developer.nvidia.com
- Cosmos 3: newsroom launch · NVIDIA blog
- Papers: Cosmos 2501.03575 · Cosmos Policy 2601.16163 · DreamZero 2602.15922 · robustness study 2603.22078
- Wiki companions: Review-WAM-vs-VLA-Robustness · Review-World-Models · Cosmos Policy · GR00T Series · DreamGen
← Back to Home