Review DreamZero - Heungwoo/research GitHub Wiki
In-Depth Review — DreamZero: World Action Models are Zero-shot Policies
Paper: World Action Models are Zero-shot Policies · NVIDIA (× UC Berkeley, CMU collaborators) Authors: Seonghyeon Ye†, Yunhao Ge*, Kaiyuan Zheng*, Shenyuan Gao*, Sihyun Yu*, … Yuke Zhu†, Linxi "Jim" Fan†, Joel Jang† (project leads †, core contributors *) arXiv: 2602.15922 (v1 Feb 17, 2026, cs.RO) · Project: dreamzero0.github.io · Code/weights: github.com/dreamzero0/dreamzero (open-sourced) OpenReview: https://openreview.net/forum?id=cd33uUB609
Companion reviews: World Models · WAM vs VLA Robustness · mimic-video · LDA-1B · Real-Time Execution · Human Video → Robot Transfer · GR00T series.
1. TL;DR
- The flagship "World Action Model" (WAM). DreamZero is a 14B robot foundation model built on a pretrained image-to-video diffusion backbone (Wan2.1-I2V-14B) that jointly predicts future video and continuous actions. It reframes control as inverse dynamics — align motor commands with a predicted visual future — rather than dense state→action imitation, and coins WAM (over "VAM") to signal that video is only one possible world-modeling target (tactile/force/latent futures could follow).
- The headline result: >2× generalization over SOTA VLAs. On a real AgiBot G1 bimanual mobile manipulator evaluated in unseen environments with unseen objects, DreamZero reaches 62.2% average task progress on seen-task categories vs 27.4% for the best pretrained VLA baseline (GR00T N1.6 / π0.5) — and both baselines were pretrained on thousands of hours of cross-embodiment robot data while DreamZero trained only on ~500 h. On 10 fully unseen tasks, DreamZero hits 39.5% vs 16.3% (pretrained VLA) and <1% (from-scratch VLA).
- Real-time from a 14B video diffusion model. A stack of optimizations — asynchronous closed-loop execution, CFG parallelism, DiT caching, torch.compile+CUDA graphs, NVFP4 quantization, and the DreamZero-Flash decoupled noise schedule — delivers a 38× speedup (5.7 s → 150 ms per chunk), enabling 7 Hz closed-loop control. This is the paper's systems contribution and what makes WAMs deployable at all.
- Data-efficient cross-embodiment transfer — from video only. Video-only demonstrations from another robot (YAM, 20 min) or humans (12 min) lift unseen-task progress from 38.3% to 55.4% / 54.3% (relative +42%+). And a model pretrained on AgiBot G1 adapts to an entirely new robot (YAM) with only 30 minutes of play data while retaining zero-shot generalization — a new bar for data-efficient embodiment adaptation.
- "Improving robotics reduces to improving video generation." The paper's central empirical claim: larger video backbones → better video prediction → better actions (14B beats 5B 50% vs 21%; VLAs at 5B/14B/32B all score ~0% on the same diverse data), and most DreamZero failures are video-generation errors, not action-extraction errors — the policy faithfully executes whatever the video predicts.
2. Why this paper matters
- It is the strongest hardware evidence yet for the WAM thesis. The 2026 world-model-vs-VLA debate (see [Review-World-Models]], [[Review-WAM-vs-VLA-Robustness]]) had challengers — [mimic-video (10× sample efficiency), LDA-1B (+48% dexterous) — but DreamZero pushes the argument to a 14B, real-time, open-weights system on a mobile bimanual robot with an explicit ablation showing VLAs cannot learn its diverse data at any scale (0% at 5B/14B/32B). It is the reference the wiki cites (Ye et al. 2026) whenever WAMs-as-policies come up, e.g. in Qwen-RobotWorld's framing.
- It attacks the exact axis VLAs are weakest on: new motions, not new objects. VLAs inherit VLM semantics ("move coke can to Taylor Swift") but "lack representations of how actions should be executed." DreamZero's task granularity is defined by motion + object type, so "fold socks" is unseen relative to "fold shirt" — a deliberately harder generalization axis than the object/scene generalization VLAs usually report.
- It reframes data collection. Because video prediction is inherited from web-scale pretraining and only the inverse-dynamics mapping must be learned from robot data, DreamZero learns better from diverse, non-repetitive trajectories (500 h across 22 environments, ~4.4 min / ~42 subtasks per episode) than from repetitive task-focused demos — inverting the VLA convention. Ablation: diverse 500 h beats repetitive 500 h, 50% vs 33%.
- It makes the video-only cross-embodiment path concrete, connecting directly to the human-video transfer fork: WAMs can absorb human/other-robot experience without action labels, opening the "abundant human video, no teleop" scaling pathway that the emergence/decoupling/synthesis camps all chase from different angles.
3. Architecture

Figure 4 of the paper. Left (training): three inputs — video (VAE-encoded), language (text encoder), proprioception (state encoder) — feed a causal DiT backbone that jointly denoises noisy video and action latents under a shared flow-matching objective, with teacher forcing (denoise the current chunk conditioned on clean previous chunks). Right (inference): the same backbone runs closed-loop with a KV cache; after each action chunk executes asynchronously in the real world, ground-truth observations replace the predicted frames in the KV cache ("Update with Real Observation"), which eliminates the compounding error of free-running autoregressive video rollout while preserving native frame rate for tight video-action alignment.
3.1 The WAM formulation
DreamZero models the joint distribution of video and action, which factorizes exactly into autoregressive video prediction × an inverse-dynamics model:
π(o_{l:l+H}, a_{l:l+H} | o_{0:l}, c, q_l) = π(o_{l:l+H} | o_{0:l}, c, q_l) · π(a_{l:l+H} | o_{0:l+H}, q_l) ⌊ video prediction ⌋ ⌊ IDM ⌋
Crucially it trains one end-to-end model for both factors (not a separate video model + IDM as in some prior WAMs), for deep cross-modal integration. Since the video-prediction half is largely inherited from web-scale pretraining, "DreamZero only needs to additionally learn to predict videos for the robot embodiment and extract corresponding actions from the generated videos."
3.2 Design choices
- Backbone: Wan2.1-I2V-14B-480P image-to-video diffusion transformer. Only state/action encoders and action decoder are added (minimal new params); text encoder, image encoder, and VAE are frozen; all DiT blocks are updated. Multi-view robot inputs are concatenated into one frame rather than changing the backbone. (LoRA was tried and gave suboptimal results.)
- Autoregressive, not bidirectional. AR enables KV-cache reuse (3–4× faster inference), lets the policy use visual history as guidance, preserves native FPS (bidirectional diffusion needs fixed-length windows → subsampling that distorts FPS and hurts video-action alignment), and produces substantially smoother motions. AR is applied only to the video modality to avoid closed-loop action error propagation. Task progress is similar to a bidirectional variant (50% vs 50%) but AR wins on smoothness + speed.
- Chunk-wise teacher-forced flow matching. Video is predicted in chunks of K latent frames matched to the action horizon; a shared denoising timestep across video and action speeds early convergence; teacher forcing denoises the current chunk from clean prior chunks (enabling variable-length training like LLM token training).
- Closed-loop KV-cache correction. The defining WAM trick: replacing predicted frames with real observations in the KV cache after each execution kills the autoregressive-video compounding-error problem — an advantage unavailable to pure video generators.
3.3 Real-time execution — the 38× stack
| Layer | Optimization | Effect |
|---|---|---|
| Structure | Asynchronous closed-loop — controller runs the latest chunk while inference runs concurrently | Target: latency < ~200 ms (chunk = 48 steps @ 30 Hz = 1.6 s) |
| System | CFG parallelism (2 GPUs) | −47% per-step latency |
| System | DiT caching (reuse velocities when cosine-similar) | 16 → 4 effective DiT steps |
| Impl. | torch.compile + CUDA graphs; cuDNN attention; GPU-side scheduler | up to ~10.9× (GB200) |
| Impl. | NVFP4 quantization (QKV/Softmax in FP8, non-linear in FP16) | 16.6× cumulative (GB200) |
| Model | DreamZero-Flash | 38× cumulative → 150 ms |
DreamZero-Flash is the model-level key: standard training couples video and action to the same noise level, but few-step inference must predict clean actions from still-noisy video. Flash fixes the train-test mismatch by biasing video timesteps toward high noise (t_video = 1−η, η~Beta(7,1), E[t_video]=0.125) while keeping action timesteps uniform — training the model to extract clean actions from noisy visual context. This cuts denoising from 4 steps to 1 with far less quality loss: on table bussing, 1-step DreamZero collapses 83% → 52%, but Flash holds 74% at 1 step (2.33× faster). Cumulative speedups: ~9× (H100 system+impl), ~16× (GB200), 38× (GB200 + Flash); everything except DiT caching and quantization is mathematically loss-free.
4. Data & training
- Pretraining data (AgiBot G1): ~500 hours / 7,193 episodes teleoperated across 22 environments (homes, restaurants, supermarkets, coffee shops, offices), ~4.4 min / ~42 subtasks per episode — deliberately long-horizon and diverse over repetitive. Also validated on Franka via DROID (one of the most heterogeneous public datasets) for reproducibility.
- Training: 100K steps, global batch 128, per embodiment (separate pretraining; multi-embodiment left to future work). Idle actions filtered; relative joint positions as default action representation.
- Evaluation philosophy — "first-citizens of generalization evals": the default setting is unseen environment + unseen objects (eval sites are in a different geographic location than training), so every benchmark tests OOD generalization. Baselines: GR00T N1.6 and π0.5, each in from-scratch (fair, same VLM weights, same data) and from-pretrained (official cross-embodiment checkpoints) variants, all trained on identical data with matched compute.
5. Results
5.1 Generalization (real-robot, unseen env + objects)
| Setting | DreamZero | Best pretrained VLA | From-scratch VLA |
|---|---|---|---|
| Seen tasks — AVG task progress (AgiBot G1) | 62.2% | 27.4% | ~0% |
| Unseen tasks — AVG task progress (AgiBot G1) | 39.5% | 16.3% | <1% |
| DROID-Franka — unseen, task progress / success | 49% / 22.5% | 31% / 12.5% (GR00T), 33% / 7.5% (π0.5) | — |
Per-category on seen tasks: PnP-Easy 93.8, PnP-Hard 48.4, Contact-Rich 49.0 (DreamZero) — from-scratch VLAs near-zero across all. Standout unseen tasks: Remove Hat from Mannequin 85.7%, Shake Hands 59.2%. Qualitatively, pretrained VLAs "reach toward objects and attempt grasping regardless of instruction" (overfitting to pick-and-place), while DreamZero does visual planning per task.
5.2 Post-training retention
After task-specific fine-tuning (shirt folding 33 h, fruit packing 12 h, table bussing 40 h), DreamZero matches or beats VLA baselines and retains environment generalization — outperforming SOTA VLAs by ~10% average task progress, biggest gap on fruit packing. The point: world-modeling generalization survives specialization, which VLA post-training typically erodes.
5.3 Cross-embodiment transfer (video-only)
| Method | Unseen-task progress |
|---|---|
| DreamZero (baseline) | 38.3% ± 7.6% |
| + Human→Robot transfer (12 min video-only) | 54.3% ± 10.4% |
| + Robot→Robot transfer (YAM, 20 min video-only) | 55.4% ± 9.5% |
Transfer uses only the video-prediction objective on the cross-embodiment data (no action labels), co-trained 1:1 with pretraining data for 10K steps. Robot→robot edges out human→robot (narrower embodiment gap; both bimanual parallel grippers), but both give >42% relative gains from ≤20 min of video.
5.4 Few-shot new-embodiment adaptation
DreamZero-AgiBot adapts to a new robot (YAM) with 55 trajectories / 11 tasks / ~30 min of play data, retaining strong language following and generalizing to novel objects (pumpkins, teddy bears, cup noodles). The paper credits (1) visual similarity of the two grippers and (2) the hypothesis that learning an implicit IDM from predicted video is inherently more sample-efficient than direct policy learning.
5.5 Ablations (PnP-Easy, 50K steps / batch 32)
| Axis | Result |
|---|---|
| Data diversity | Diverse 500 h 50% vs repetitive 500 h 33% |
| Model scale | 14B 50% vs 5B 21% (smaller model hallucinates → bad actions); VLAs 0% at both 5B and 14B (and 32B) |
| Architecture | AR 50% vs BD 50% task progress — but AR is smoother + 3–4× faster |
6. Significance & positioning
- Where it sits among 2026 WAMs. DreamZero explicitly differentiates from prior WAMs by (a) systematically exploiting data diversity and scale, (b) the autoregressive + KV-cache-correction design for long-horizon closed loop, (c) SOTA cross-embodiment transfer both directions, and (d) real-time 14B deployment. Versus mimic-video (video backbone + separate IDM decoder) and LDA-1B (unified WM in DINO latent space), DreamZero is the single-model, pixel-video, real-robot, largest-scale point in the design space.
- It sharpens the World Models review verdict. The "VLM backbones are blind to physical causality" thesis now has its strongest datapoint: at matched data and compute, VLAs score ~0% where a WAM scores 50–62% — on diverse data specifically. The corollary "improving robotics = improving video generation" gives the WAM camp a clean scaling story (bigger/better video model → better policy) that VLAs lack for the action head.
- Open weights + open eval. Model, inference code, and PolaRiS / Genie-Sim-3.0 eval harnesses are released — a rare fully-open flagship in a field where π and the Qwen suite stay closed. It even shows non-trivial Genie-Sim-3.0 (100 tasks) performance without training on the 10k h of sim data, from only ~500 h real.
7. Limitations
7.1 Authors' stated
- Single embodiment per model. Pretraining is per-robot (AgiBot G1 or Franka); multi-embodiment joint training is left to future work.
- Human-transfer scale is tiny. The cross-embodiment human experiments use only ~12 min of in-lab egocentric data; large-scale in-the-wild human video is hypothesized to help but untested.
- Memory tasks not evaluated. Although the stateful KV-cache design supports history, no memory-requiring task is evaluated or post-trained.
- No scaling laws yet. Diversity and backbone size help, but WAM scaling laws (model × data × compute) are unquantified.
7.2 Reviewer's concerns
- Absolute success rates are still moderate. Unseen-task progress of 39.5% and cross-embodiment 54–55% are relative wins over near-zero VLAs; task-completion success (not partial progress) on hard/contact-rich tasks remains low — the headline "2×" is over weak baselines on OOD, not near-ceiling numbers.
- Failures bottleneck on video generation. The paper frames "most failures are video errors" as a positive (improving the backbone helps), but it also means DreamZero inherits every video-model weakness — contact-rich physics, deformable objects, and long-horizon coherence are exactly where video diffusion is weakest.
- Compute/serving cost. 7 Hz needs a 14B DiT with multi-GPU CFG parallelism and Blackwell NVFP4; the 38× speedup is impressive but the floor is still a data-center-class inference budget, unlike a 2–4B VLA on a single edge GPU.
- The generalization protocol is the paper's own. "First-citizens of generalization evals" is admirable, but seen/unseen granularity (motion+object) and the geographic-split OOD setting are self-defined; cross-paper comparison awaits shared protocols like PolaRiS (which the authors do release evals for).
- VLA-0% ablation may be pessimistic for VLAs. Truncating 8B/32B VLMs to half their blocks + attaching DiT heads is a reasonable scale-match but not how production VLAs are built; the 0% result is a strong claim resting on a constructed baseline.
8. Links & related pages
- arXiv: https://arxiv.org/abs/2602.15922 · Project: https://dreamzero0.github.io · Code/weights: https://github.com/dreamzero0/dreamzero · OpenReview: https://openreview.net/forum?id=cd33uUB609
- World Models — the taxonomy DreamZero anchors as the flagship WAM-as-policy
- WAM vs VLA Robustness — the controlled robustness comparison this result extends
- mimic-video · LDA-1B — the two RSS 2026 WAM/VAM siblings
- Real-Time Execution — where DreamZero-Flash / async closed-loop belong
- Human Video → Robot Transfer — the video-only transfer path DreamZero validates
- Qwen-RobotWorld — the generation-side (language-actioned) world model; DreamZero is the policy-side counterpart
- GR00T series · π0.5 — the VLA baselines it more-than-doubles