Review World Models - Heungwoo/research GitHub Wiki

In-Depth Review — World Models for Robot Learning

Compiled June 2026 · Focus: learned models that predict the future of the world (pixels, latents, geometry, or reward) and how 2024–2026 robotics uses them — as policy backbones, RL environments, data factories, planners, and evaluators.

This review takes a model-centric lens: what the world model is and predicts. It is the sibling of VLA Architectures §5 Category E, which slices the same papers by how the world model is wired into the policy (the E1–E5 usage taxonomy). Read that page for the action-decoder view; read this one for the world-model view. Companion reviews: WAM vs VLA Robustness · Goal-Image Conditioning · RL for VLA · Cross-Embodiment.


1. Why a world model — and the "one model, three roles" idea

A world model (WM) answers: given the current observation and an action, what happens next? For robotics that single capability unlocks four otherwise-expensive things:

  • Data without robots — roll the WM forward to synthesize training trajectories (the teleop-replacement bet).
  • RL without real rollouts — treat the WM as a simulator; optimize the policy against imagined returns.
  • Evaluation without a test rig — rank policy checkpoints by imagined success before touching hardware.
  • Planning without a hand-built model — search action sequences toward an imagined goal.

The strong claim several 2026 papers make is convergence: the same action-conditioned video model can serve as RL environment (WMPO), evaluator (WorldGym, Ctrl-World), and even the policy backbone (Cosmos Policy). World-model quality, not policy architecture, becomes the lever. This is why Category E is the largest growth area in the ICLR 2026 VLA cohort.


2. Two orthogonal axes

flowchart TB
  subgraph WHAT["AXIS 1 — what the model predicts / represents"]
    P1["Pixels<br/>video diffusion"]
    P2["Discrete tokens<br/>autoregressive"]
    P3["Latent embeddings<br/>JEPA / feature-space"]
    P4["3D / 4D geometry<br/>pointmaps · splats · flow"]
    P5["Structured cues<br/>mask + depth + features"]
  end
  subgraph HOW["AXIS 2 — how robotics uses it"]
    U1["Policy backbone (VAM)"]
    U2["RL environment"]
    U3["Data factory"]
    U4["Planner / MPC"]
    U5["Evaluator"]
    U6["Auxiliary loss"]
  end
  WHAT --> HOW
Loading

The two axes are largely independent: a pixel video model can be a backbone (Cosmos-Policy), an RL env (WMPO), a data factory (DreamGen), or an evaluator (WorldGym); a latent model can plan (DINO-WM) or back a policy (CoWVLA). §3 organizes by Axis 1 (the model-centric contribution of this page); §4 maps onto Axis 2 (cross-referencing Category E).


3. The model families (Axis 1 — what it predicts)

A. Pixel video-diffusion world models (the dominant family)

Fine-tune an internet-scale video generator (Cosmos, Wan2.x, LTX-Video, OpenSora, Stable Video Diffusion, DynamiCrafter, CogVideoX) into an action-conditioned future-frame predictor. The defining move of 2024–26 is "video generation → world model": add action/pose conditioning + long-horizon consistency + a downstream control role.

  • General WFMs: NVIDIA Cosmos (platform paper arXiv 2501.03575; Predict / Transfer / Reason branches; action-conditioned Predict2 variants) and DeepMind Genie 2 (Dec 2024, playable 3D worlds from one image) → Genie 3 (Aug 2025, real-time 24 fps / 720p, minutes-long consistency, promptable world events). These are simulators-for-agents; robotics consumes them as priors.
  • Robot instantiations: Cosmos Policy (Cosmos-Predict2-2B, latent-frame injection of action+future+value → 98.5% LIBERO, 93.6% real ALOHA vs π0.5 88.6%), Genie Envisioner (LTX-Video-2B GE-Base + 160M GE-Act, multi-view, ~1M AgiBot episodes, 200 ms @ 4090), Vid2World (DynamiCrafter→causal interactive WM via Diffusion Forcing), VideoVLA (CogVideoX-5B as the VLA itself, 80.4% SIMPLER).
  • Pros: inherits billions of frames of physical-motion prior; interpretable (you can watch the rollout); strong on visual perturbations and deformables; data-efficient at fine-tune (Cosmos-Policy needed 185 task trajectories). Cons: heavy (multi-billion-param), slow (Cosmos-Policy 390 ms = 6.2× π0.5), drift over long rollouts, weak contact physics.

B. Autoregressive token world models

Tokenize frames (VQGAN/VQ-VAE) and predict the next tokens with a causal transformer — cheaper and more controllable than diffusion, weaker fidelity.

  • Exemplars: iVideoGPT (arXiv 2405.15223), VLA-RFT (compact ~138M AR WM as a verified-reward RL env: 86.6→91.1% LIBERO in 400 steps), WorldGym (~609M DiT + Diffusion Forcing; in-WM vs real Pearson r = 0.78), and the 1X World Model (1x.tech, eval-by-bits-not-atoms; uses the Cosmos 8×8×8 tokenizer).
  • Pros: cheap rollouts, LLM-style scaling/streaming, easy action interleaving. Cons: tokenization caps fidelity; AR multi-step generation degrades without attention fixes (see WorldVLA).

C. Latent / JEPA world models (predict representations, not pixels)

LeCun's thesis: generating pixels wastes capacity and destabilizes long rollouts; predict in a learned latent instead.

  • Exemplars: V-JEPA 2 / V-JEPA 2-AC (arXiv 2506.09985; >1M h video pretrain → ~300M action-conditioned head on <62 h Droid → zero-shot Franka pick-and-place by latent goal-matching), DINO-WM (arXiv 2411.04983; predict future DINOv2 patch features → zero-shot planning), Sparse Imagination (DINO-ViT latent rollouts, token-dropout → −52.6% planning time), CoWVLA (Emu3 predicts a motion-latent chain via VidTwin video-VAE).
  • Pros: cheap, planning-aligned, data-efficient, genuine zero-shot transfer. Cons: quality capped by the frozen/learned feature space; demonstrated mostly on shorter-horizon, constrained tasks; not interpretable.

D. 3D / 4D geometry world models

Predict explicit geometry instead of (or alongside) appearance — sidesteps the cross-embodiment action-vocabulary problem by handing a planner a 3D target.

  • Exemplars: Geometry-aware 4D Video (SVD-2.4B jointly predicts RGB + pointmaps → FoundationPose → EE trajectory; 0.64 novel-view vs DP 0.12), FlowDreamer (arXiv 2505.10075; explicit 3D scene flow + diffusion, +7–11%), Re³Sim (arXiv 2502.08645; Gaussian-splat real-to-sim), Avi (point-cloud→IK).
  • Pros: physically grounded, embodiment-agnostic target, strong on novel viewpoints. Cons: open-loop and slow (Geometry-4D ~30 s/chunk), depends on a perception stack (FoundationPose/SAM2/CAD).

E. Structured-cue forecasters

Don't predict full RGB — predict a compact bundle of world-knowledge cues as an auxiliary signal.

  • Exemplar: DreamVLA (forecasts dynamic-region mask + depth + DINOv2/SAM features, not pixels; CALVIN 4.44, real 76.7%).
  • Pros: cheapest WM signal, no pixel-decode at inference, a free representation regularizer. Cons: not a simulator — can't roll out for RL/eval; a regularizer, not a substitute for the action head.

F. World-model + inverse-dynamics action decoding (compositional)

Not a distinct prediction target but a composition that cuts across families A–D: a world model predicts the future (frames / latents / geometry) and a separate inverse-dynamics model (IDM) — or latent-action model — reads consecutive predicted states and outputs the action that bridges them. "Generate the future, then decode the action that gets there." Three sub-variants:

  • Explicit video-WM → IDM (generate-then-decode). A text/goal-conditioned video diffusion model produces an H-frame plan; an inverse-dynamics net maps each consecutive frame pair → action, so the policy is the IDM. UniPi (NeurIPS 2023), HiP (NeurIPS 2023; LLM subgoal → video → VC-1-init IDM), VLP (Video-Language Planning), RoboDreamer (ICML 2024; compositional video → IDM, closed-loop replan). The data-factory use of the same pattern is DreamGen, which runs an IDM / latent-action model post-hoc on generated video to recover pseudo-actions ("neural trajectories") for imitation learning.

  • Disentangled forward + inverse pretraining. Train the world (forward) model and the IDM as explicitly-paired modules. DeFI (ICLR 2026): forward-dynamics GFDM (video pretraining) + inverse-dynamics GIDM (self-supervised latent actions from action-free video); CALVIN ABC-D 4.51, real-world 81.3%.

  • Latent-action IDM (from action-free video). A latent-action model z_t = IDM(o_t, o_{t+K}) paired with a forward model reconstructs the future; the latent action is later grounded to real actuation. villa-X (adds a proprio-FDM), UniVLA, ViPRA, Human-Video Pretraining. External precedents: LAPA, OpenAI VPT.

  • Pros: decouples what-to-do (world model — learnable from cheap action-free / human video) from how-to-act (IDM), which unlocks internet/human-video pretraining and cross-embodiment (the IDM is the only embodiment-specific part); latent-action variants need zero action labels. Cons: inverse-dynamics drift — any pixel/latent error in the predicted future compounds into actuation error; open-loop variants (UniPi) fail on contact-rich tasks; a full video rollout per replan is expensive. This is precisely why the field has shifted toward co-generation (actions + frames denoised jointly — CoT-VLA, dVLA, VideoVLA) and goal-pose decoding (Goal-VLA, collapsing the generated image to an object pose for a planner — decoupling at the pose level rather than the frame-pair level), and why π0.7 deliberately avoids IDM (its flow-matching expert consumes subgoal tokens and outputs actions directly). Full cluster: Goal-Image review §4.B.


4. The usage roles (Axis 2 — cross-ref Category E)

Role What the WM does Canonical robot exemplars
Backbone (VAM) — E4 the video model is the policy; actions co-generated Cosmos Policy, Genie Envisioner, VideoVLA, LingBot-VA, DreamZero
RL environment — E3 imagine rollouts + reward → policy gradient WMPO (GRPO, real 53→70%), VLA-RFT, GigaBrain RAMP
Data factory — E2 offline-generate trajectories → standard IL DreamGen (~10× teleop, log-linear scaling), RIGVid (85% zero-demo), AnchorDream (2512.11797, +36.4% sim)
Evaluator — (E3-adjacent) rank checkpoints by imagined success WorldGym (r=0.78), Ctrl-World (+44.7% via imagined-rollout SFT), 1X World Model
Planner / MPC — E5/F search actions toward an imagined goal Sparse Imagination, TMoW, Goal-VLA, DINO-WM, NWM (2412.03572)
Auxiliary loss — E1 forecast as a representation regularizer DreamVLA, Geometry-4D (E1+E5)
Predictive reference for reflexive control (new) forecast short-horizon contact; servo predicted-vs-observed at sensor rate OmniVTA (visuo-tactile WM + 60 Hz reflexive controller)

5. Latest trends (2024 → 2026)

  1. "Video generation → world model" went mainstream. Open video DiTs (Wan2.x, LTX, OpenSora, Cosmos) are now cheap enough to fine-tune into action-conditioned simulators; this seeded the entire E4 backbone wave and made E2/E3 affordable.
  2. Genie 3's real-time interactivity (24 fps, promptable events, Aug 2025) and GameNGen (DOOM at >20 fps, 2408.14837) proved a neural net can be the engine — the conceptual ceiling robot WMs are climbing toward.
  3. The pixels-vs-latent split hardened. Pixel WFMs (Cosmos, Genie, Cosmos-Policy) win interpretability, sim2real data, and visual robustness; latent/JEPA (V-JEPA 2-AC, DINO-WM, NWM) win cost, data efficiency, and zero-shot planning. Both shipped credible 2025 robot results; neither dominates.
  4. World-model-as-evaluator matured into a research program — WorldGym's r=0.78, Ctrl-World's checkpoint ranking, and the 1X World Model Challenge ("evaluate bits, not atoms").
  5. Data engines got embodiment-aware. DreamGen → AnchorDream (anchor the robot embodiment to avoid hallucinated arms) → Re³Sim/Phys2Real (real-to-sim Gaussian-splat + VLM physics priors) — the synthetic-data line is the most deployment-ready use today.
  6. The first controlled robustness audit landed. WAM vs VLA Robustness is the load-bearing empirical check on the whole family (see §6).

6. Pros & cons — the honest synthesis

The case for WMs: the only known route to (a) RL/eval without hardware, (b) teleop-free data at ~10× scale, (c) cross-embodiment via geometry/latent targets, and (d) interpretable foresight. When the WM is good, it is a force multiplier on every downstream axis.

The six tensions the 2025–26 literature flags:

  1. Hallucination / physical implausibility. Generated rollouts can violate kinematics and contact. Most dangerous as RL reward or evaluator — PPO will happily optimize against a fantasy. WMPO mitigates by fine-tuning the WM on current-policy rollouts so it can imagine failures; WorldGym adds a VLM judge.
  2. Compounding error over long horizons. AR/latent errors amplify → blur, drift, corrupted reward. Mitigations: chunk-wise joint denoising (vs free-running), bounded-horizon planning, Diffusion Forcing.
  3. Latency / cost. High-fidelity action-conditioned prediction is expensive; the robustness study measured WAM inference at 4.8×–83× π0.5 (π0.5 63 ms; Cosmos-Policy 390 ms; LingBot-VA up to 5,230 ms in sim). On-robot real-time multi-view prediction is unsolved.
  4. Evaluation is hard. PSNR/FVD measure pixels, not task-relevant dynamics or controllability — a weak proxy for usefulness.
  5. Pixels vs latent (trend #3) is an unresolved architectural bet, not a settled choice.
  6. Sim-to-real of generated data is unproven at the limit — hallucinated artifacts can inject spurious correlations; Fast-WAM trained clean-only collapses 97.6→51.5% on LIBERO, showing the video prior is necessary but not sufficient — data diversity remains the dominant lever, not the WM alone.

The empirical verdict (from the controlled benchmark): video-generation WAMs win visual-perturbation robustness (Cosmos-Policy 82.2%, LingBot-VA 74.2%) but lose camera-viewpoint and robot-state geometry, and a data-diverse plain VLA (π0.5, 85.7%) still leads overall on LIBERO-Plus. WMs are not yet a free win over a well-trained VLA — they trade latency and geometry for visual robustness and synthetic data.


7. Comparison table

Family (Axis 1) Predicts Typical size Strengths Weaknesses Best role
A. Pixel video diffusion future frames 1.5–14 B visual robustness, interpretable, sim2real data slow, drift, contact physics backbone · data factory · eval
B. AR token frame tokens 0.1–0.6 B cheap, streamable, controllable fidelity cap, AR degradation RL env · eval
C. Latent / JEPA embeddings 0.3–1 B cheap, zero-shot planning, data-efficient fidelity cap, short-horizon, opaque planner
D. 3D / 4D geometry pointmaps · flow · splats 1–2.4 B physically grounded, embodiment-agnostic open-loop, slow, perception-stack dep. planner (geometry→IK)
E. Structured cues mask+depth+features adds to base cheapest, free regularizer not a simulator auxiliary loss
F. WM + inverse-dynamics (compositional) future + decoded action WM + small IDM action-free/human-video pretraining, cross-embodiment inverse-dynamics drift, open-loop on contact, replan cost data factory · planner · pretraining

8. Open questions

  • Q1. Is the right substrate pixels or latents? Latent JEPA is cheaper and zero-shot but unproven long-horizon; pixels are interpretable and data-generating but slow. No head-to-head on the same manipulation suite exists yet.
  • Q2. Can a WM be trusted as an RL reward at scale? WMPO's policy-behavior alignment and WorldGym's VLM judge are patches; the hallucination-as-reward failure mode is unsolved in general.
  • Q3. Does WM-generated data transfer as well as physics-sim or real data? AnchorDream/Re³Sim suggest yes for visual diversity; no controlled study isolates it from data-diversity confounds.
  • Q4. Latency. Can multi-view action-conditioned prediction hit real-time on-robot, or is the WM forever a training-/eval-time tool? Genie 3 shows game-resolution real-time is possible; robot-fidelity is not there.
  • Q5. Evaluation. What replaces FVD/PSNR with a controllability-and-dynamics metric? (1X's "bits not atoms" is a start.)

9. Practical decision guide

  1. No robot data for the target task? → Data factory (A): DreamGen / AnchorDream; or zero-shot geometry/latent planning (Goal-VLA, DINO-WM, V-JEPA 2-AC).
  2. Expensive/unsafe embodiment, want RL without a sim? → RL environment (E3): WMPO (pixel) or VLA-RFT (cheap AR-token).
  3. Need to rank checkpoints before hardware? → Evaluator: WorldGym / Ctrl-World / 1X-WM.
  4. Video-rich domain, want one model for everything? → Backbone VAM (A): Cosmos Policy / Genie Envisioner — but budget for latency and verify geometry robustness (robustness review).
  5. Cross-embodiment is the primary axis? → Geometry (D): predict 4D/pointmaps → IK (Geometry-4D).
  6. Just want a free accuracy bump on an existing VLA? → Auxiliary forecast (E): DreamVLA — no inference cost, no simulator claims.
  7. Latency-critical deployment? → Don't put a billion-param video model in the control loop; use it offline (data/eval/RL) and ship a lean policy.

10. Links


🗓 State of the Field (updated Aug 2026)

Verdict: world models graduated from serving VLAs to challenging them — dynamics-pretrained models now beat π0.5-class VLAs on hardware, though VLAs still outnumber them ~2:1 at the venues.

📈 Trend

2025: supporting roles (data engines, evaluators, aux losses). H1 2026, on hardware: LDA-1B (dynamics+policy+forecasting in DINO latent space, 30k h) beats π0.5 by +21/+48/+23% (contact-rich/dexterous/long-horizon) and gains +10% from low-quality data BC discards; mimic-video (Cosmos-Predict2 + IDM decoder) claims 10× sample efficiency. "VLM backbones are blind to physical causality" is now a testable, partially-supported thesis. Census: RSS 2026 ran ~33 VLA vs ~15 WM papers; ICML 2026 fields a mature 12-paper WM cluster — DreamDojo (44,000 h egocentric human video → continuous-latent-action world model, distilled to real-time 10.9 FPS), LAC-WM (latent-action conditioning beats explicit-action by up to +46.7% on unseen robots), and VLAW (policy↔WM co-improvement loop, +39.2% absolute).

⚖️ Approaches & trade-offs

Role Exemplars Pros Cons
Data engine DreamGen, GigaWorld, Qwen-RobotWorld Mature, safe, scalable Synthesis-fidelity ceiling
Evaluator Ctrl-World, WorldGym, RobotWorld (stated); dWorldEval (ICML 2026) Kills benchmark drift Language-actioned variants can't consume continuous policy actions; dWorldEval partially closes this by making actions first-class tokens in a discrete-diffusion WM
VAM / WAM backbone mimic-video, Cosmos-Policy, DreamZero (14B pixel-video, >2× over VLAs), ω-0 (humanoid, reconstruction-free latent future — video dropped from the action path); DYNA-2 (industry: ~1M h human-video-only; video-prediction is a co-training signal dropped at inference → reactive policy; published 1k→1M h transfer fits R²≈0.88–0.93 — self-reported) Dynamics priors, sample efficiency, real-time Loses VLM instruction depth; failures bottleneck on video quality; data-center inference floor
WAM as critic/evaluator (IROS 2026 🆕) AtomVLA — a predictive latent world model scores VLA action chunks against LLM-derived subtasks → enables offline GRPO (no robot rollouts); LIBERO 97.0%. Also RoboDream (compositional WM as data factory), Scaling Cross-Embodiment WMs for dexterous manip Latent WM as cheap post-training critic; no online rollout cost; long-horizon robustness Critic fidelity is the ceiling; sim-anchored eval
Unified WM+policy LDA-1B Uses all data tiers incl. failures Heaviest training
Aux losses FLARE, DreamVLA Nearly free Weakest signal

⚠️ Limitations & open problems

  • Contact-rich physics fidelity is the universal weak spot (and the blocker for WM-as-evaluator on manipulation).
  • 20B-class video backbones price out most labs; no small-model VAM recipe yet.
  • Zero-shot WAM policies still underperform fine-tuned VLAs on standard suites; the robustness-vs-raw-success trade-off (Review-WAM-vs-VLA-Robustness) is unresolved.

← Back to Home

⚠️ **GitHub.com Fallback** ⚠️