Review World Models - Heungwoo/research GitHub Wiki
Compiled June 2026 · Focus: learned models that predict the future of the world (pixels, latents, geometry, or reward) and how 2024–2026 robotics uses them — as policy backbones, RL environments, data factories, planners, and evaluators.
This review takes a model-centric lens: what the world model is and predicts. It is the sibling of VLA Architectures §5 Category E, which slices the same papers by how the world model is wired into the policy (the E1–E5 usage taxonomy). Read that page for the action-decoder view; read this one for the world-model view. Companion reviews: WAM vs VLA Robustness · Goal-Image Conditioning · RL for VLA · Cross-Embodiment.
A world model (WM) answers: given the current observation and an action, what happens next? For robotics that single capability unlocks four otherwise-expensive things:
- Data without robots — roll the WM forward to synthesize training trajectories (the teleop-replacement bet).
- RL without real rollouts — treat the WM as a simulator; optimize the policy against imagined returns.
- Evaluation without a test rig — rank policy checkpoints by imagined success before touching hardware.
- Planning without a hand-built model — search action sequences toward an imagined goal.
The strong claim several 2026 papers make is convergence: the same action-conditioned video model can serve as RL environment (WMPO), evaluator (WorldGym, Ctrl-World), and even the policy backbone (Cosmos Policy). World-model quality, not policy architecture, becomes the lever. This is why Category E is the largest growth area in the ICLR 2026 VLA cohort.
flowchart TB
subgraph WHAT["AXIS 1 — what the model predicts / represents"]
P1["Pixels<br/>video diffusion"]
P2["Discrete tokens<br/>autoregressive"]
P3["Latent embeddings<br/>JEPA / feature-space"]
P4["3D / 4D geometry<br/>pointmaps · splats · flow"]
P5["Structured cues<br/>mask + depth + features"]
end
subgraph HOW["AXIS 2 — how robotics uses it"]
U1["Policy backbone (VAM)"]
U2["RL environment"]
U3["Data factory"]
U4["Planner / MPC"]
U5["Evaluator"]
U6["Auxiliary loss"]
end
WHAT --> HOW
The two axes are largely independent: a pixel video model can be a backbone (Cosmos-Policy), an RL env (WMPO), a data factory (DreamGen), or an evaluator (WorldGym); a latent model can plan (DINO-WM) or back a policy (CoWVLA). §3 organizes by Axis 1 (the model-centric contribution of this page); §4 maps onto Axis 2 (cross-referencing Category E).
Fine-tune an internet-scale video generator (Cosmos, Wan2.x, LTX-Video, OpenSora, Stable Video Diffusion, DynamiCrafter, CogVideoX) into an action-conditioned future-frame predictor. The defining move of 2024–26 is "video generation → world model": add action/pose conditioning + long-horizon consistency + a downstream control role.
- General WFMs: NVIDIA Cosmos (platform paper arXiv 2501.03575; Predict / Transfer / Reason branches; action-conditioned Predict2 variants) and DeepMind Genie 2 (Dec 2024, playable 3D worlds from one image) → Genie 3 (Aug 2025, real-time 24 fps / 720p, minutes-long consistency, promptable world events). These are simulators-for-agents; robotics consumes them as priors.
- Robot instantiations: Cosmos Policy (Cosmos-Predict2-2B, latent-frame injection of action+future+value → 98.5% LIBERO, 93.6% real ALOHA vs π0.5 88.6%), Genie Envisioner (LTX-Video-2B GE-Base + 160M GE-Act, multi-view, ~1M AgiBot episodes, 200 ms @ 4090), Vid2World (DynamiCrafter→causal interactive WM via Diffusion Forcing), VideoVLA (CogVideoX-5B as the VLA itself, 80.4% SIMPLER).
- Pros: inherits billions of frames of physical-motion prior; interpretable (you can watch the rollout); strong on visual perturbations and deformables; data-efficient at fine-tune (Cosmos-Policy needed 185 task trajectories). Cons: heavy (multi-billion-param), slow (Cosmos-Policy 390 ms = 6.2× π0.5), drift over long rollouts, weak contact physics.
Tokenize frames (VQGAN/VQ-VAE) and predict the next tokens with a causal transformer — cheaper and more controllable than diffusion, weaker fidelity.
-
Exemplars: iVideoGPT (arXiv 2405.15223), VLA-RFT (compact ~138M AR WM as a verified-reward RL env: 86.6→91.1% LIBERO in 400 steps), WorldGym (~609M DiT + Diffusion Forcing; in-WM vs real Pearson r = 0.78), and the 1X World Model (
1x.tech, eval-by-bits-not-atoms; uses the Cosmos 8×8×8 tokenizer). - Pros: cheap rollouts, LLM-style scaling/streaming, easy action interleaving. Cons: tokenization caps fidelity; AR multi-step generation degrades without attention fixes (see WorldVLA).
LeCun's thesis: generating pixels wastes capacity and destabilizes long rollouts; predict in a learned latent instead.
- Exemplars: V-JEPA 2 / V-JEPA 2-AC (arXiv 2506.09985; >1M h video pretrain → ~300M action-conditioned head on <62 h Droid → zero-shot Franka pick-and-place by latent goal-matching), DINO-WM (arXiv 2411.04983; predict future DINOv2 patch features → zero-shot planning), Sparse Imagination (DINO-ViT latent rollouts, token-dropout → −52.6% planning time), CoWVLA (Emu3 predicts a motion-latent chain via VidTwin video-VAE).
- Pros: cheap, planning-aligned, data-efficient, genuine zero-shot transfer. Cons: quality capped by the frozen/learned feature space; demonstrated mostly on shorter-horizon, constrained tasks; not interpretable.
Predict explicit geometry instead of (or alongside) appearance — sidesteps the cross-embodiment action-vocabulary problem by handing a planner a 3D target.
- Exemplars: Geometry-aware 4D Video (SVD-2.4B jointly predicts RGB + pointmaps → FoundationPose → EE trajectory; 0.64 novel-view vs DP 0.12), FlowDreamer (arXiv 2505.10075; explicit 3D scene flow + diffusion, +7–11%), Re³Sim (arXiv 2502.08645; Gaussian-splat real-to-sim), Avi (point-cloud→IK).
- Pros: physically grounded, embodiment-agnostic target, strong on novel viewpoints. Cons: open-loop and slow (Geometry-4D ~30 s/chunk), depends on a perception stack (FoundationPose/SAM2/CAD).
Don't predict full RGB — predict a compact bundle of world-knowledge cues as an auxiliary signal.
- Exemplar: DreamVLA (forecasts dynamic-region mask + depth + DINOv2/SAM features, not pixels; CALVIN 4.44, real 76.7%).
- Pros: cheapest WM signal, no pixel-decode at inference, a free representation regularizer. Cons: not a simulator — can't roll out for RL/eval; a regularizer, not a substitute for the action head.
Not a distinct prediction target but a composition that cuts across families A–D: a world model predicts the future (frames / latents / geometry) and a separate inverse-dynamics model (IDM) — or latent-action model — reads consecutive predicted states and outputs the action that bridges them. "Generate the future, then decode the action that gets there." Three sub-variants:
-
Explicit video-WM → IDM (generate-then-decode). A text/goal-conditioned video diffusion model produces an H-frame plan; an inverse-dynamics net maps each consecutive frame pair → action, so the policy is the IDM. UniPi (NeurIPS 2023), HiP (NeurIPS 2023; LLM subgoal → video → VC-1-init IDM), VLP (Video-Language Planning), RoboDreamer (ICML 2024; compositional video → IDM, closed-loop replan). The data-factory use of the same pattern is DreamGen, which runs an IDM / latent-action model post-hoc on generated video to recover pseudo-actions ("neural trajectories") for imitation learning.
-
Disentangled forward + inverse pretraining. Train the world (forward) model and the IDM as explicitly-paired modules. DeFI (ICLR 2026): forward-dynamics GFDM (video pretraining) + inverse-dynamics GIDM (self-supervised latent actions from action-free video); CALVIN ABC-D 4.51, real-world 81.3%.
-
Latent-action IDM (from action-free video). A latent-action model
z_t = IDM(o_t, o_{t+K})paired with a forward model reconstructs the future; the latent action is later grounded to real actuation. villa-X (adds a proprio-FDM), UniVLA, ViPRA, Human-Video Pretraining. External precedents: LAPA, OpenAI VPT. -
Pros: decouples what-to-do (world model — learnable from cheap action-free / human video) from how-to-act (IDM), which unlocks internet/human-video pretraining and cross-embodiment (the IDM is the only embodiment-specific part); latent-action variants need zero action labels. Cons: inverse-dynamics drift — any pixel/latent error in the predicted future compounds into actuation error; open-loop variants (UniPi) fail on contact-rich tasks; a full video rollout per replan is expensive. This is precisely why the field has shifted toward co-generation (actions + frames denoised jointly — CoT-VLA, dVLA, VideoVLA) and goal-pose decoding (Goal-VLA, collapsing the generated image to an object pose for a planner — decoupling at the pose level rather than the frame-pair level), and why π0.7 deliberately avoids IDM (its flow-matching expert consumes subgoal tokens and outputs actions directly). Full cluster: Goal-Image review §4.B.
| Role | What the WM does | Canonical robot exemplars |
|---|---|---|
| Backbone (VAM) — E4 | the video model is the policy; actions co-generated | Cosmos Policy, Genie Envisioner, VideoVLA, LingBot-VA, DreamZero |
| RL environment — E3 | imagine rollouts + reward → policy gradient | WMPO (GRPO, real 53→70%), VLA-RFT, GigaBrain RAMP |
| Data factory — E2 | offline-generate trajectories → standard IL | DreamGen (~10× teleop, log-linear scaling), RIGVid (85% zero-demo), AnchorDream (2512.11797, +36.4% sim) |
| Evaluator — (E3-adjacent) | rank checkpoints by imagined success | WorldGym (r=0.78), Ctrl-World (+44.7% via imagined-rollout SFT), 1X World Model |
| Planner / MPC — E5/F | search actions toward an imagined goal | Sparse Imagination, TMoW, Goal-VLA, DINO-WM, NWM (2412.03572) |
| Auxiliary loss — E1 | forecast as a representation regularizer | DreamVLA, Geometry-4D (E1+E5) |
| Predictive reference for reflexive control (new) | forecast short-horizon contact; servo predicted-vs-observed at sensor rate | OmniVTA (visuo-tactile WM + 60 Hz reflexive controller) |
- "Video generation → world model" went mainstream. Open video DiTs (Wan2.x, LTX, OpenSora, Cosmos) are now cheap enough to fine-tune into action-conditioned simulators; this seeded the entire E4 backbone wave and made E2/E3 affordable.
- Genie 3's real-time interactivity (24 fps, promptable events, Aug 2025) and GameNGen (DOOM at >20 fps, 2408.14837) proved a neural net can be the engine — the conceptual ceiling robot WMs are climbing toward.
- The pixels-vs-latent split hardened. Pixel WFMs (Cosmos, Genie, Cosmos-Policy) win interpretability, sim2real data, and visual robustness; latent/JEPA (V-JEPA 2-AC, DINO-WM, NWM) win cost, data efficiency, and zero-shot planning. Both shipped credible 2025 robot results; neither dominates.
- World-model-as-evaluator matured into a research program — WorldGym's r=0.78, Ctrl-World's checkpoint ranking, and the 1X World Model Challenge ("evaluate bits, not atoms").
- Data engines got embodiment-aware. DreamGen → AnchorDream (anchor the robot embodiment to avoid hallucinated arms) → Re³Sim/Phys2Real (real-to-sim Gaussian-splat + VLM physics priors) — the synthetic-data line is the most deployment-ready use today.
- The first controlled robustness audit landed. WAM vs VLA Robustness is the load-bearing empirical check on the whole family (see §6).
The case for WMs: the only known route to (a) RL/eval without hardware, (b) teleop-free data at ~10× scale, (c) cross-embodiment via geometry/latent targets, and (d) interpretable foresight. When the WM is good, it is a force multiplier on every downstream axis.
The six tensions the 2025–26 literature flags:
- Hallucination / physical implausibility. Generated rollouts can violate kinematics and contact. Most dangerous as RL reward or evaluator — PPO will happily optimize against a fantasy. WMPO mitigates by fine-tuning the WM on current-policy rollouts so it can imagine failures; WorldGym adds a VLM judge.
- Compounding error over long horizons. AR/latent errors amplify → blur, drift, corrupted reward. Mitigations: chunk-wise joint denoising (vs free-running), bounded-horizon planning, Diffusion Forcing.
- Latency / cost. High-fidelity action-conditioned prediction is expensive; the robustness study measured WAM inference at 4.8×–83× π0.5 (π0.5 63 ms; Cosmos-Policy 390 ms; LingBot-VA up to 5,230 ms in sim). On-robot real-time multi-view prediction is unsolved.
- Evaluation is hard. PSNR/FVD measure pixels, not task-relevant dynamics or controllability — a weak proxy for usefulness.
- Pixels vs latent (trend #3) is an unresolved architectural bet, not a settled choice.
- Sim-to-real of generated data is unproven at the limit — hallucinated artifacts can inject spurious correlations; Fast-WAM trained clean-only collapses 97.6→51.5% on LIBERO, showing the video prior is necessary but not sufficient — data diversity remains the dominant lever, not the WM alone.
The empirical verdict (from the controlled benchmark): video-generation WAMs win visual-perturbation robustness (Cosmos-Policy 82.2%, LingBot-VA 74.2%) but lose camera-viewpoint and robot-state geometry, and a data-diverse plain VLA (π0.5, 85.7%) still leads overall on LIBERO-Plus. WMs are not yet a free win over a well-trained VLA — they trade latency and geometry for visual robustness and synthetic data.
| Family (Axis 1) | Predicts | Typical size | Strengths | Weaknesses | Best role |
|---|---|---|---|---|---|
| A. Pixel video diffusion | future frames | 1.5–14 B | visual robustness, interpretable, sim2real data | slow, drift, contact physics | backbone · data factory · eval |
| B. AR token | frame tokens | 0.1–0.6 B | cheap, streamable, controllable | fidelity cap, AR degradation | RL env · eval |
| C. Latent / JEPA | embeddings | 0.3–1 B | cheap, zero-shot planning, data-efficient | fidelity cap, short-horizon, opaque | planner |
| D. 3D / 4D geometry | pointmaps · flow · splats | 1–2.4 B | physically grounded, embodiment-agnostic | open-loop, slow, perception-stack dep. | planner (geometry→IK) |
| E. Structured cues | mask+depth+features | adds to base | cheapest, free regularizer | not a simulator | auxiliary loss |
| F. WM + inverse-dynamics (compositional) | future + decoded action | WM + small IDM | action-free/human-video pretraining, cross-embodiment | inverse-dynamics drift, open-loop on contact, replan cost | data factory · planner · pretraining |
- Q1. Is the right substrate pixels or latents? Latent JEPA is cheaper and zero-shot but unproven long-horizon; pixels are interpretable and data-generating but slow. No head-to-head on the same manipulation suite exists yet.
- Q2. Can a WM be trusted as an RL reward at scale? WMPO's policy-behavior alignment and WorldGym's VLM judge are patches; the hallucination-as-reward failure mode is unsolved in general.
- Q3. Does WM-generated data transfer as well as physics-sim or real data? AnchorDream/Re³Sim suggest yes for visual diversity; no controlled study isolates it from data-diversity confounds.
- Q4. Latency. Can multi-view action-conditioned prediction hit real-time on-robot, or is the WM forever a training-/eval-time tool? Genie 3 shows game-resolution real-time is possible; robot-fidelity is not there.
- Q5. Evaluation. What replaces FVD/PSNR with a controllability-and-dynamics metric? (1X's "bits not atoms" is a start.)
- No robot data for the target task? → Data factory (A): DreamGen / AnchorDream; or zero-shot geometry/latent planning (Goal-VLA, DINO-WM, V-JEPA 2-AC).
- Expensive/unsafe embodiment, want RL without a sim? → RL environment (E3): WMPO (pixel) or VLA-RFT (cheap AR-token).
- Need to rank checkpoints before hardware? → Evaluator: WorldGym / Ctrl-World / 1X-WM.
- Video-rich domain, want one model for everything? → Backbone VAM (A): Cosmos Policy / Genie Envisioner — but budget for latency and verify geometry robustness (robustness review).
- Cross-embodiment is the primary axis? → Geometry (D): predict 4D/pointmaps → IK (Geometry-4D).
- Just want a free accuracy bump on an existing VLA? → Auxiliary forecast (E): DreamVLA — no inference cost, no simulator claims.
- Latency-critical deployment? → Don't put a billion-param video model in the control loop; use it offline (data/eval/RL) and ship a lean policy.
- Wiki world-model pages (backbone/policy): Cosmos Policy · Genie Envisioner · Vid2World · VideoVLA · CoWVLA
- Wiki world-model pages (RL/eval/data): WMPO · VLA-RFT · Ctrl-World · WorldGym · GigaBrain-0.5M · DreamGen · RIGVid
- Wiki world-model pages (geometry/latent/aux): Geometry-4D Video · DreamVLA · Sparse Imagination · TMoW · Goal-VLA
- Companion reviews: VLA Architectures (Category E) · WAM vs VLA Robustness · Goal-Image Conditioning · RL for VLA
- Key external WMs: Cosmos (2501.03575) · Genie 3 (DeepMind blog) · V-JEPA 2 (2506.09985) · DINO-WM (2411.04983) · NWM (2412.03572)
- Key external WMs (cont.): WorldVLA (2506.21539) · iVideoGPT (2405.15223) · UniSim (2310.06114) · GameNGen (2408.14837)
Verdict: world models graduated from serving VLAs to challenging them — dynamics-pretrained models now beat π0.5-class VLAs on hardware, though VLAs still outnumber them ~2:1 at the venues.
2025: supporting roles (data engines, evaluators, aux losses). H1 2026, on hardware: LDA-1B (dynamics+policy+forecasting in DINO latent space, 30k h) beats π0.5 by +21/+48/+23% (contact-rich/dexterous/long-horizon) and gains +10% from low-quality data BC discards; mimic-video (Cosmos-Predict2 + IDM decoder) claims 10× sample efficiency. "VLM backbones are blind to physical causality" is now a testable, partially-supported thesis. Census: RSS 2026 ran ~33 VLA vs ~15 WM papers; ICML 2026 fields a mature 12-paper WM cluster — DreamDojo (44,000 h egocentric human video → continuous-latent-action world model, distilled to real-time 10.9 FPS), LAC-WM (latent-action conditioning beats explicit-action by up to +46.7% on unseen robots), and VLAW (policy↔WM co-improvement loop, +39.2% absolute).
| Role | Exemplars | Pros | Cons |
|---|---|---|---|
| Data engine | DreamGen, GigaWorld, Qwen-RobotWorld | Mature, safe, scalable | Synthesis-fidelity ceiling |
| Evaluator | Ctrl-World, WorldGym, RobotWorld (stated); dWorldEval (ICML 2026) | Kills benchmark drift | Language-actioned variants can't consume continuous policy actions; dWorldEval partially closes this by making actions first-class tokens in a discrete-diffusion WM |
| VAM / WAM backbone | mimic-video, Cosmos-Policy, DreamZero (14B pixel-video, >2× over VLAs), ω-0 (humanoid, reconstruction-free latent future — video dropped from the action path); DYNA-2 (industry: ~1M h human-video-only; video-prediction is a co-training signal dropped at inference → reactive policy; published 1k→1M h transfer fits R²≈0.88–0.93 — self-reported) | Dynamics priors, sample efficiency, real-time | Loses VLM instruction depth; failures bottleneck on video quality; data-center inference floor |
| WAM as critic/evaluator (IROS 2026 🆕) | AtomVLA — a predictive latent world model scores VLA action chunks against LLM-derived subtasks → enables offline GRPO (no robot rollouts); LIBERO 97.0%. Also RoboDream (compositional WM as data factory), Scaling Cross-Embodiment WMs for dexterous manip | Latent WM as cheap post-training critic; no online rollout cost; long-horizon robustness | Critic fidelity is the ceiling; sim-anchored eval |
| Unified WM+policy | LDA-1B | Uses all data tiers incl. failures | Heaviest training |
| Aux losses | FLARE, DreamVLA | Nearly free | Weakest signal |
- Contact-rich physics fidelity is the universal weak spot (and the blocker for WM-as-evaluator on manipulation).
- 20B-class video backbones price out most labs; no small-model VAM recipe yet.
- Zero-shot WAM policies still underperform fine-tuned VLAs on standard suites; the robustness-vs-raw-success trade-off (Review-WAM-vs-VLA-Robustness) is unresolved.
← Back to Home