ICLR 2026 Vid2World - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Category: World Model / Video Diffusion Trend tag: Trend 8 (world models from internet-scale video) Affiliation: Tsinghua University · Chongqing University
flowchart LR
VD[Pretrained DynamiCrafter<br/>1.1B U-Net video diffusion] --> Caus["Causalization<br/>1) Causal mask on temporal attn<br/>2) Extrapolative weight transfer<br/> for temporal conv"]
Caus --> DF[Diffusion Forcing<br/>k_t ~ U 0,K independently per frame]
DF --> CAI[Causal Action Injection<br/>per-frame action embed via MLP]
CAI --> CAG[Causal Action Guidance<br/>action dropout p +<br/>classifier-free ε_guided = 1+λ ε_cond − λ ε_uncond]
CAG --> WM[Interactive autoregressive<br/>world model]
Two camps with complementary weaknesses:
- Domain-specific world models (DreamerV3, DIAMOND, NWM) predict accurately but require costly per-domain action-labeled data and produce low-fidelity rollouts.
- Pretrained video diffusion models (DynamiCrafter, Sora, Veo) generate high-fidelity video at internet scale but are non-interactive — they denoise full sequences with bidirectional temporal context, have no frame-level action conditioning, and cannot do autoregressive rollouts.
The paper argues the right move is "model-level transfer" of internet-scale video priors into a world model, rather than pretraining yet another world model on cross-domain action-labeled data (which is still data-hungry and yields low-fidelity output). Two technical barriers must be cleared: causal generation and frame-level action conditioning.
DynamiCrafter — a 1.1B U-Net-based video diffusion model pretrained on internet-scale videos (Xing et al., 2024). Vid2World post-trains it under a causal training objective and adds frame-level action conditioning.
Temporal attention layers. Causalize via causal masking (no parameter changes — attention is content-based, so restricting receptive field to past tokens is "free").
Temporal convolution layers. Symmetric kernels {w_t}_{t=-m}^{m} aggregate from past and future frames; the original kernels are non-causal. Three weight-transfer strategies are studied:
-
Shift Weight Transfer — shift the entire kernel m steps into the past, getting {w't}{t=-2m}^{0}. Preserves all weights but introduces temporal misalignment.
-
Masked Weight Transfer — keep only the {w_t}_{t≤0} weights and zero the rest (hard causal mask at init). Causal but throws away future-facing weights.
-
Extrapolative Weight Transfer (proposed). Posits a linear feature relationship z_{t+k} ≈ Σ γ_{k,j} z_{t-j} + β_k, then redistributes the future-side weights {w_i}_{i>0} onto the past side to preserve the original convolution output:
w'j = 1[j≥-m] · w_j + 1[-p+1≤j≤0] · Σ γ{i,-j} w_i, b' = b + Σ w_i β_i.
Detailed derivation in Appendix A.2; error bound (Proposition 1) in Appendix A.3.
Training objective for causal generation: Diffusion Forcing. Standard video diffusion uses a homogeneous noise level across frames (all frames share k). For causal autoregressive sampling, history frames must be clean (k=0) while the current frame is being denoised — a noise-level distribution the original training never sees. Vid2World adopts Diffusion Forcing (Chen et al., 2024): sample noise level independently per frame k_t ~ U(0,K). This exposes the model to all noise-level combinations and enables flexible causal rollouts.
Causal Action Injection. Frame-level action a_{t-1} is encoded via a lightweight MLP and added to the model's latent representation at temporal position t. This binds each predicted frame to its preceding action — the basis for fine-grained interactive control.
Action Dropout for Classifier-Free Guidance. Each timestep's action is independently dropped with probability p:
L(θ) = E [ Σ_t ‖ ε_t − ε_θ([x^{k_τ}τ]{≤t}, [ã_τ]{<t}, [k_τ]{≤t}) ‖² ], ã_t = ∅ w.p. p, else a_t.
This gives both ε_cond = ε_θ(…, [a_τ]{τ<t}, …) (full action context) and ε_uncond = ε_θ(…, [a{τ<t-1}, ∅], …) (most recent action masked). Inference uses CFG-style:
ε_guided = (1+λ) · ε_cond − λ · ε_uncond.
Theorem 4.1 (Causal Action Guidance as Probability Steering; proven in Appendix A.4): with H_t := ([x_τ]{τ<t}, [a_τ]{τ<t-1}) the history excluding the current action, this score composition is equivalent to sampling from a steered posterior p̃(x_t | a_{t-1}, H_t) ∝ p(x_t | H_t) · p(a_{t-1} | x_t, H_t)^ω with ω ∝ (1+λ) — a history-consistent prior times an action-alignment classifier term. The guidance scale λ ∈ ℝ⁺ is thus a knob trading off action responsiveness vs generation fidelity. (The appendix also contains Proposition 1 in A.3, the Extrapolative Weight Transfer error bound.)
On RT-1 robot manipulation: post-trained for 100k gradient steps, ~7 days on 4× A100 GPUs (extrapolative-weight-transfer variant). Two inference modes:
- Vid2World-NAR — denoise all frames simultaneously (matches baselines).
- Vid2World — autoregressive denoising with causal action guidance.
Ablation models in Table 2 are trained for only 30k gradient steps due to compute budget.
| Model | FVD ↓ | FID ↓ | SSIM ↑ | LPIPS ↓ | PSNR ↑ | DreamSim ↓ |
|---|---|---|---|---|---|---|
| Pre-trained Base Model† | 237.6 | 5.432 | 0.712 | 0.228 | 20.6 | — |
| Classifier Guidance† | 213.1 | 6.005 | 0.683 | 0.250 | 19.8 | 0.054 |
| ControlNet† | 27.1 | 3.248 | 0.836 | 0.148 | 24.5 | — |
| Action-Conditioned† | 24.2 | 2.965 | 0.852 | 0.134 | 25.6 | — |
| Language-Conditioned† | 33.7 | 3.511 | 0.812 | 0.177 | 22.1 | — |
| AVID† | 39.3 | 3.436 | 0.842 | 0.142 | 25.3 | — |
| Vid2World-NAR† | 18.7 | 5.871 | 0.856 | 0.140 | 25.8 | 0.048 |
| Vid2World* | 18.5 | 5.806 | 0.842 | 0.152 | 24.6 | 0.054 |
†non-autoregressive prediction; *autoregressive prediction. Vid2World wins FVD in both modes, often by large margins; SSIM/LPIPS/PSNR competitive with or matching the best baselines.
3D Game Simulation — CS:GO (4 conditioning frames → autoregress to length 16; Pearce-Zhu 5.5M frames / 95h)
| Model | FVD ↓ | FID ↓ | SSIM ↑ | LPIPS ↓ |
|---|---|---|---|---|
| DIAMOND-Fast* | 577.1 | 115.6 | 0.449 | 0.547 |
| DIAMOND-HQ* | 368.5 | 87.2 | 0.447 | 0.510 |
| Vid2World* | 106.6 | 17.5 | 0.481 | 0.404 |
Relative gains over best DIAMOND configuration: −71.1% FVD, −79.9% FID.
| Model | FVD ↓ | FID ↓ | SSIM ↑ | LPIPS ↓ | PSNR ↑ | DreamSim ↓ |
|---|---|---|---|---|---|---|
| NWM (1B)‡ single-step | 31.2 | 34.1 | 0.389 | 0.295 ± 0.002 | 15.343 ± 0.060 | 0.091 ± 0.001 |
| NWM + Ego4D (1B)‡ single-step | 41.0 | 34.9 | 0.361 | 0.368 ± 0.003 | 14.072 ± 0.075 | 0.138 ± 0.002 |
| Vid2World* autoregressive | 59.4 | 42.9 | 0.481 | 0.3236 | 16.10 | 0.108 |
‡single-step prediction (NWM is conditioned on the prediction timestep t directly, so it avoids autoregressive error accumulation). Vid2World is autoregressive over 16 frames + 4 history (context length 20 > training horizon of 16, demonstrating temporal generalization). On par with NWM and surpasses NWM+Ego4D on 4 of the 6 metrics (the paper's own claim), even under autoregressive error accumulation — it wins SSIM/LPIPS/PSNR/DreamSim (losing only FVD/FID) and posts the best SSIM in the block.
Vid2World rolls out three RT-1 checkpoints (Begin / 15% / Converged) inside the world model on the close-drawer task. Human evaluators annotate trajectory success. The success-rate ranking inside Vid2World matches the real-world ranking — a direct demonstration that the world model is good enough to use as a SIMPLER-style evaluation environment (Algorithm 3 in paper).
| WT variant | AG | FVD ↓ | FID ↓ | SSIM ↑ | LPIPS ↓ | PSNR ↑ |
|---|---|---|---|---|---|---|
| Shift | — | 29.9 | 7.85 | 0.799 | 0.185 | 21.5 |
| Masked | — | 29.4 | 7.07 | 0.824 | 0.169 | 22.9 |
| Extrapolative | — | 28.6 | 7.52 | 0.832 | 0.162 | 23.4 |
| Masked | ✓ | 25.8 | 6.84 | 0.840 | 0.159 | 23.9 |
| Extrapolative | ✓ | 22.4 | 6.16 | 0.839 | 0.159 | 23.9 |
Both causalization (Masked > Shift, Extrapolative ≈ Masked but slightly better) and action guidance (✓ rows beat — rows by 4–6 FVD) contribute independently.
PSNR / SSIM / LPIPS / DreamSim plotted as a function of λ ∈ [1.0, 4.0]. Performance improves with λ initially (alignment to action helps) but degrades at high λ due to over-sharpening. Sweet spot is around λ ≈ 2.
- Backbone scale. DynamiCrafter at 1.1B is "relatively lightweight." Authors expect larger video-diffusion backbones (NVIDIA Cosmos, Wan, etc.) to deliver better world-model fidelity but did not test due to compute.
- Training time. Causalization + action conditioning take 100k gradient steps × ~7 days × 4 A100. Authors hope future methods can do this in fewer steps.
- No explicit policy learning. Vid2World demonstrates Real2Sim evaluation, but the paper does not yet show that policies learned inside Vid2World transfer back to the real world (RL or model-based planning is left as future work).
- vs DIAMOND (Alonso et al. 2024): DIAMOND trains a domain-specific autoregressive diffusion world model for CS:GO. Vid2World, starting from a generic video diffusion checkpoint, beats DIAMOND-HQ by −71.1% FVD / −79.9% FID — a large gap that argues for model-level transfer of video priors over from-scratch domain training.
- vs NWM (Bar et al. 2025): NWM is a 1B-parameter dedicated navigation world model, conditioned directly on prediction timestep (so it skips error accumulation). Vid2World matches NWM's single-step performance from autoregressive rollouts, despite being a generic transfer.
- vs AVID (Rigter et al. 2024): AVID adapts a frozen video model with an action adapter; Vid2World's full causalization + action injection beats it across robot manipulation metrics.
- vs Genie 2, Cosmos, Wan as world models: these are larger generic video models that could potentially be Vid2World-ized; the paper explicitly flags this as future work.
- vs Genie Envisioner: GE trains a unified video+action stack from scratch on manipulation data. Vid2World takes the opposite path — preserve a pretrained video diffusion checkpoint, bolt on causality and action guidance. Different bets on where the training cost should be paid.
- vs Ctrl-World / WMPO: these use world models for RL / policy improvement; Vid2World provides the substrate their successors might use.
- First systematic study of video-diffusion-to-world-model transfer — the three weight-transfer schemes (with the Extrapolative variant's error bound, Proposition 1 / Appendix A.3) and the causal action guidance mechanism (formalized as probability steering in Theorem 4.1 / Appendix A.4) are the lasting methodological contributions.
- OpenReview: https://openreview.net/forum?id=pFyzqbUiF9
- Project page: https://knightnemo.github.io/vid2world/
← Back to ICLR-2026