ICLR 2026 Vid2World - Heungwoo/research GitHub Wiki

Vid2World — Crafting Video Diffusion Models to Interactive World Models

Venue: ICLR 2026 Category: World Model / Video Diffusion Trend tag: Trend 8 (world models from internet-scale video) Affiliation: Tsinghua University · Chongqing University

Approach diagram

flowchart LR
  VD[Pretrained DynamiCrafter<br/>1.1B U-Net video diffusion] --> Caus["Causalization<br/>1) Causal mask on temporal attn<br/>2) Extrapolative weight transfer<br/>   for temporal conv"]
  Caus --> DF[Diffusion Forcing<br/>k_t ~ U 0,K independently per frame]
  DF --> CAI[Causal Action Injection<br/>per-frame action embed via MLP]
  CAI --> CAG[Causal Action Guidance<br/>action dropout p +<br/>classifier-free ε_guided = 1+λ ε_cond − λ ε_uncond]
  CAG --> WM[Interactive autoregressive<br/>world model]
Loading

Problem

Two camps with complementary weaknesses:

  • Domain-specific world models (DreamerV3, DIAMOND, NWM) predict accurately but require costly per-domain action-labeled data and produce low-fidelity rollouts.
  • Pretrained video diffusion models (DynamiCrafter, Sora, Veo) generate high-fidelity video at internet scale but are non-interactive — they denoise full sequences with bidirectional temporal context, have no frame-level action conditioning, and cannot do autoregressive rollouts.

The paper argues the right move is "model-level transfer" of internet-scale video priors into a world model, rather than pretraining yet another world model on cross-domain action-labeled data (which is still data-hungry and yields low-fidelity output). Two technical barriers must be cleared: causal generation and frame-level action conditioning.

Detailed Method

Base model

DynamiCrafter — a 1.1B U-Net-based video diffusion model pretrained on internet-scale videos (Xing et al., 2024). Vid2World post-trains it under a causal training objective and adds frame-level action conditioning.

4.1 Video Diffusion Causalization

Temporal attention layers. Causalize via causal masking (no parameter changes — attention is content-based, so restricting receptive field to past tokens is "free").

Temporal convolution layers. Symmetric kernels {w_t}_{t=-m}^{m} aggregate from past and future frames; the original kernels are non-causal. Three weight-transfer strategies are studied:

  1. Shift Weight Transfer — shift the entire kernel m steps into the past, getting {w't}{t=-2m}^{0}. Preserves all weights but introduces temporal misalignment.

  2. Masked Weight Transfer — keep only the {w_t}_{t≤0} weights and zero the rest (hard causal mask at init). Causal but throws away future-facing weights.

  3. Extrapolative Weight Transfer (proposed). Posits a linear feature relationship z_{t+k} ≈ Σ γ_{k,j} z_{t-j} + β_k, then redistributes the future-side weights {w_i}_{i>0} onto the past side to preserve the original convolution output:

    w'j = 1[j≥-m] · w_j + 1[-p+1≤j≤0] · Σ γ{i,-j} w_i, b' = b + Σ w_i β_i.

    Detailed derivation in Appendix A.2; error bound (Proposition 1) in Appendix A.3.

Training objective for causal generation: Diffusion Forcing. Standard video diffusion uses a homogeneous noise level across frames (all frames share k). For causal autoregressive sampling, history frames must be clean (k=0) while the current frame is being denoised — a noise-level distribution the original training never sees. Vid2World adopts Diffusion Forcing (Chen et al., 2024): sample noise level independently per frame k_t ~ U(0,K). This exposes the model to all noise-level combinations and enables flexible causal rollouts.

4.2 Causal Action Guidance

Causal Action Injection. Frame-level action a_{t-1} is encoded via a lightweight MLP and added to the model's latent representation at temporal position t. This binds each predicted frame to its preceding action — the basis for fine-grained interactive control.

Action Dropout for Classifier-Free Guidance. Each timestep's action is independently dropped with probability p:

L(θ) = E [ Σ_t ‖ ε_t − ε_θ([x^{k_τ}τ]{≤t}, [ã_τ]{<t}, [k_τ]{≤t}) ‖² ], ã_t = ∅ w.p. p, else a_t.

This gives both ε_cond = ε_θ(…, [a_τ]{τ<t}, …) (full action context) and ε_uncond = ε_θ(…, [a{τ<t-1}, ∅], …) (most recent action masked). Inference uses CFG-style:

ε_guided = (1+λ) · ε_cond − λ · ε_uncond.

Theorem 4.1 (Causal Action Guidance as Probability Steering; proven in Appendix A.4): with H_t := ([x_τ]{τ<t}, [a_τ]{τ<t-1}) the history excluding the current action, this score composition is equivalent to sampling from a steered posterior p̃(x_t | a_{t-1}, H_t) ∝ p(x_t | H_t) · p(a_{t-1} | x_t, H_t)^ω with ω ∝ (1+λ) — a history-consistent prior times an action-alignment classifier term. The guidance scale λ ∈ ℝ⁺ is thus a knob trading off action responsiveness vs generation fidelity. (The appendix also contains Proposition 1 in A.3, the Extrapolative Weight Transfer error bound.)

Training resources

On RT-1 robot manipulation: post-trained for 100k gradient steps, ~7 days on 4× A100 GPUs (extrapolative-weight-transfer variant). Two inference modes:

  • Vid2World-NAR — denoise all frames simultaneously (matches baselines).
  • Vid2World — autoregressive denoising with causal action guidance.

Ablation models in Table 2 are trained for only 30k gradient steps due to compute budget.

Comprehensive Results (Table 1 of paper)

Robot manipulation — RT-1 (Brohan et al. 2023)

Model FVD ↓ FID ↓ SSIM ↑ LPIPS ↓ PSNR ↑ DreamSim ↓
Pre-trained Base Model† 237.6 5.432 0.712 0.228 20.6 —
Classifier Guidance† 213.1 6.005 0.683 0.250 19.8 0.054
ControlNet† 27.1 3.248 0.836 0.148 24.5 —
Action-Conditioned† 24.2 2.965 0.852 0.134 25.6 —
Language-Conditioned† 33.7 3.511 0.812 0.177 22.1 —
AVID† 39.3 3.436 0.842 0.142 25.3 —
Vid2World-NAR† 18.7 5.871 0.856 0.140 25.8 0.048
Vid2World* 18.5 5.806 0.842 0.152 24.6 0.054

†non-autoregressive prediction; *autoregressive prediction. Vid2World wins FVD in both modes, often by large margins; SSIM/LPIPS/PSNR competitive with or matching the best baselines.

3D Game Simulation — CS:GO (4 conditioning frames → autoregress to length 16; Pearce-Zhu 5.5M frames / 95h)

Model FVD ↓ FID ↓ SSIM ↑ LPIPS ↓
DIAMOND-Fast* 577.1 115.6 0.449 0.547
DIAMOND-HQ* 368.5 87.2 0.447 0.510
Vid2World* 106.6 17.5 0.481 0.404

Relative gains over best DIAMOND configuration: −71.1% FVD, −79.9% FID.

Open-World Navigation — RECON (1B-param NWM baseline; 3D x,y,yaw action)

Model FVD ↓ FID ↓ SSIM ↑ LPIPS ↓ PSNR ↑ DreamSim ↓
NWM (1B)‡ single-step 31.2 34.1 0.389 0.295 ± 0.002 15.343 ± 0.060 0.091 ± 0.001
NWM + Ego4D (1B)‡ single-step 41.0 34.9 0.361 0.368 ± 0.003 14.072 ± 0.075 0.138 ± 0.002
Vid2World* autoregressive 59.4 42.9 0.481 0.3236 16.10 0.108

‡single-step prediction (NWM is conditioned on the prediction timestep t directly, so it avoids autoregressive error accumulation). Vid2World is autoregressive over 16 frames + 4 history (context length 20 > training horizon of 16, demonstrating temporal generalization). On par with NWM and surpasses NWM+Ego4D on 4 of the 6 metrics (the paper's own claim), even under autoregressive error accumulation — it wins SSIM/LPIPS/PSNR/DreamSim (losing only FVD/FID) and posts the best SSIM in the block.

Real2Sim policy evaluation

Vid2World rolls out three RT-1 checkpoints (Begin / 15% / Converged) inside the world model on the close-drawer task. Human evaluators annotate trajectory success. The success-rate ranking inside Vid2World matches the real-world ranking — a direct demonstration that the world model is good enough to use as a SIMPLER-style evaluation environment (Algorithm 3 in paper).

Ablation Studies

Weight transfer × action guidance (Table 2; 30k-step training budget)

WT variant AG FVD ↓ FID ↓ SSIM ↑ LPIPS ↓ PSNR ↑
Shift — 29.9 7.85 0.799 0.185 21.5
Masked — 29.4 7.07 0.824 0.169 22.9
Extrapolative — 28.6 7.52 0.832 0.162 23.4
Masked ✓ 25.8 6.84 0.840 0.159 23.9
Extrapolative ✓ 22.4 6.16 0.839 0.159 23.9

Both causalization (Masked > Shift, Extrapolative ≈ Masked but slightly better) and action guidance (✓ rows beat — rows by 4–6 FVD) contribute independently.

Guidance scale λ sweep on CS:GO (Figure 8)

PSNR / SSIM / LPIPS / DreamSim plotted as a function of λ ∈ [1.0, 4.0]. Performance improves with λ initially (alignment to action helps) but degrades at high λ due to over-sharpening. Sweet spot is around λ ≈ 2.

Limitations (as stated by authors, Sec. 6)

  1. Backbone scale. DynamiCrafter at 1.1B is "relatively lightweight." Authors expect larger video-diffusion backbones (NVIDIA Cosmos, Wan, etc.) to deliver better world-model fidelity but did not test due to compute.
  2. Training time. Causalization + action conditioning take 100k gradient steps × ~7 days × 4 A100. Authors hope future methods can do this in fewer steps.
  3. No explicit policy learning. Vid2World demonstrates Real2Sim evaluation, but the paper does not yet show that policies learned inside Vid2World transfer back to the real world (RL or model-based planning is left as future work).

Significance & Positioning

  • vs DIAMOND (Alonso et al. 2024): DIAMOND trains a domain-specific autoregressive diffusion world model for CS:GO. Vid2World, starting from a generic video diffusion checkpoint, beats DIAMOND-HQ by −71.1% FVD / −79.9% FID — a large gap that argues for model-level transfer of video priors over from-scratch domain training.
  • vs NWM (Bar et al. 2025): NWM is a 1B-parameter dedicated navigation world model, conditioned directly on prediction timestep (so it skips error accumulation). Vid2World matches NWM's single-step performance from autoregressive rollouts, despite being a generic transfer.
  • vs AVID (Rigter et al. 2024): AVID adapts a frozen video model with an action adapter; Vid2World's full causalization + action injection beats it across robot manipulation metrics.
  • vs Genie 2, Cosmos, Wan as world models: these are larger generic video models that could potentially be Vid2World-ized; the paper explicitly flags this as future work.
  • vs Genie Envisioner: GE trains a unified video+action stack from scratch on manipulation data. Vid2World takes the opposite path — preserve a pretrained video diffusion checkpoint, bolt on causality and action guidance. Different bets on where the training cost should be paid.
  • vs Ctrl-World / WMPO: these use world models for RL / policy improvement; Vid2World provides the substrate their successors might use.
  • First systematic study of video-diffusion-to-world-model transfer — the three weight-transfer schemes (with the Extrapolative variant's error bound, Proposition 1 / Appendix A.3) and the causal action guidance mechanism (formalized as probability steering in Theorem 4.1 / Appendix A.4) are the lasting methodological contributions.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️