ICLR 2026 Geometry 4D Video - Heungwoo/research GitHub Wiki

Geometry-aware 4D Video Generation for Robot Manipulation

Venue: ICLR 2026 (Poster) Authors / affiliations: Zeyi Liu, Shuang Li, Eric Cousineau, Siyuan Feng, Benjamin Burchfiel, Shuran Song — Stanford University + Toyota Research Institute Category: World Model / Video-for-Actions Trend tag: Trend 8 (geometry-aware video world models)

Approach diagram

flowchart LR
  RGBD_n[RGB-D from view v_n] --> EncN[RGB VAE + Pointmap VAE]
  RGBD_m[RGB-D from view v_m] --> EncM[RGB VAE + Pointmap VAE]
  EncN --> Unet[Latent Video Diffusion U-Net<br/>SVD backbone, 2.4B params]
  EncM --> Unet
  Unet --> DecN[Decoder for v_n]
  Unet --> DecM[Decoder for v_m<br/>+12 cross-attn layers from v_n]
  DecN --> RGB_n[Future RGB v_n] & PM_n[Pointmap X^n_t']
  DecM --> RGB_m[Future RGB v_m] & PM_mn[Pointmap X^m→n_t'<br/>projected to v_n frame]
  PM_n --> Loss[L_3D-diff:<br/>cross-view alignment]
  PM_mn --> Loss
  RGB_n --> FP[FoundationPose<br/>6DoF gripper tracker]
  PM_n --> FP
  FP --> Action[End-effector trajectory<br/>open-loop control]
Loading

Problem

Video generators for manipulation rarely respect 3D geometry across viewpoints — frames may be temporally coherent yet inconsistent between cameras, undermining their use as a planning substrate. Existing approaches also typically depend on explicit camera poses. The authors target two failure modes simultaneously: temporal coherence (smooth causal motion) and 3D consistency (the same object should occupy the same world-space point regardless of which camera renders it). Pixel-only video diffusion (Stable Video Diffusion, etc.) achieves the former but not the latter; static 3D-aware methods (NeRF/4D Gaussian) reverse the trade-off.

Detailed Method

The model is a Stable-Video-Diffusion (SVD) backbone fine-tuned to jointly predict RGB videos and pointmap videos for two camera views, with cross-view supervision adapted from DUSt3R.

Inputs and tokens

At inference the model takes a single RGB-D frame from each of two views (v_n: native, v_m: secondary), repeats it h=10 times to form a pseudo-history (matching SVD), and predicts h future RGB frames and h future pointmaps per view.

Each conditioning frame is encoded by a frozen RGB VAE (image features) and a separately fine-tuned Pointmap VAE (initialised from the RGB VAE then re-trained on pointmap data). VAE encodings are h × c × w′ × h′ with c=4, w′=32, h′=40 (h=10), and concatenated channel-wise with noisy future latents to give h × 16 × 32 × 40 inputs to the U-Net.

Diffusion objective (RGB)

Standard SVD loss with predict-clean parameterisation:

L_diff(t′) = E_{ε,z(0),k} [ ‖ z_{t′}(0) − f_θ(z_{t′}(k), k) ‖² ]

where z_{t′}(k) = √α_k z_{t′}(0) + √(1−α_k) ε.

Cross-view pointmap loss

For native view v_n, predict pointmap latent z^n_{t′}; for v_m, predict the pointmap latent expressed in v_n's coordinate frame, z^{m→n}_{t′}. The 3D loss is:

L_3D-diff(t′) = E[ ‖ z^n_{t′}(0) − f_θ(z^n_{t′}(k), k, c^n) ‖² ] + E[ ‖ z^{m→n}{t′}(0) − f_θ(z^{m→n}{t′}(k), k, c^m) ‖² ]

Camera poses are needed at training (to define the projection v_m → v_n), but not at inference — the model has internalised the cross-view geometry.

Joint loss with gripper-region reweighting

Final loss (Eq. 4 in the paper) doubles the weight on gripper pixels via a downsampled binary mask (segmentations from sim labels or SAM2 in real):

L = Σ_{t′=t+1..t+h} (1 + 1{w_g(t′)=1}) · ( L^n_diff(t′) + L^m_diff(t′) ) + λ · L_3D-diff(t′)

with λ = 1.

Multi-view cross-attention

Two separate U-Net decoders (identical architecture, independent weights) are used because the asymmetry of "predict in v_n's frame" vs "project from v_m to v_n's frame" cannot be served by weight sharing. After each decoder block in the v_m branch, a cross-attention layer is inserted whose key/value come from the corresponding v_n decoder block. Total 12 added cross-attention layers.

Robot pose extraction

Generated 4D videos are passed to FoundationPose with a binary gripper mask (from SAM2 in the first frame), camera intrinsics, and a CAD model of the gripper. Pose is estimated in both views; the higher-confidence prediction is taken and transformed to global frame using v_n's extrinsics. Gripper state (open/close) is computed from finger-pixel point-cloud centroid distance with task-specific thresholds: 0.10 m for StoreCerealBoxUnderShelf, 0.06 m for PutSpatulaOnTable, 0.12 m for PlaceAppleFromBowlIntoBin.

Hyperparameters

  • Backbone: SVD U-Net, 2.4 B trainable parameters total.
  • Optimizer: AdamW (β₁=0.95, β₂=0.999, ε=10⁻⁸, weight decay 10⁻⁶).
  • LR: 1 × 10⁻⁶, batch size 4.
  • Training compute: 4× NVIDIA RTX A6000 (48 GB), 60 epochs per task; ~15 k extra steps for real-world fine-tuning.
  • Inference: EulerEDMSampler, 25 denoising steps; ~30 s per 10-frame chunk on a single RTX 4090.

Comprehensive Results

Multi-view 4D generation quality (Table 1)

mIoU = cross-view 3D-mask consistency; FVD-n / FVD-m = Fréchet Video Distance per view; AbsRel / δ₁ = depth error and threshold accuracy.

Task 1: StoreCerealBoxUnderShelf

Method mIoU↑ FVD-n↓ FVD-m↓ AbsRel-n↓ AbsRel-m↓ δ₁-n↑ δ₁-m↑
OURS 0.70 411.20 561.43 0.06 0.11 0.95 0.92
OURS w/o MV attn 0.41 497.43 607.73 0.15 0.31 0.75 0.66
4D Gaussian (SoM) 0.39 1208.00 1094.98 0.20 0.31 0.74 0.63
SVD – 977.06 743.25 – – – –
SVD w/ MV attn – 941.73 653.44 – – – –

Task 2: PutSpatulaOnTable

Method mIoU↑ FVD-n↓ FVD-m↓ AbsRel-n↓ AbsRel-m↓ δ₁-n↑ δ₁-m↑
OURS 0.69 377.68 257.70 0.03 0.07 0.98 0.97
OURS w/o MV attn 0.44 451.54 302.29 0.10 0.33 0.89 0.41
4D Gaussian 0.46 1241.13 815.77 0.33 0.30 0.43 0.37
SVD – 370.92 417.56 – – – –
SVD w/ MV attn – 536.02 445.68 – – – –

Task 3: PlaceAppleFromBowlIntoBin

Method mIoU↑ FVD-n↓ FVD-m↓ AbsRel-n↓ AbsRel-m↓ δ₁-n↑ δ₁-m↑
OURS 0.64 490.88 366.98 0.06 0.07 0.95 0.96
OURS w/o MV attn 0.26 597.05 573.73 0.14 0.49 0.76 0.30
4D Gaussian 0.44 1396.10 1191.40 0.18 0.16 0.80 0.81
SVD – 659.52 628.01 – – – –
SVD w/ MV attn – 812.94 766.52 – – – –

Multi-task real-world dataset

Method mIoU↑ FVD-n↓ FVD-m↓ AbsRel-n↓ AbsRel-m↓ δ₁-n↑ δ₁-m↑
OURS 0.56 384.26 331.00 0.10 0.20 0.93 0.89
OURS w/o MV attn 0.32 694.23 601.93 0.14 0.34 0.90 0.81
4D Gaussian 0.00 2102.52 2608.94 0.32 1.80 0.78 0.22
SVD – 605.10 648.41 – – – –
SVD w/ MV attn – 1151.05 931.33 – – – –

The model wins on every task on cross-view consistency and on most depth metrics, while remaining competitive on RGB FVD.

Manipulation success on novel viewpoints (Table 2)

30 rollouts/task, novel object poses + held-out cameras.

Method Task 1 Task 2 Task 3 Avg
Dreamitate 0.10 0.17 0.10 0.12
Diffusion Policy (DP) 0.10 0.27 0.00 0.12
DP3 (3D point cloud) 0.23 0.27 0.00 0.25
OURS 0.73 0.67 0.53 0.64

The paper's headline claim is >50 % absolute average gain over the video-generation baseline Dreamitate (0.64 vs 0.12). Over the strongest behaviour-cloning baseline DP3 (0.25) the margin is +0.39 absolute.

Compute / parameter cost (Table 3)

Method Inference Train memory Trainable params
OURS 30.0 s 47 GB 2.4 B
OURS w/o MV attn 29.3 s 46.5 GB 2.38 B
4D Gaussian 2 s 2.8 GB 856 k
SVD (Dreamitate) 13.4 s 45.8 GB 1.54 B
SVD w/ MV attn 15.1 s 46.3 GB 2.4 B

Scaling to more views (Table 4)

Without retraining, the two-view model extends to a 3rd+ view at inference by running an extra forward pass between the reference view and each new view (cost grows ~linearly in #views). On StoreCerealBoxUnderShelf, generating three views jointly degrades only gracefully: View 2 (w/ View 1) mIoU 0.62 / FVD 566.86 / AbsRel 0.09 / δ₁ 0.95; View 3 (w/ View 1) mIoU 0.54 / FVD 662.83 / AbsRel 0.12 / δ₁ 0.89.

Ablation Studies

  • Multi-view cross-attention (OURS w/o MV attn). Removing cross-attention drops mIoU from 0.70 → 0.41 on Task 1 and 0.64 → 0.26 on Task 3, and inflates AbsRel-m from 0.11 → 0.31. The model fails to learn the v_m → v_n projection without explicit cross-decoder feature flow.
  • Adding cross-attention to SVD only on RGB (SVD w/ MV attn). RGB-side cross-attention without pointmap supervision still yields 3D-inconsistent gripper pose, confirming cross-attention is necessary but insufficient — pointmap supervision is the load-bearing piece.
  • 4D Gaussian baseline. Optimising a Shape-of-Motion 4D Gaussian on a single SVD-generated video underperforms drastically on novel views (e.g. mIoU = 0.00 on real-world data) because the underlying generated video is fixed-view.
  • DP / DP3 baselines. Diffusion Policy fails on novel viewpoints despite multi-view training inputs because it learns no explicit cross-view geometry. DP3's global point-cloud helps slightly on cereal-box but not on smaller objects (spatula, apple).
  • Gripper-region reweighting (Eq. 4). Doubles loss on gripper pixels for higher pose-tracking accuracy; presented as an architecture choice rather than separately ablated.

Limitations

The authors explicitly call out:

  • Multi-view RGB-D training data with varied camera poses is easy to obtain in simulation but hard in the real world due to hardware and calibration constraints. Real-world depth quality is also a problem; they suggest leveraging RGB-only depth estimators (FoundationStereo, VGGT) in future work.
  • Inference is slow (~30 s per 10-frame chunk) compared to end-to-end behaviour cloning, making closed-loop control awkward. Faster substrates suggested: pyramidal flow matching, autoregressive video transformers, one-step diffusion (Diffusion Adversarial Post-Training).
  • Open-loop deployment: the robot executes generated trajectories without re-planning until a chunk completes — drift and contact errors are not corrected mid-chunk.
  • Action extraction depends on FoundationPose + SAM2 with a CAD model of the gripper; failures in any of these propagate.

Significance & Positioning

This work pushes "video-as-world-model" from 2D-temporal toward genuinely 4D / geometry-aware. Three things distinguish it from concurrent video-policy work:

  1. No camera poses at inference. A practical point: deploying a video world model in a new lab requires only an RGB-D image per camera, not extrinsics.
  2. Pointmap-domain supervision rather than RGB-only multi-view supervision — the authors show in ablation that RGB cross-attention alone (SVD w/ MV attn) is not enough for 3D consistency.
  3. Direct gripper-pose recovery via FoundationPose sidesteps the noisy "video → inverse-dynamics" step that has plagued video-conditioned policies (Dreamitate, Du et al. UniPi style).

Within the 2026 video-world-model cluster:

  • ViPRA and Vid2World focus on action prediction from 2D video; this work supplies the 3D substrate they were missing.
  • Genie Envisioner is a generalist video-as-world agent — adding pointmap heads in this style would plausibly improve its viewpoint generalisation.
  • Ctrl-World is closed-loop and faster; this work is open-loop and 3D-aware. The two are complementary.

The 64 % average success rate on novel viewpoints (vs 12–25 % for diffusion-policy / DP3 baselines) is the headline empirical result and the strongest evidence yet that geometric video supervision pays off for downstream control rather than just generation quality.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️