ICLR 2026 Geometry 4D Video - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 (Poster) Authors / affiliations: Zeyi Liu, Shuang Li, Eric Cousineau, Siyuan Feng, Benjamin Burchfiel, Shuran Song — Stanford University + Toyota Research Institute Category: World Model / Video-for-Actions Trend tag: Trend 8 (geometry-aware video world models)
flowchart LR
RGBD_n[RGB-D from view v_n] --> EncN[RGB VAE + Pointmap VAE]
RGBD_m[RGB-D from view v_m] --> EncM[RGB VAE + Pointmap VAE]
EncN --> Unet[Latent Video Diffusion U-Net<br/>SVD backbone, 2.4B params]
EncM --> Unet
Unet --> DecN[Decoder for v_n]
Unet --> DecM[Decoder for v_m<br/>+12 cross-attn layers from v_n]
DecN --> RGB_n[Future RGB v_n] & PM_n[Pointmap X^n_t']
DecM --> RGB_m[Future RGB v_m] & PM_mn[Pointmap X^m→n_t'<br/>projected to v_n frame]
PM_n --> Loss[L_3D-diff:<br/>cross-view alignment]
PM_mn --> Loss
RGB_n --> FP[FoundationPose<br/>6DoF gripper tracker]
PM_n --> FP
FP --> Action[End-effector trajectory<br/>open-loop control]
Video generators for manipulation rarely respect 3D geometry across viewpoints — frames may be temporally coherent yet inconsistent between cameras, undermining their use as a planning substrate. Existing approaches also typically depend on explicit camera poses. The authors target two failure modes simultaneously: temporal coherence (smooth causal motion) and 3D consistency (the same object should occupy the same world-space point regardless of which camera renders it). Pixel-only video diffusion (Stable Video Diffusion, etc.) achieves the former but not the latter; static 3D-aware methods (NeRF/4D Gaussian) reverse the trade-off.
The model is a Stable-Video-Diffusion (SVD) backbone fine-tuned to jointly predict RGB videos and pointmap videos for two camera views, with cross-view supervision adapted from DUSt3R.
At inference the model takes a single RGB-D frame from each of two views (v_n: native, v_m: secondary), repeats it h=10 times to form a pseudo-history (matching SVD), and predicts h future RGB frames and h future pointmaps per view.
Each conditioning frame is encoded by a frozen RGB VAE (image features) and a separately fine-tuned Pointmap VAE (initialised from the RGB VAE then re-trained on pointmap data). VAE encodings are h × c × w′ × h′ with c=4, w′=32, h′=40 (h=10), and concatenated channel-wise with noisy future latents to give h × 16 × 32 × 40 inputs to the U-Net.
Standard SVD loss with predict-clean parameterisation:
L_diff(t′) = E_{ε,z(0),k} [ ‖ z_{t′}(0) − f_θ(z_{t′}(k), k) ‖² ]
where z_{t′}(k) = √α_k z_{t′}(0) + √(1−α_k) ε.
For native view v_n, predict pointmap latent z^n_{t′}; for v_m, predict the pointmap latent expressed in v_n's coordinate frame, z^{m→n}_{t′}. The 3D loss is:
L_3D-diff(t′) = E[ ‖ z^n_{t′}(0) − f_θ(z^n_{t′}(k), k, c^n) ‖² ] + E[ ‖ z^{m→n}{t′}(0) − f_θ(z^{m→n}{t′}(k), k, c^m) ‖² ]
Camera poses are needed at training (to define the projection v_m → v_n), but not at inference — the model has internalised the cross-view geometry.
Final loss (Eq. 4 in the paper) doubles the weight on gripper pixels via a downsampled binary mask (segmentations from sim labels or SAM2 in real):
L = Σ_{t′=t+1..t+h} (1 + 1{w_g(t′)=1}) · ( L^n_diff(t′) + L^m_diff(t′) ) + λ · L_3D-diff(t′)
with λ = 1.
Two separate U-Net decoders (identical architecture, independent weights) are used because the asymmetry of "predict in v_n's frame" vs "project from v_m to v_n's frame" cannot be served by weight sharing. After each decoder block in the v_m branch, a cross-attention layer is inserted whose key/value come from the corresponding v_n decoder block. Total 12 added cross-attention layers.
Generated 4D videos are passed to FoundationPose with a binary gripper mask (from SAM2 in the first frame), camera intrinsics, and a CAD model of the gripper. Pose is estimated in both views; the higher-confidence prediction is taken and transformed to global frame using v_n's extrinsics. Gripper state (open/close) is computed from finger-pixel point-cloud centroid distance with task-specific thresholds: 0.10 m for StoreCerealBoxUnderShelf, 0.06 m for PutSpatulaOnTable, 0.12 m for PlaceAppleFromBowlIntoBin.
- Backbone: SVD U-Net, 2.4 B trainable parameters total.
- Optimizer: AdamW (β₁=0.95, β₂=0.999, ε=10⁻⁸, weight decay 10⁻⁶).
- LR: 1 × 10⁻⁶, batch size 4.
- Training compute: 4× NVIDIA RTX A6000 (48 GB), 60 epochs per task; ~15 k extra steps for real-world fine-tuning.
- Inference: EulerEDMSampler, 25 denoising steps; ~30 s per 10-frame chunk on a single RTX 4090.
mIoU = cross-view 3D-mask consistency; FVD-n / FVD-m = Fréchet Video Distance per view; AbsRel / δ₁ = depth error and threshold accuracy.
Task 1: StoreCerealBoxUnderShelf
| Method | mIoU↑ | FVD-n↓ | FVD-m↓ | AbsRel-n↓ | AbsRel-m↓ | δ₁-n↑ | δ₁-m↑ |
|---|---|---|---|---|---|---|---|
| OURS | 0.70 | 411.20 | 561.43 | 0.06 | 0.11 | 0.95 | 0.92 |
| OURS w/o MV attn | 0.41 | 497.43 | 607.73 | 0.15 | 0.31 | 0.75 | 0.66 |
| 4D Gaussian (SoM) | 0.39 | 1208.00 | 1094.98 | 0.20 | 0.31 | 0.74 | 0.63 |
| SVD | – | 977.06 | 743.25 | – | – | – | – |
| SVD w/ MV attn | – | 941.73 | 653.44 | – | – | – | – |
Task 2: PutSpatulaOnTable
| Method | mIoU↑ | FVD-n↓ | FVD-m↓ | AbsRel-n↓ | AbsRel-m↓ | δ₁-n↑ | δ₁-m↑ |
|---|---|---|---|---|---|---|---|
| OURS | 0.69 | 377.68 | 257.70 | 0.03 | 0.07 | 0.98 | 0.97 |
| OURS w/o MV attn | 0.44 | 451.54 | 302.29 | 0.10 | 0.33 | 0.89 | 0.41 |
| 4D Gaussian | 0.46 | 1241.13 | 815.77 | 0.33 | 0.30 | 0.43 | 0.37 |
| SVD | – | 370.92 | 417.56 | – | – | – | – |
| SVD w/ MV attn | – | 536.02 | 445.68 | – | – | – | – |
Task 3: PlaceAppleFromBowlIntoBin
| Method | mIoU↑ | FVD-n↓ | FVD-m↓ | AbsRel-n↓ | AbsRel-m↓ | δ₁-n↑ | δ₁-m↑ |
|---|---|---|---|---|---|---|---|
| OURS | 0.64 | 490.88 | 366.98 | 0.06 | 0.07 | 0.95 | 0.96 |
| OURS w/o MV attn | 0.26 | 597.05 | 573.73 | 0.14 | 0.49 | 0.76 | 0.30 |
| 4D Gaussian | 0.44 | 1396.10 | 1191.40 | 0.18 | 0.16 | 0.80 | 0.81 |
| SVD | – | 659.52 | 628.01 | – | – | – | – |
| SVD w/ MV attn | – | 812.94 | 766.52 | – | – | – | – |
Multi-task real-world dataset
| Method | mIoU↑ | FVD-n↓ | FVD-m↓ | AbsRel-n↓ | AbsRel-m↓ | δ₁-n↑ | δ₁-m↑ |
|---|---|---|---|---|---|---|---|
| OURS | 0.56 | 384.26 | 331.00 | 0.10 | 0.20 | 0.93 | 0.89 |
| OURS w/o MV attn | 0.32 | 694.23 | 601.93 | 0.14 | 0.34 | 0.90 | 0.81 |
| 4D Gaussian | 0.00 | 2102.52 | 2608.94 | 0.32 | 1.80 | 0.78 | 0.22 |
| SVD | – | 605.10 | 648.41 | – | – | – | – |
| SVD w/ MV attn | – | 1151.05 | 931.33 | – | – | – | – |
The model wins on every task on cross-view consistency and on most depth metrics, while remaining competitive on RGB FVD.
30 rollouts/task, novel object poses + held-out cameras.
| Method | Task 1 | Task 2 | Task 3 | Avg |
|---|---|---|---|---|
| Dreamitate | 0.10 | 0.17 | 0.10 | 0.12 |
| Diffusion Policy (DP) | 0.10 | 0.27 | 0.00 | 0.12 |
| DP3 (3D point cloud) | 0.23 | 0.27 | 0.00 | 0.25 |
| OURS | 0.73 | 0.67 | 0.53 | 0.64 |
The paper's headline claim is >50 % absolute average gain over the video-generation baseline Dreamitate (0.64 vs 0.12). Over the strongest behaviour-cloning baseline DP3 (0.25) the margin is +0.39 absolute.
| Method | Inference | Train memory | Trainable params |
|---|---|---|---|
| OURS | 30.0 s | 47 GB | 2.4 B |
| OURS w/o MV attn | 29.3 s | 46.5 GB | 2.38 B |
| 4D Gaussian | 2 s | 2.8 GB | 856 k |
| SVD (Dreamitate) | 13.4 s | 45.8 GB | 1.54 B |
| SVD w/ MV attn | 15.1 s | 46.3 GB | 2.4 B |
Without retraining, the two-view model extends to a 3rd+ view at inference by running an extra forward pass between the reference view and each new view (cost grows ~linearly in #views). On StoreCerealBoxUnderShelf, generating three views jointly degrades only gracefully: View 2 (w/ View 1) mIoU 0.62 / FVD 566.86 / AbsRel 0.09 / δ₁ 0.95; View 3 (w/ View 1) mIoU 0.54 / FVD 662.83 / AbsRel 0.12 / δ₁ 0.89.
- Multi-view cross-attention (OURS w/o MV attn). Removing cross-attention drops mIoU from 0.70 → 0.41 on Task 1 and 0.64 → 0.26 on Task 3, and inflates AbsRel-m from 0.11 → 0.31. The model fails to learn the v_m → v_n projection without explicit cross-decoder feature flow.
- Adding cross-attention to SVD only on RGB (SVD w/ MV attn). RGB-side cross-attention without pointmap supervision still yields 3D-inconsistent gripper pose, confirming cross-attention is necessary but insufficient — pointmap supervision is the load-bearing piece.
- 4D Gaussian baseline. Optimising a Shape-of-Motion 4D Gaussian on a single SVD-generated video underperforms drastically on novel views (e.g. mIoU = 0.00 on real-world data) because the underlying generated video is fixed-view.
- DP / DP3 baselines. Diffusion Policy fails on novel viewpoints despite multi-view training inputs because it learns no explicit cross-view geometry. DP3's global point-cloud helps slightly on cereal-box but not on smaller objects (spatula, apple).
- Gripper-region reweighting (Eq. 4). Doubles loss on gripper pixels for higher pose-tracking accuracy; presented as an architecture choice rather than separately ablated.
The authors explicitly call out:
- Multi-view RGB-D training data with varied camera poses is easy to obtain in simulation but hard in the real world due to hardware and calibration constraints. Real-world depth quality is also a problem; they suggest leveraging RGB-only depth estimators (FoundationStereo, VGGT) in future work.
- Inference is slow (~30 s per 10-frame chunk) compared to end-to-end behaviour cloning, making closed-loop control awkward. Faster substrates suggested: pyramidal flow matching, autoregressive video transformers, one-step diffusion (Diffusion Adversarial Post-Training).
- Open-loop deployment: the robot executes generated trajectories without re-planning until a chunk completes — drift and contact errors are not corrected mid-chunk.
- Action extraction depends on FoundationPose + SAM2 with a CAD model of the gripper; failures in any of these propagate.
This work pushes "video-as-world-model" from 2D-temporal toward genuinely 4D / geometry-aware. Three things distinguish it from concurrent video-policy work:
- No camera poses at inference. A practical point: deploying a video world model in a new lab requires only an RGB-D image per camera, not extrinsics.
- Pointmap-domain supervision rather than RGB-only multi-view supervision — the authors show in ablation that RGB cross-attention alone (SVD w/ MV attn) is not enough for 3D consistency.
- Direct gripper-pose recovery via FoundationPose sidesteps the noisy "video → inverse-dynamics" step that has plagued video-conditioned policies (Dreamitate, Du et al. UniPi style).
Within the 2026 video-world-model cluster:
- ViPRA and Vid2World focus on action prediction from 2D video; this work supplies the 3D substrate they were missing.
- Genie Envisioner is a generalist video-as-world agent — adding pointmap heads in this style would plausibly improve its viewpoint generalisation.
- Ctrl-World is closed-loop and faster; this work is open-loop and 3D-aware. The two are complementary.
The 64 % average success rate on novel viewpoints (vs 12–25 % for diffusion-policy / DP3 baselines) is the headline empirical result and the strongest evidence yet that geometric video supervision pays off for downstream control rather than just generation quality.
- OpenReview: https://openreview.net/forum?id=18gC6pZVVc
- Project: https://robot4dgen.github.io/
← Back to ICLR-2026