ICML 2026 MVISTA 4D - Heungwoo/research GitHub Wiki

MVISTA-4D โ€” view-consistent 4D world model with test-time action inference

Venue: ICML 2026 (Poster) Category: World Model Traction (2026-06): 2 citations (arXiv)

Overview of the MVISTA-4D pipeline (Figure 1 from Wang et al., 2026)

Problem

World-model-based "imagine-then-act" is a promising paradigm for robotic manipulation, but existing models either forecast in pure image space โ€” producing geometrically implausible futures โ€” or reason over partial 3D geometry that is sparse and weak in appearance cues. Single-view RGB-D world models suffer from incomplete geometry and brittleness under occlusion, while monocular depth drifts in scale and time. A second gap is action recovery: standard inverse-dynamics models are ill-posed because many actions explain the same perceptual transition, and most pipelines treat actions as per-step signals, ignoring the low-dimensional, temporally correlated structure of real trajectories.

Method

MVISTA-4D is a trajectory-conditioned, geometry-consistent multi-view 4D generative world model built on the WAN2.2 TI2V backbone. From a single-view RGB-D observation plus a text instruction it "imagines" the remaining viewpoints, which are back-projected and fused into a more complete 3D point trajectory.

  • Cross-modality fusion: a learnable modality token tags appearance vs. geometry tokens, and a lightweight local cross-modality attention module (radius-r neighborhood, cost O(Nk)) exchanges RGBโ†”depth features with channel-wise gated residual updates.
  • Cross-view consistency: each camera is encoded in spherical coordinates around a shared look-at point (estimated by least squares), then Fourier-featured (K=2) into a compact 13-D view token. Views are concatenated along the height dimension, naturally supporting a variable number of views.
  • Trajectory conditioning: the full action trajectory is compressed into a latent style code via a frozen TCN-VAE.
  • Test-time action inference: the model first generates a text-only 4D rollout, freezes it, then optimizes a random trajectory latent z (100 backprop steps) so that the trajectory-conditioned generation matches the fixed rollout; the TCN decoder turns z into an action sequence.
  • Residual inverse dynamics (R-IDM): a residual module corrects the decoded trajectory prior using consecutive 3D point sets, anchoring the ill-posed IDM around a plausible trajectory.

Qualitative 4D generation on the RoboTwin dataset; colored boxes mark different viewpoints (Figure 2 from Wang et al., 2026)

Results

Evaluated on RLBench, RoboTwin2, and a real 4-camera robot platform (8,000 / 10,000 collected trajectories; 14 real tasks), with success rate over 100 episodes. On manipulation (Table 3), the full model reaches 72.6% on RLBench and 43.0% on RoboTwin, beating P-ACT (60.4 / 20.5), UniPi* (34.6 / 16.3), 4DGen (47.0 / 40.2), and TesserAct (67.3 / 33.9). Ablations show test-time latent optimization beats an action-head prior, and the residual IDM beats both "w/o R-IDM" and a full point-based IDM. On the real robot it outperforms TesserAct on most tasks (e.g., Open Drawer 56 vs. 37, Place Fruits 23 vs. 17).

Significance

MVISTA-4D unifies multi-view 4D scene imagination with trajectory-level action recovery, addressing two long-standing weaknesses of world-model manipulation โ€” occlusion-induced geometry holes and ill-posed inverse dynamics. Its variable-view design and test-time trajectory optimization point toward more robust, geometry-grounded imagine-then-act policies in the 2026 VLA landscape.

Links

โ† Back to ICML-2026