ICML 2026 MVISTA 4D - Heungwoo/research GitHub Wiki
MVISTA-4D โ view-consistent 4D world model with test-time action inference
Venue: ICML 2026 (Poster) Category: World Model Traction (2026-06): 2 citations (arXiv)

Problem
World-model-based "imagine-then-act" is a promising paradigm for robotic manipulation, but existing models either forecast in pure image space โ producing geometrically implausible futures โ or reason over partial 3D geometry that is sparse and weak in appearance cues. Single-view RGB-D world models suffer from incomplete geometry and brittleness under occlusion, while monocular depth drifts in scale and time. A second gap is action recovery: standard inverse-dynamics models are ill-posed because many actions explain the same perceptual transition, and most pipelines treat actions as per-step signals, ignoring the low-dimensional, temporally correlated structure of real trajectories.
Method
MVISTA-4D is a trajectory-conditioned, geometry-consistent multi-view 4D generative world model built on the WAN2.2 TI2V backbone. From a single-view RGB-D observation plus a text instruction it "imagines" the remaining viewpoints, which are back-projected and fused into a more complete 3D point trajectory.
- Cross-modality fusion: a learnable modality token tags appearance vs. geometry tokens, and a lightweight local cross-modality attention module (radius-r neighborhood, cost O(Nk)) exchanges RGBโdepth features with channel-wise gated residual updates.
- Cross-view consistency: each camera is encoded in spherical coordinates around a shared look-at point (estimated by least squares), then Fourier-featured (K=2) into a compact 13-D view token. Views are concatenated along the height dimension, naturally supporting a variable number of views.
- Trajectory conditioning: the full action trajectory is compressed into a latent style code via a frozen TCN-VAE.
- Test-time action inference: the model first generates a text-only 4D rollout, freezes it, then optimizes a random trajectory latent z (100 backprop steps) so that the trajectory-conditioned generation matches the fixed rollout; the TCN decoder turns z into an action sequence.
- Residual inverse dynamics (R-IDM): a residual module corrects the decoded trajectory prior using consecutive 3D point sets, anchoring the ill-posed IDM around a plausible trajectory.

Results
Evaluated on RLBench, RoboTwin2, and a real 4-camera robot platform (8,000 / 10,000 collected trajectories; 14 real tasks), with success rate over 100 episodes. On manipulation (Table 3), the full model reaches 72.6% on RLBench and 43.0% on RoboTwin, beating P-ACT (60.4 / 20.5), UniPi* (34.6 / 16.3), 4DGen (47.0 / 40.2), and TesserAct (67.3 / 33.9). Ablations show test-time latent optimization beats an action-head prior, and the residual IDM beats both "w/o R-IDM" and a full point-based IDM. On the real robot it outperforms TesserAct on most tasks (e.g., Open Drawer 56 vs. 37, Place Fruits 23 vs. 17).
Significance
MVISTA-4D unifies multi-view 4D scene imagination with trajectory-level action recovery, addressing two long-standing weaknesses of world-model manipulation โ occlusion-induced geometry holes and ill-posed inverse dynamics. Its variable-view design and test-time trajectory optimization point toward more robust, geometry-grounded imagine-then-act policies in the 2026 VLA landscape.
Links
- arXiv: 2602.09878
- ICML 2026: https://icml.cc/virtual/2026/poster/65571
โ Back to ICML-2026