ICML 2026 Structured 4D Latent World Model - Heungwoo/research GitHub Wiki
Structured 4D Latent World Model for Robot Planning — Predicting how 3D scene structure evolves, then planning through it
Venue: ICML 2026 (Poster) Category: World Model Affiliations: Zhiyi Li, Peilin Wu, Xiaoshen Han, Ruojin Cai, Yilun Du
Problem
Learned world models are emerging as a powerful paradigm in robotics, offering a promising path toward task generalization, long-horizon planning, and flexible decision-making. However, most learned world models operate in 2D video space. Predicting raw pixel futures couples scene dynamics to view-specific appearance, which can produce futures that look plausible frame-by-frame but are physically inconsistent or incoherent across viewpoints — a poor substrate for planning that must reason about where objects actually are in 3D.
Method
This work represents scenes in a structured latent space that encodes the scene's 3D structure rather than relying solely on 2D video. The model predicts how that 3D structure evolves over time (4D), conditioned on observations and a text instruction. Because the representation captures the scene holistically, it can be decoded into diverse 3D formats, enabling a more complete and physically consistent scene understanding than pixel-space rollouts.
For control, the model is used as a planner: it imagines future structured-latent states for a given instruction, and a goal-conditioned inverse dynamics module converts those predicted futures into executable robot actions.
flowchart LR
O[Observations] --> E[Encode to structured 4D latent]
T[Text instruction] --> E
E --> W[World model: predict 3D structure evolution]
W --> D[Decode to 3D formats]
W --> F[Predicted future latent state]
F --> IDM[Goal-conditioned inverse dynamics]
O --> IDM
IDM --> A[Robot actions]
Results
The authors report that predicted scenes have superior visual quality, physical consistency, and multi-view coherence compared to state-of-the-art video-based planners. The structured-latent planner achieves strong performance on manipulation tasks, generalizes to novel visual conditions, and is demonstrated in a successful real-world robotic implementation. (Quantitative tables are not available from the ICML abstract page; numbers will be added once the full paper is released.)
Significance
By moving the world model's prediction target from 2D pixels to a structured 3D latent that can be decoded into multiple 3D formats, this approach aims to give planning a physically grounded, view-consistent imagination. Pairing that imagination with a goal-conditioned inverse dynamics module yields an actionable planner — a concrete step toward world-model-based robot control that holds up under novel viewpoints and transfers to real hardware.
Links
- ICML 2026: https://icml.cc/virtual/2026/poster/63054
← Back to ICML-2026