ICML 2026 Predicting What Matters - Heungwoo/research GitHub Wiki

Predicting What Matters — Robust Generalist Robot Policies via Future Semantic Masks

Venue: ICML 2026 (Poster) Category: World Model Affiliations: Authors: Yunfan Lou, Xiaowei Chi, Xiaojie Zhang, Zezhong Qian, Chengxuan Li, Rongyu Zhang, Yaoxu Lyu, Guoyu Song, Chuyao Fu, Haoxuan Xu, Pengwei Wang, Shanghang Zhang

Problem

World models derived from large-scale video generative pre-training have emerged as a promising paradigm for generalist robot policy learning. However, such models often focus on high-fidelity RGB video prediction, and this objective can cause overfitting to irrelevant factors — dynamic backgrounds, illumination changes, and other visual noise — rather than the physical dynamics that actually matter for control. The result is brittle generalization when textures or appearances shift between training and deployment.

Method

The paper introduces the Mask World Model (MWM), which leverages video diffusion architectures to predict the evolution of semantic masks instead of pixels. By forecasting how segmentation masks change over time rather than reconstructing full RGB frames, MWM imposes a geometric information bottleneck that forces the model to capture essential physical dynamics and contact relations while filtering out visual noise. This mask-based world model is then integrated with a diffusion-based policy head for end-to-end robotic control.

flowchart LR
    O[Observation frames] --> SM[Semantic mask extraction]
    SM --> VD[Video-diffusion Mask World Model<br/>predicts future mask evolution]
    VD -->|geometric information bottleneck| F[Future semantic masks<br/>physical dynamics + contacts]
    F --> DP[Diffusion policy head]
    O --> DP
    DP --> A[Robot actions]
Loading

By predicting what matters — object geometry, contact relations, and dynamics — instead of what is rendered, the model decouples control-relevant structure from appearance-level distractors.

Results

The authors demonstrate effectiveness on simulation benchmarks LIBERO and RLBench, where MWM outperforms existing RGB-based world-model approaches. In real-world testing and robustness evaluations, the model exhibits superior generalization with robust resilience to texture-information loss, confirming that the semantic-mask bottleneck retains the dynamics needed for control while discarding appearance noise. (Detailed per-task numbers were not available from the abstract source.)

Significance

The work argues that for generalist robot policies, what to predict is as important as how well it is predicted. Replacing pixel-level RGB forecasting with semantic-mask forecasting reframes the world-model objective around geometry and physical interaction, yielding policies that are markedly more robust to visual distribution shift — a key obstacle for deploying video-pretrained world models on real robots.

Links

← Back to ICML-2026

⚠️ **GitHub.com Fallback** ⚠️