ICML 2026 Scaling Real World Robot Policy Evaluation - Heungwoo/research GitHub Wiki
dWorldEval — Scaling robot policy evaluation with an action-centric discrete-diffusion world model
Venue: ICML 2026 (Poster) Category: World Model Affiliations: Yaxuan Li, Zhongyi Zhou, Yefei Chen, Yaokai Xue, Yichen Zhu (Current Robotics; University of Toronto) Traction (2026-06): 1 citation (arXiv)

Problem
Evaluating generalist robot policies across thousands of tasks and environments is infeasible with real-world rollouts or asset-heavy simulators. Generative world models are a scalable alternative, but they are not yet reliable evaluation proxies for two reasons. First, they struggle on out-of-distribution actions: trained mostly on successful demonstrations, they ignore erroneous actions and hallucinate success. Second, physical inconsistency produces unrealistic artifacts (objects warping or vanishing on contact). The authors argue the bottleneck is architectural: most prior models adapt video-generation backbones not natively designed to take actions as input, so actions are injected only as weak auxiliary conditions (cross-attention, AdaLN) and get overridden by dominant visual priors.
Method
dWorldEval is a world model based on Masked Discrete Diffusion (MDD). It maps all modalities — vision, language, and robot actions — into a unified token space and denoises them with a single transformer network, making actions first-class tokens rather than auxiliary conditions. Three components support reliable evaluation:
- Action control via a unified token sequence so the model strictly adheres to control signals.
- Sparse keyframe memory (K=4 keyframes) to maintain spatiotemporal consistency over long horizons.
- A discrete progress token (Progress-as-text) indicating degree of task completion; at inference the model jointly denoises future observations and the progress score, automatically declaring success when progress reaches 1 — removing the need for an external oracle.
Models predict multi-view outcomes at 256² resolution conditioned on 128² keyframes, with progress-token loss weight 2 vs. 1 for visual tokens, prediction horizon matched to the action chunk (2–8), and 16-step parallel decoding at inference. The authors also propose a Δ-LPIPS metric to quantify action controllability.

Results
Evaluated on LIBERO, RoboTwin (ARX arm), and a real bimanual AgileX platform (datasets augmented with ~1k failed rollouts each to enable failure-aware scoring), against WorldEval, WorldGym, and Ctrl-World:
- Policy-ranking reliability: dWorldEval achieves strong correlation with real execution success and minimal rank violation (MMRV = 0.013) on LIBERO single-view, where baselines reach MMRV up to 0.039. Correlations with true success rates are r = 0.910 (LIBERO multi-view), r = 0.927 (RoboTwin), and r = 0.918 (real world); ablating history drops multi-view r to 0.786.
- Consistency (round-trip LPIPS, lower better): at action horizon H=20, dWorldEval = 0.243 vs. Ctrl-World 0.370, WorldGym 0.482, WorldEval 0.531.
- Fidelity (Table 3): consistently low Δ-LPIPS (~0.30–0.37) across LIBERO, RoboTwin, and real-world, with third-person/wrist views, real-world remaining comparable to simulation.
- Controllability: where WorldGym self-corrects a missed grasp and WorldEval hallucinates objects, dWorldEval faithfully reproduces the failure under suboptimal actions.
Significance
By making actions first-class tokens in a discrete-diffusion world model — rather than weak conditioning on a video backbone — dWorldEval becomes faithful enough to render execution failures and rank heterogeneous policies (π0 checkpoints, DexVLA, Diffusion Policy) across simulation and real robots, with automatic success scoring. It points toward a new architectural paradigm for building scalable robotics evaluation simulators.
Links
- arXiv: 2604.22152
- ICML 2026: https://icml.cc/virtual/2026/poster/65898
← Back to ICML-2026