RSS 2026 Interactive World Simulator for Robot - Heungwoo/research GitHub Wiki

Interactive World Simulator for Robot Policy Training and Evaluation

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: World Models & Memory · paper #18 Authors: Yixuan Wang, Rhythm Syed, Fangyu Wu, Mengchao Zhang, Aykut Onol, Jose Barreiros, Hooshang Nayyeri, Tony Dear, Huan Zhang, Yunzhu Li arXiv: 2603.08546 · program page

Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Interactive World Simulator overview (Figure 1 of arXiv 2603.08546, © the authors)

Figure 1 in three panels: (left) real-world ALOHA robot interaction data across six tasks (mug grasping, rope routing, rope collecting, T pushing, box packing, pile sweeping); (middle) the action-conditioned video model rolled out autoregressively, with a realism-vs-speed scatter showing "Ours (15 FPS)" above Dreamer4, Cosmos, UVA and DINO-WM in PSNR, and 10-minute stable rollouts (t = 0 to 6000); (right) the two applications — near-flat success curves across 100% world-sim to 100% real training-data mixtures, and a task-score correlation plot between world-sim and real evaluation.

Problem

Action-conditioned video prediction models ("world models") are promising for robot policy training and evaluation, but existing ones are either too slow for interactive use (heavy multi-step diffusion needing enterprise GPUs) or drift and accumulate errors over long-horizon rollouts, so they cannot serve as faithful surrogates for demonstration collection or reproducible policy evaluation.

Method

The Interactive World Simulator is built from a moderate-sized robot interaction dataset in two stages: (1) an autoencoder with a CNN encoder and a consistency-model decoder (CTM-style training) maps 128×128 RGB frames to compact 2D latents; (2) with the autoencoder frozen, an action-conditioned latent dynamics model — also a consistency model, instantiated as 3D-conv blocks with FiLM modulation and spatiotemporal attention — is trained with next-frame supervision, injecting small noise into observation contexts for robustness. Inference is autoregressive with a shifting fixed-length context window, giving stable rollouts of over 10 minutes at 15 FPS on a single RTX 4090. Data: ~600 play episodes per real task (~6 hours of collection each) on an ALOHA bimanual robot, plus 10,000 scripted episodes for a MuJoCo T-pushing task; the mug-grasping model is only 176.02 MB, trained in ~6 h (stage 1) + ~12 h (stage 2) on one H200. Humans teleoperate inside the simulator (keyboard or kinematic device) to collect synthetic demonstrations, and policies can be rolled out inside it for evaluation.

Results

Over 192-step action-conditioned rollouts aggregated across seven tasks, it beats Cosmos, UVA, Dreamer4 and DINO-WM on all metrics, e.g. PSNR 25.82 vs 20.81 (Dreamer4) and 17.79 (DINO-WM), and FVD 243.20 vs 799.34 (Cosmos) and 1747–2213 for the rest. Policies (DP, ACT, π0, π0.5) trained on 100-episode mixtures from 100% simulator data to 100% real data perform comparably: DP scores 87.9% with pure simulator data vs 90.3% with pure real data; ACT 76.2% vs 73.6%; π0.5 rises from 73.1% to 88.8% with more real data. Policy scores in the simulator correlate strongly with real-world scores across four tasks (r = 0.8455–0.9908).

Significance

Shows that lightweight consistency-model world models can replace both the physical robot for demonstration collection and much of real-world evaluation, with quantified sim-to-real ranking fidelity — a practical answer to the evaluation bottleneck highlighted by the Large Behavior Models analysis. Connects to the wiki's [Review-World-Models]] thread and to evaluation-methodology discussions in [RSS 2026 survey.

← Back to RSS 2026 survey · RSS-2026-Papers · Home