RSS 2026 Simulation Distillation - Heungwoo/research GitHub Wiki

Simulation Distillation: Pretraining World Models in Simulation for Rapid Real-World Adaptation

Venue: RSS 2026 (Sydney, Jul 13โ€“17) ยท Session: World Models & Memory ยท paper #17 Authors: Jacob Levy, Tyler Westenbroek, Kevin Huang, Fernando Palafox, Patrick Yin, Shayegan Omidshafiei, Dong-Ki Kim, Abhishek Gupta, David Fridovich-Keil arXiv: 2603.15759 ยท program page

Summary compiled from the arXiv paper (v2); all numbers quoted from the paper. Trend context: RSS 2026 survey.

SimDist zero-shot failures vs real-world improvement (Figure 1 of arXiv 2603.15759, ยฉ the authors)

Figure 1 contrasts zero-shot sim-to-real failures (left column: peg insertion, table leg, slippery slope, foam) with successful executions after SimDist adaptation (right): the UR5e completes both precise assembly tasks and the Unitree Go2 traverses low-friction PTFE panels and memory foam, after only 15โ€“30 minutes of real-world data.

Problem

End-to-end policy finetuning in new real-world environments is inefficient and brittle โ€” model-free RL finetuning often collapses via catastrophic forgetting, especially on long-horizon contact-rich tasks. World models enable planning by counterfactual reasoning, but training action-conditioned robot world models directly in the real world requires diverse data at impractical scale. The question is how to obtain the coverage and supervision needed for planning-grade world models without collecting it in the real world.

Method

SimDist pretrains a planning-oriented latent world model entirely in simulation and reduces real-world adaptation to supervised system identification. (1) A privileged state-based expert policy, its intermediate checkpoints, and its value function are trained with RL; diverse rollouts are generated by mixing expert and sub-optimal policies with contiguous action perturbations, giving dense reward/value supervision (100k trajectories for manipulation, 100M for the quadruped). (2) The world model โ€” encoder, history encoder, transformer-based chunked latent dynamics predicting T future states in one forward pass, transformer sequence-to-sequence reward/value heads, base-policy head, and no pixel reconstruction โ€” is pretrained on this data. (3) At deployment, MPPI planning runs with the encoder, reward model, and value function frozen; only the latent dynamics model is finetuned on real-world prediction losses, iterating collection and finetuning. Manipulation uses a UR5e with three 224ร—224 RGB cameras, ResNet-18 encoders, a 64-d latent, and H=T=5 at 5 Hz; the Go2 quadruped uses H=T=25 planned at 50 Hz on an RTX 4090M laptop.

Results

Across four real-world tasks โ€” Peg Insertion and Table Leg assembly (narrow/wide initial-condition grids), quadruped Slippery Slope (PTFE panels), and Foam โ€” SimDist reliably improves with only 15โ€“30 minutes of real-world data and typically reaches scores about 2ร— higher than any baseline (RLPD, IQL, SGFT-SAC, Diffusion Policy, ฯ€0.5), which stagnate or collapse during finetuning. Task throughput improves ~1.5โ€“2ร— over zero-shot. On Slippery Slope the latent-dynamics loss on a held-out real trajectory drops from 0.076 (pretrained) to 0.019 (finetuned). Simulation ablations: full SimDist reaches 0.90/0.85 success (Peg/Table Leg) vs 0.10/0.05 with expert-only pretraining data; unfreezing the encoder or the value function during adaptation destroys performance, and adding pixel reconstruction drops manipulation to 0.32/0.21.

Significance

A crisp decomposition result for sim-to-real: task structure (representations, rewards, values) transfers across the dynamics gap, so only the dynamics model needs real-world correction โ€” turning unstable real-world RL into stable supervised finetuning. Directly relevant to the planning-oriented vs generative world-model discussion in Review-World-Models.

โ† Back to RSS 2026 survey ยท RSS-2026-Papers ยท Home