RSS 2026 Simulation Distillation - Heungwoo/research GitHub Wiki
Simulation Distillation: Pretraining World Models in Simulation for Rapid Real-World Adaptation
Venue: RSS 2026 (Sydney, Jul 13โ17) ยท Session: World Models & Memory ยท paper #17 Authors: Jacob Levy, Tyler Westenbroek, Kevin Huang, Fernando Palafox, Patrick Yin, Shayegan Omidshafiei, Dong-Ki Kim, Abhishek Gupta, David Fridovich-Keil arXiv: 2603.15759 ยท program page
Summary compiled from the arXiv paper (v2); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Figure 1 contrasts zero-shot sim-to-real failures (left column: peg insertion, table leg, slippery slope, foam) with successful executions after SimDist adaptation (right): the UR5e completes both precise assembly tasks and the Unitree Go2 traverses low-friction PTFE panels and memory foam, after only 15โ30 minutes of real-world data.
Problem
End-to-end policy finetuning in new real-world environments is inefficient and brittle โ model-free RL finetuning often collapses via catastrophic forgetting, especially on long-horizon contact-rich tasks. World models enable planning by counterfactual reasoning, but training action-conditioned robot world models directly in the real world requires diverse data at impractical scale. The question is how to obtain the coverage and supervision needed for planning-grade world models without collecting it in the real world.
Method
SimDist pretrains a planning-oriented latent world model entirely in simulation and reduces real-world adaptation to supervised system identification. (1) A privileged state-based expert policy, its intermediate checkpoints, and its value function are trained with RL; diverse rollouts are generated by mixing expert and sub-optimal policies with contiguous action perturbations, giving dense reward/value supervision (100k trajectories for manipulation, 100M for the quadruped). (2) The world model โ encoder, history encoder, transformer-based chunked latent dynamics predicting T future states in one forward pass, transformer sequence-to-sequence reward/value heads, base-policy head, and no pixel reconstruction โ is pretrained on this data. (3) At deployment, MPPI planning runs with the encoder, reward model, and value function frozen; only the latent dynamics model is finetuned on real-world prediction losses, iterating collection and finetuning. Manipulation uses a UR5e with three 224ร224 RGB cameras, ResNet-18 encoders, a 64-d latent, and H=T=5 at 5 Hz; the Go2 quadruped uses H=T=25 planned at 50 Hz on an RTX 4090M laptop.
Results
Across four real-world tasks โ Peg Insertion and Table Leg assembly (narrow/wide initial-condition grids), quadruped Slippery Slope (PTFE panels), and Foam โ SimDist reliably improves with only 15โ30 minutes of real-world data and typically reaches scores about 2ร higher than any baseline (RLPD, IQL, SGFT-SAC, Diffusion Policy, ฯ0.5), which stagnate or collapse during finetuning. Task throughput improves ~1.5โ2ร over zero-shot. On Slippery Slope the latent-dynamics loss on a held-out real trajectory drops from 0.076 (pretrained) to 0.019 (finetuned). Simulation ablations: full SimDist reaches 0.90/0.85 success (Peg/Table Leg) vs 0.10/0.05 with expert-only pretraining data; unfreezing the encoder or the value function during adaptation destroys performance, and adding pixel reconstruction drops manipulation to 0.32/0.21.
Significance
A crisp decomposition result for sim-to-real: task structure (representations, rewards, values) transfers across the dynamics gap, so only the dynamics model needs real-world correction โ turning unstable real-world RL into stable supervised finetuning. Directly relevant to the planning-oriented vs generative world-model discussion in Review-World-Models.
โ Back to RSS 2026 survey ยท RSS-2026-Papers ยท Home