RSS 2026 PolaRiS - Heungwoo/research GitHub Wiki
PolaRiS โ Scalable Real-to-Sim Evaluations for Generalist Robot Policies
Venue: RSS 2026 (Manipulation session) ยท Authors: Arhan Jain, Mingtong Zhang, Kanav Arora, William Chen, โฆ Sergey Levine, Chelsea Finn, Wei-Chiu Ma, Dhruv Shah, Abhishek Gupta, Karl Pertsch (UW / Stanford / PI orbit) Category: Evaluation infrastructure (real-to-sim) Trend tag: RSS 2026 thread 6 โ the evaluation crisis arXiv: 2512.16881 ยท project
Compiled from the verified RSS 2026 abstract and the paper's Fig. 1.
Key figure

Figure 1 of the paper โ the four-part system. (1) Tools: a short video of a real scene passes through the PolaRiS Scene Builder (browser-based) into a simulated evaluation environment via neural reconstruction. (2) Dataset: a small sim-data co-training step turns any generalist policy into a "PolaRiS-ready" one, closing the residual real-to-sim gap. (3) Evaluation: the scatter plot shows the money result โ PolaRiS sim performance vs real-world performance for ฯ0.5, ฯ0-FAST, ฯ0, and PaliGemma policies hugging the diagonal (strong rank correlation). (4) Hub: environments are shared on a HuggingFace hub, the "distributed, democratized evaluation" piece.
Problem
Generalist policies must be evaluated across many scenes and tasks, but real rollouts are slow, stochastic, and irreproducible โ while existing sim benchmarks correlate too weakly with real performance to guide development (the RobotManip manifesto's complaint, from the infrastructure side).
Method
- Neural reconstruction โ interactive simulation: short video scans of real scenes become high-fidelity simulated evaluation environments.
- A simple sim-data co-training recipe bridges residual real-to-sim gaps, enabling zero-shot policy evaluation in unseen simulated environments (no per-environment fine-tuning of the policies under test).
Results (as reported)
- Validation at scale: 600 real-world rollouts paired against 93,000+ simulation rollouts.
- PolaRiS rankings correlate with real-world, non-finetuned generalist policy rankings much more strongly than existing simulated benchmarks.
- Rapid environment creation โ a path to distributed, democratized evaluation of robot foundation models.
Significance
The most credible answer yet to the field's evaluation bottleneck: instead of arguing about which static benchmark to trust, make benchmark creation cheap and anchored to real scenes. The paired-rollout validation methodology is itself a contribution (most sim benchmarks never measure their own real-world correlation). Complements LIBERO-X (perturbation protocol on the standard benchmark) and RoboLab (#96), and gives the Qwen-RobotManip-style OOD agenda the scalable infrastructure it lacked. The PI-orbit author list suggests this is how the frontier labs intend to evaluate their own generalists.
โ RSS 2026 survey ยท Home