RSS 2026 PolaRiS - Heungwoo/research GitHub Wiki

PolaRiS โ€” Scalable Real-to-Sim Evaluations for Generalist Robot Policies

Venue: RSS 2026 (Manipulation session) ยท Authors: Arhan Jain, Mingtong Zhang, Kanav Arora, William Chen, โ€ฆ Sergey Levine, Chelsea Finn, Wei-Chiu Ma, Dhruv Shah, Abhishek Gupta, Karl Pertsch (UW / Stanford / PI orbit) Category: Evaluation infrastructure (real-to-sim) Trend tag: RSS 2026 thread 6 โ€” the evaluation crisis arXiv: 2512.16881 ยท project

Compiled from the verified RSS 2026 abstract and the paper's Fig. 1.

Key figure

PolaRiS overview (Figure 1 of arXiv 2512.16881, ยฉ the authors)

Figure 1 of the paper โ€” the four-part system. (1) Tools: a short video of a real scene passes through the PolaRiS Scene Builder (browser-based) into a simulated evaluation environment via neural reconstruction. (2) Dataset: a small sim-data co-training step turns any generalist policy into a "PolaRiS-ready" one, closing the residual real-to-sim gap. (3) Evaluation: the scatter plot shows the money result โ€” PolaRiS sim performance vs real-world performance for ฯ€0.5, ฯ€0-FAST, ฯ€0, and PaliGemma policies hugging the diagonal (strong rank correlation). (4) Hub: environments are shared on a HuggingFace hub, the "distributed, democratized evaluation" piece.

Problem

Generalist policies must be evaluated across many scenes and tasks, but real rollouts are slow, stochastic, and irreproducible โ€” while existing sim benchmarks correlate too weakly with real performance to guide development (the RobotManip manifesto's complaint, from the infrastructure side).

Method

  • Neural reconstruction โ†’ interactive simulation: short video scans of real scenes become high-fidelity simulated evaluation environments.
  • A simple sim-data co-training recipe bridges residual real-to-sim gaps, enabling zero-shot policy evaluation in unseen simulated environments (no per-environment fine-tuning of the policies under test).

Results (as reported)

  • Validation at scale: 600 real-world rollouts paired against 93,000+ simulation rollouts.
  • PolaRiS rankings correlate with real-world, non-finetuned generalist policy rankings much more strongly than existing simulated benchmarks.
  • Rapid environment creation โ†’ a path to distributed, democratized evaluation of robot foundation models.

Significance

The most credible answer yet to the field's evaluation bottleneck: instead of arguing about which static benchmark to trust, make benchmark creation cheap and anchored to real scenes. The paired-rollout validation methodology is itself a contribution (most sim benchmarks never measure their own real-world correlation). Complements LIBERO-X (perturbation protocol on the standard benchmark) and RoboLab (#96), and gives the Qwen-RobotManip-style OOD agenda the scalable infrastructure it lacked. The PI-orbit author list suggests this is how the frontier labs intend to evaluate their own generalists.

โ† RSS 2026 survey ยท Home