ICLR 2026 RoboArena - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Category: Benchmark Trend tag: Trend 8 (evaluation) Authors: Yash Jangir, Yidi Zhang, Pang-Chi Lo, Kashu Yamazaki, Chenyu Zhang, Kuan-Hsun Tu, Tsung-Wei Ke, Lei Ke, Yonatan Bisk, Katerina Fragkiadaki (Carnegie Mellon University, National Taiwan University) arXiv: 2510.23571 — project: robotarenainf.github.io
Note: This is RobotArena ∞ (real-to-sim, CMU/NTU). Do not confuse with the separate RoboArena (Atreya, Pertsch et al., arXiv 2506.18123), which is a distributed real-world double-blind evaluation network.
flowchart LR
Data[Robot video demos<br/>BridgeV2 / DROID / RH20T] --> Auto[Auto real-to-sim translation<br/>VLM + 2D→3D generation<br/>+ differentiable rendering]
Auto --> Sim[Simulated counterpart<br/>digital twin]
Sim --> Pert[Perturb textures /<br/>object placements]
Pol[Test policy] --> Pert
Pert --> Roll[Rollout]
Roll --> VLM[VLM-guided scoring<br/>task progression]
Roll --> Human[Crowd preference judgments<br/>Bradley–Terry ranking]
VLM --> Score[Benchmark score / leaderboard]
Human --> Score
Score -. scalable — translate more demos on demand .-> Auto
Real-world evaluation of generalist policies is labor-intensive, slow, unsafe at scale, and hard to reproduce, while fixed sim suites (LIBERO, CALVIN, SIMPLER) are saturating — methods like PLD push success into the high 90s — and were never designed for broad cross-environment evaluation. Judging "success" also hinges on nuanced human judgments of execution quality. The field needs benchmarks that scale without per-scene physical setup.
Real-to-sim translation pipeline: automatically converts video demonstrations from existing robot datasets (BridgeV2 → BridgeSim, DROID → DROIDSim, RH20T → RH20TSim) into simulated counterparts, leveraging vision-language models, 2D-to-3D generative modeling, and differentiable rendering to build digital twins of tabletop manipulation scenes. Environments are systematically perturbed (textures, object placements) to probe generalization. Policies are then evaluated two ways: VLM-guided scoring of task progression, and scalable crowdsourced human preference judgments aggregated into Bradley–Terry rankings — turning heavy human labor into lightweight pairwise comparisons. The pipeline is automated, so the benchmark is continuously evolving / scalable (the ∞).
Translates three major robot datasets into simulated counterparts and produces VLM-based progression scores plus Bradley–Terry human-preference rankings. A real-vs-sim consistency check on a manipulation task showed agreement across policies (e.g., RoboVLM and SpatialVLA succeeded in both real and sim, Octo failed in both); comprehensive numerical real-to-sim correlation coefficients are not reported. Positioned by the authors as among the most extensive robot evaluations to date.
Breaks the "fixed benchmark suite" paradigm by making evaluation environments themselves generated from real demos. Combines automated VLM scoring with human preference signal, addressing the difficulty of defining manipulation "success" by automated metrics alone. Validity of sim rankings against real-robot rankings (large-scale correlation) remains the main open question.
- arXiv: 2510.23571 — https://arxiv.org/abs/2510.23571
- Project / leaderboard: https://robotarenainf.github.io/
- OpenReview: https://openreview.net/forum?id=OutljIofvS
← Back to ICLR-2026