RSS 2026 RoboLab - Heungwoo/research GitHub Wiki

RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Datasets and Benchmarks · paper #96 Authors: Xuning Yang, Rishit Dagli, Alex Zook, Hugo Hadfield, Ankit Goyal, Stan Birchfield, Fabio Ramos, Jonathan Tremblay (NVIDIA; University of Toronto; University of Sydney) arXiv: 2604.09860 · program page

Summary compiled from the arXiv paper (v3); all numbers quoted from the paper. The camera-ready benchmark is named RoboLab-120 (120 tasks) — an expansion of the ~80-task version in the original program abstract. Trend context: RSS 2026 survey.

RoboLab benchmarking framework and policy-analysis overview (Figure 1 of arXiv 2604.09860, © the authors)

Figure 1 overview. Top row: RoboLab's robot- and policy-agnostic pipeline for human-authored and LLM-enabled scene, task, and environment generation. Bottom row: the policy-performance analysis tools — multiple evaluation axes, discrete failure-event metrics, Neural-Posterior-Estimation sensitivity analysis, and benchmark-level correlation with real-world results (Spearman ρ = 1, Pearson r = 0.68).

Problem

Simulation benchmarks for generalist robot policies saturate quickly and usually share the same domain between training and evaluation, which trivializes success rates and hides real generalization and failure modes. RoboLab asks (1) how well simulation behavior predicts real-world policy performance and (2) which external factors most strongly perturb that behavior.

Method

RoboLab is a benchmarking framework built on IsaacLab that enables human-authored and AI-enabled generation of photorealistic, physically realistic scenes and tasks in a robot- and policy-agnostic way, deferring embodiment-specific binding to runtime. The accompanying RoboLab-120 benchmark contains 120 hand-curated pick-and-place tasks spanning three difficulty levels (65 simple, 38 moderate, 18 complex) and three competency axes (91 visual, 36 procedural, 44 relational). To avoid rewarding sim-domain overfitting, all evaluated policies are fine-tuned only on the real-world DROID dataset (7-DOF Franka, Robotiq 2F-85), and each task is run for N=10 episodes. A suite of analysis tools reports subtask "score" (partial credit), discrete failure events (wrong object grasped, object dropped, gripper collision), and Bayesian sensitivity analysis via Neural Posterior Estimation.

Results

On RoboLab-120, current SOTA policies perform poorly: π0.5 leads with 28.0% overall success / 0.43 score, followed by π0-FAST (15.5%), GR00T N1.6 (7.2%), π0 (5.0%), and PaliGemma (3.4%). The success/score gap is large on hard tasks (π0.5 drops to 13.5% success but keeps a 0.44 score on complex tasks), indicating policies reach partial milestones but fail late. Policies are sensitive to language: π0.5 falls from 28.0% (default) to 15.3% (vague) instructions. They are robust to lighting (90–100%) and table-texture changes (<5% degradation). Across the four policies with both measurements, RoboLab-120 success rankings match real-world RoboArena Elo (Spearman ρ = 1.00; Pearson r = 0.68), supporting simulation as a real-world proxy.

Significance

RoboLab argues that decoupling training and evaluation domains and adding granular, event-level diagnostics turns high-fidelity simulation into a meaningful, scalable proxy for real-world generalist-policy evaluation. Related: Review-VLA-Evaluation.

← Back to RSS 2026 survey · RSS-2026-Papers · Home