ICLR 2026 RoboCasa365 - Heungwoo/research GitHub Wiki

RoboCasa365 — Large-Scale Simulation Framework for Generalist Robots

Venue: ICLR 2026 · Project: robocasa.ai Category: Data / Benchmark Trend tag: Trend 5

Approach diagram

flowchart LR
  Proc[Procedural generation] --> K[2,500 kitchen scenes]
  K --> Tasks[365 everyday tasks]
  Tasks --> Tele[~600 h human teleop]
  Tasks --> Synth[~1,600 h synthetic demos]
  Tele --> Train[Train generalist policy]
  Synth --> Train
  Train --> Bench[Evaluation:<br/>multi-task · foundation training · lifelong]
Loading

Problem

Real teleop data is expensive and hard to scale. Simulation is cheap but historically too brittle and unrealistic to produce policies that transfer to real robots. The gap is closing — but no publicly available simulation dataset existed at the scale required for generalist-policy training.

Method

365 tasks (65 atomic + 300 composite tasks, spanning 60 distinct kitchen activities; ~220 require mobile manipulation) across 2,500 procedurally-generated kitchen environments modeled from 50 real U.S. homes (Zillow listings), totaling ~2,200 hours of demonstration data: 600+ h of human teleop + 1,600+ h of synthetic (MimicGen) demonstrations. Atomic tasks build on 10 foundational skills (pick-and-place, open/close doors and drawers, twist knobs, turn levers, press buttons, insertion, navigation, slide racks, open/close lids). Benchmark supports evaluation in multi-task, foundation-model-training, and lifelong-learning settings.

Results

  • Sim-to-real co-training (DROID Panda arm, 4 real kitchen tasks): combining sim pretraining with real fine-tuning ("sim-and-real") raised average success from 61.8% → 79.8% (+18.1%) vs. real-only.
  • Pretraining on the 300-task set gives ~3× downstream data efficiency and large gains on unseen composite tasks; training on 300 tasks beats training on 50 (task diversity matters).
  • Best foundation-model baseline: GR00T N1.5 reached 43.0% (atomic), 9.6% (composite-seen), 4.4% (composite-unseen) over 300 pretraining tasks. Supported policies include Diffusion Policy, π, and GR00T.
  • Synthetic data caveat: naively adding MimicGen synthetic demos did not help — Human300 alone outperformed Human300 + MimicGen (MG60), because synthetic-demo quality is uneven.

Significance

First publicly available simulation dataset with the scale, diversity, and quality needed for foundation-model-style training. Key empirical insight: large-scale synthetic data does not straightforwardly augment human teleop — naive mixing can degrade performance, so leveraging it effectively remains an open problem. The benchmark's main value is enabling controlled study (data scale, composition, co-training) of what actually drives generalist-robot performance, for groups without large teleop farms.

Links

Authors: Soroush Nasiriany, Sepehr Nasiriany, Abhiram Maddukuri, Yuke Zhu (UT Austin / NVIDIA Research).

Related pages

  • EgoDex (the human-video data scaling alternative)

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️