ICLR 2026 RoboCasa365 - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 · Project: robocasa.ai Category: Data / Benchmark Trend tag: Trend 5
flowchart LR
Proc[Procedural generation] --> K[2,500 kitchen scenes]
K --> Tasks[365 everyday tasks]
Tasks --> Tele[~600 h human teleop]
Tasks --> Synth[~1,600 h synthetic demos]
Tele --> Train[Train generalist policy]
Synth --> Train
Train --> Bench[Evaluation:<br/>multi-task · foundation training · lifelong]
Real teleop data is expensive and hard to scale. Simulation is cheap but historically too brittle and unrealistic to produce policies that transfer to real robots. The gap is closing — but no publicly available simulation dataset existed at the scale required for generalist-policy training.
365 tasks (65 atomic + 300 composite tasks, spanning 60 distinct kitchen activities; ~220 require mobile manipulation) across 2,500 procedurally-generated kitchen environments modeled from 50 real U.S. homes (Zillow listings), totaling ~2,200 hours of demonstration data: 600+ h of human teleop + 1,600+ h of synthetic (MimicGen) demonstrations. Atomic tasks build on 10 foundational skills (pick-and-place, open/close doors and drawers, twist knobs, turn levers, press buttons, insertion, navigation, slide racks, open/close lids). Benchmark supports evaluation in multi-task, foundation-model-training, and lifelong-learning settings.
- Sim-to-real co-training (DROID Panda arm, 4 real kitchen tasks): combining sim pretraining with real fine-tuning ("sim-and-real") raised average success from 61.8% → 79.8% (+18.1%) vs. real-only.
- Pretraining on the 300-task set gives ~3× downstream data efficiency and large gains on unseen composite tasks; training on 300 tasks beats training on 50 (task diversity matters).
- Best foundation-model baseline: GR00T N1.5 reached 43.0% (atomic), 9.6% (composite-seen), 4.4% (composite-unseen) over 300 pretraining tasks. Supported policies include Diffusion Policy, π, and GR00T.
- Synthetic data caveat: naively adding MimicGen synthetic demos did not help — Human300 alone outperformed Human300 + MimicGen (MG60), because synthetic-demo quality is uneven.
First publicly available simulation dataset with the scale, diversity, and quality needed for foundation-model-style training. Key empirical insight: large-scale synthetic data does not straightforwardly augment human teleop — naive mixing can degrade performance, so leveraging it effectively remains an open problem. The benchmark's main value is enabling controlled study (data scale, composition, co-training) of what actually drives generalist-robot performance, for groups without large teleop farms.
- Project: https://robocasa.ai/
- arXiv: https://arxiv.org/abs/2603.04356
- OpenReview: https://openreview.net/forum?id=tQJYKwc3n4
- ICLR poster: https://iclr.cc/virtual/2026/poster/10006981
- GitHub: https://github.com/robocasa/robocasa
Authors: Soroush Nasiriany, Sepehr Nasiriany, Abhiram Maddukuri, Yuke Zhu (UT Austin / NVIDIA Research).
- EgoDex (the human-video data scaling alternative)
← Back to ICLR-2026