ICML 2026 RoboTwin 20 - Heungwoo/research GitHub Wiki
RoboTwin 2.0 — A Scalable Data Generator and Benchmark with Strong Domain Randomization for Bimanual Manipulation
Venue: ICML 2026 (Poster) Category: Benchmark Traction (2026-06): 239 citations (arXiv)

Problem
Simulation-based data synthesis has become a powerful way to boost real-world robotic manipulation, but existing synthetic datasets are still inadequate for robust bimanual manipulation for two reasons: (1) there is no efficient, scalable method to generate data for novel tasks, and (2) simulation environments are oversimplified and fail to capture real-world complexity. The result is poor sim-to-real transfer and weak generalization to cluttered, visually varied real scenes.
Method
RoboTwin 2.0 is a scalable closed-loop simulation framework for automated, large-scale generation of diverse, realistic data plus unified evaluation protocols for dual-arm manipulation.
- RoboTwin-OD object library. A large-scale dataset of 731 object instances across 147 categories, each annotated with semantic and manipulation-relevant labels.
- Expert data synthesis pipeline. Grounded on RoboTwin-OD and a predefined skill API, an MLLM agent automatically synthesizes executable task programs from natural-language instructions, refined with simulation-in-the-loop feedback. The paper uses DeepSeek-V3 for program synthesis and moonshot-v1-32k-vision-preview for multimodal error localization and verification.
- Five-axis structured domain randomization. To improve sim-to-real transfer, trajectories are randomized along clutter, lighting, background, tabletop height (up to 3 cm), and language instructions, raising data diversity and policy robustness.
- Scale. The framework is instantiated across 50 dual-arm tasks spanning five robot embodiments, with over 100,000 pre-collected domain-randomized expert trajectories.
Results
- Code generation. RoboTwin 2.0's MLLM + multimodal-feedback refinement yields a 10.9% gain in code-generation success rate over the baseline, consistently matching or beating RoboTwin 1.0 across the per-task breakdown.
- Few-shot real-world transfer. A VLA model fine-tuned on RoboTwin 2.0 data achieves a 367% relative improvement (42.0% vs. 9.0%) on unseen-scene real-world tasks — i.e., the ~3.6× few-shot gain over a low-demo baseline.
- Zero-shot generalization. Models trained exclusively on the synthetic data attain a 228% relative gain (~2.2×), demonstrating strong generalization with no real-world supervision.
- Simulation benchmark. Across the 50-task benchmark (e.g., a 13-task sampled table reporting ACT, DP, DP3, RDT, and Pi0 under Easy/Hard settings), domain randomization markedly improves robustness, especially in the harder, cluttered settings.
The data generator, benchmark, pre-collected dataset, and code are all released.
Significance
RoboTwin 2.0 turns synthetic-data generation for bimanual manipulation into a largely automated, scalable pipeline: an MLLM writes and verifies task code in the loop, and aggressive five-axis domain randomization closes much of the sim-to-real gap. By coupling a 731-object library, 50 tasks across five embodiments, 100K+ trajectories, and a unified evaluation protocol, it provides a foundation for reproducible benchmarks and scalable sim-to-real pipelines — and the strong few-shot/zero-shot transfer results show synthetic data alone can drive robust real-world policies.
Links
- arXiv: 2506.18088
- ICML 2026: https://icml.cc/virtual/2026/poster/62192
← Back to ICML-2026