CoRL 2026 MolmoBOT - Heungwoo/research GitHub Wiki
CoRL 2026 โ MolmoBOT: Large-Scale Simulation Enables Zero-Shot Manipulation
Venue: CoRL 2026 (Austin, TX, Nov 9โ12) ยท Allen Institute for AI (Ai2). Paper: arXiv 2603.16861. Representative of: the sim-only zero-shot-transfer thesis โ enough simulation diversity โ real-world manipulation with no real training data. Companions: World Models ยท Cross-Embodiment ยท CoRL 2026 survey.

1. Problem
Real robot data is expensive and task-specific fine-tuning is the norm. MolmoBOT asks whether simulation alone โ if made large and diverse enough โ can produce a manipulation policy that transfers zero-shot to physical robots, with no real training data and no real fine-tuning.
2. Method
- MolmoBot-Engine โ an open-source procedural pipeline that generates robots, tasks, and scenes in MolmoSpaces (200k+ pre-built houses) on top of the MuJoCo simulator. It randomizes layouts, 6-DoF object poses, lighting, textures/materials, friction/mass/joint damping, camera extrinsics, and injects action noise. Assets come from iTHOR + Objaverse (filtered for graspable, watertight colliders); an expert planner does phase-based grasp/place trajectories (RB-Y1 uses CuRobo).
- MolmoBot-Data โ ~1.8M expert trajectories (~300M frames, ~5,817 robot-hours) across 94,300 unique environments, 11,400+ pickup assets and 9,400+ receptacles. Generated at ~1,024 episodes/GPU-hour (100ร A100-80GB, ~4,500 GPU-hours total).
- Policies โ MolmoBot (Molmo2 VLM + flow-matching action head); MolmoBot-Pi0 (ฯโ architecture, for controlled comparison); MolmoBot-SPOC (lightweight, edge-deployable, RL-fine-tunable). Static (Franka FR3) and mobile (Rainbow Robotics RB-Y1) manipulation.
3. Results
- Real tabletop pick-and-place (DROID setup, 4 environments, 120 trials, zero-shot): MolmoBot 79.2% vs ฯโ.โ -DROID 39.2% and MolmoBot-Pi0 46.7%.
- Single-camera real kitchen (30 tasks): MolmoBot-Img 86.6% vs ฯโ.โ zero-shot 63.3%.
- Sim pick-and-place (200 eps): MolmoBot-Img 67.0% vs ฯโ.โ fine-tuned 46.0%, ฯโ.โ zero-shot 20.0%.
- Mobile (RB-Y1, sim): door-open specialist 77.7%; multitask 70.2% door, 44.8% pick. Real door opening was hard โ 2/9 successful openings, with underrepresented handle configs a key failure mode.
4. Why it matters
It is a strong existence proof that scale + diversity of synthetic data can beat real-data-fine-tuned baselines on real hardware for rigid/articulated pick-and-place, and it ships the data-generation engine openly.
Limitations (reviewer): confined to rigid-body + articulated tasks โ no contact-rich (insertion), deformables, or fluids/granular media; mobile-manipulation real-world transfer is still fragile (door opening low, sensitive to under-sampled geometries); real evals cover only two platforms.
5. Links
- arXiv 2603.16861
- Survey: CoRL 2026 ยท Related: World Models ยท Cross-Embodiment
โ Back to CoRL 2026 survey ยท Home