CoRL 2026 MolmoBOT - Heungwoo/research GitHub Wiki

CoRL 2026 โ€” MolmoBOT: Large-Scale Simulation Enables Zero-Shot Manipulation

Venue: CoRL 2026 (Austin, TX, Nov 9โ€“12) ยท Allen Institute for AI (Ai2). Paper: arXiv 2603.16861. Representative of: the sim-only zero-shot-transfer thesis โ€” enough simulation diversity โ†’ real-world manipulation with no real training data. Companions: World Models ยท Cross-Embodiment ยท CoRL 2026 survey.

MolmoBOT teaser: procedurally generated sim scenes, the MolmoBot policy, and zero-shot real-robot rollouts on Franka and RB-Y1 (figure from the authors, arXiv 2603.16861, ยฉ the authors)

1. Problem

Real robot data is expensive and task-specific fine-tuning is the norm. MolmoBOT asks whether simulation alone โ€” if made large and diverse enough โ€” can produce a manipulation policy that transfers zero-shot to physical robots, with no real training data and no real fine-tuning.

2. Method

  • MolmoBot-Engine โ€” an open-source procedural pipeline that generates robots, tasks, and scenes in MolmoSpaces (200k+ pre-built houses) on top of the MuJoCo simulator. It randomizes layouts, 6-DoF object poses, lighting, textures/materials, friction/mass/joint damping, camera extrinsics, and injects action noise. Assets come from iTHOR + Objaverse (filtered for graspable, watertight colliders); an expert planner does phase-based grasp/place trajectories (RB-Y1 uses CuRobo).
  • MolmoBot-Data โ€” ~1.8M expert trajectories (~300M frames, ~5,817 robot-hours) across 94,300 unique environments, 11,400+ pickup assets and 9,400+ receptacles. Generated at ~1,024 episodes/GPU-hour (100ร— A100-80GB, ~4,500 GPU-hours total).
  • Policies โ€” MolmoBot (Molmo2 VLM + flow-matching action head); MolmoBot-Pi0 (ฯ€โ‚€ architecture, for controlled comparison); MolmoBot-SPOC (lightweight, edge-deployable, RL-fine-tunable). Static (Franka FR3) and mobile (Rainbow Robotics RB-Y1) manipulation.

3. Results

  • Real tabletop pick-and-place (DROID setup, 4 environments, 120 trials, zero-shot): MolmoBot 79.2% vs ฯ€โ‚€.โ‚…-DROID 39.2% and MolmoBot-Pi0 46.7%.
  • Single-camera real kitchen (30 tasks): MolmoBot-Img 86.6% vs ฯ€โ‚€.โ‚… zero-shot 63.3%.
  • Sim pick-and-place (200 eps): MolmoBot-Img 67.0% vs ฯ€โ‚€.โ‚… fine-tuned 46.0%, ฯ€โ‚€.โ‚… zero-shot 20.0%.
  • Mobile (RB-Y1, sim): door-open specialist 77.7%; multitask 70.2% door, 44.8% pick. Real door opening was hard โ€” 2/9 successful openings, with underrepresented handle configs a key failure mode.

4. Why it matters

It is a strong existence proof that scale + diversity of synthetic data can beat real-data-fine-tuned baselines on real hardware for rigid/articulated pick-and-place, and it ships the data-generation engine openly.

Limitations (reviewer): confined to rigid-body + articulated tasks โ€” no contact-rich (insertion), deformables, or fluids/granular media; mobile-manipulation real-world transfer is still fragile (door opening low, sensitive to under-sampled geometries); real evals cover only two platforms.

5. Links

โ† Back to CoRL 2026 survey ยท Home