RSS 2026 MolmoSpaces - Heungwoo/research GitHub Wiki
MolmoSpaces: Large-Scale Open Ecosystem for Robot Manipulation and Navigation
Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Datasets and Benchmarks · paper #91 Authors: Yejin Kim, Wilbert Pumacay, Omar Rayyan, Max Argus, Winson Han, Eli Vanderbilt, Jordi Salvador, Abhay Deshpande, Rose Hendrix, Snehal Jauhri, Shuo Liu, Nur Muhammad Mahi Shafiullah, Maya Guru, Ainaz Eftekhar, Karen Farley, Donovan Clay, Jiafei Duan, Arjun Guru, Piper Wolters, et al. (Allen Institute for AI and collaborators) arXiv: 2602.11337 · program page
Summary compiled from the arXiv paper (v2, titled "MolmoSpaces: A Large-Scale Open Ecosystem for Robot Navigation and Manipulation"); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Figure 1: the ecosystem at a glance — 230K interior environments, 130k objects with asset files/text/scale-mass metadata, 42M grasps, and multiple robot embodiments, loadable into Isaac, ManiSkill, and MuJoCo with high-fidelity physics; top-right scatter shows the strong correlation between MolmoSpaces-Bench success and real-world success.
Problem
Evaluating generalist robot policies requires coverage of the long tail of scenes, objects, and instructions that physical evaluation cannot provide; existing benchmarks are near saturation, focus on short-horizon skills in single scenes, and many sim-to-real evaluation pipelines are proprietary or closed.
Method
MolmoSpaces (Allen Institute for AI and collaborators; fully open-source) comprises four parts: MolmoSpaces-Scenes — over 230k indoor environments in five datasets (120 hand-crafted single-room MSCrafted scenes, 110k procedural MSProc houses, 110k MSProcObja scenes adding Objaverse objects, 110k MSMultiType diverse layouts, and MSTwin, a digital twin of the authors' real kitchen), physics-tuned for MuJoCo, IsaacSim, and ManiSkill; MolmoSpaces-Objects — 130k+ rigid and articulated models with semantic/physical metadata; MolmoSpaces-Grasp — 42M+ annotated 6-DoF grasps over 48k interactive objects; and MolmoSpaces-Bench — eight base tasks (navigate-to, pick, pick-and-place, pick-and-place-next-to, pick-and-place-color, open, close, open-door) with verified-solvable trials, success conditions, and dense rewards. Evaluations use a Franka FR3 in DROID configuration for rigid-body manipulation and a Rainbow RB-Y1 for navigation.
Results
Zero-shot evaluation of open-source policies (π0, π0-FAST, π0.5 DROID joint-position variants, CAP contact-action policies, and navigation baselines like RING/DualVLN) shows newer policies outperform older ones (e.g., Close task up to 84% for the best policy; Pick tops out around 34% among π/CAP variants in Fig. 10). Benchmark pick results correlate with 752 real-world RoboArena pick episodes at Pearson R = 0.96 and Spearman ρ = 0.98 (R² ≈ 0.92 for object picking). Distributional analyses expose failure modes: with DROID-frequent prompt phrasing π0 comes within 1% of π0.5 versus a 14% gap otherwise; perturbing initial joint positions degrades π0.5 while lighting changes barely matter; occluding the wrist camera drops π0.5 to 2% success versus 20% for third-person-camera occlusion.
Significance
The largest open, simulator-agnostic evaluation substrate for generalist policies to date, and its diagnostic findings (prompt sensitivity, wrist-camera reliance) directly inform VLA evaluation methodology. Related threads: Review-VLA-Evaluation.
← Back to RSS 2026 survey · RSS-2026-Papers · Home