CoRL 2026 MolmoAct2 - Heungwoo/research GitHub Wiki
CoRL 2026 — MolmoAct2: Action Reasoning Models for Real-World Deployment
Venue: CoRL 2026 (Austin, TX, Nov 9–12) · Allen Institute for AI (Ai2). Paper: arXiv 2605.02881. Representative of: fully-open action-reasoning VLA stack — a completely open (weights + data + tokenizer + code) reasoning policy that stays deployable across five embodiments while cutting reasoning latency. Companions: VLA Architectures · GR00T Series · CoRL 2026 survey.

1. Problem
Robot foundation models are hard to actually deploy: frontier systems (GPT-5, Gemini Robotics ER) are closed; open alternatives demand expensive hardware; reasoning-augmented policies add prohibitive per-step latency; and fine-tuned success rates stay below practical thresholds. The field lacks a fully open, efficient, deployable action-reasoning stack that works across multiple embodiments — which is what MolmoAct2 targets, releasing weights, data, tokenizer, and training code.
2. Method
- Embodied-reasoning VLM backbone (MolmoER / Molmo2-ER): a Molmo-family VLM specialized for spatial/embodied reasoning on a ~3.3M-sample corpus with a "specialize-then-rehearse" recipe (embodied specialization, then joint refinement balancing embodied and general VQA).
- Open action tokenizer (OpenFAST / MolmoAct2-FAST): open-weight, open-data tokenizer that maps ~1-second continuous trajectories to discrete tokens via a frequency-domain transform + BPE, trained on ~1M action sequences spanning five embodiments (YAM, SO-100/101, DROID Franka, and smaller sources).
- Adaptive-depth reasoning (MolmoThink): re-predicts depth tokens only for scene regions that change between timesteps (quantized depth grid gated by RGB-patch cosine similarity), preserving geometric grounding while cutting latency.
- Per-layer KV conditioning: a flow-matching continuous-action expert is grafted onto the discrete-token VLM by conditioning each expert layer on the corresponding VLM layer's keys/values via learned adapters — no backbone modification. Trained in stages: discrete-token pre-training, flow-expert post-training, then embodiment-specific fine-tuning (YAM, DROID, SO-100/101, LIBERO).
3. Results
- Embodied reasoning: the ER backbone reaches ~63.8% average across 13 embodied-reasoning benchmarks, above GPT-5 (~57.9%) and Gemini Robotics ER-1.5.
- Manipulation: on DROID/MolmoSpace tasks the MolmoAct2-DROID policy averages ~37.7% vs π₀.₅-DROID ~34.5%, with larger gains on pick and pick-and-place subtasks.
- Data release: ~720 hours of teleoperated bimanual (YAM) trajectories, described as the largest open bimanual dataset to date; also evaluated on LIBERO, real-world YAM household tasks, and zero-shot SO-100/101 deployment.
4. Why it matters
- Pushes "reasoning VLA" from closed demos toward a reproducible, end-to-end open release (backbone + tokenizer + data + code), lowering the barrier for labs without frontier-scale infrastructure.
- Shows adaptive-depth reasoning can keep the accuracy benefit of spatial grounding while removing most of its latency cost — a concrete answer to the reasoning-vs-deployability tension.
- The five-embodiment tokenizer and large open bimanual dataset are reusable assets independent of the specific policy.
Limitations (reviewer): exact numbers vary between the abstract and HTML draft and were read from an unrefereed preprint, so treat specific percentages as provisional; real-world evaluations are on a modest task suite (household YAM tasks) rather than at large scale; success rates, while state-of-the-art here, are still short of reliability thresholds for unsupervised deployment.
5. Links
- arXiv 2605.02881
- Survey: CoRL 2026 · Related: VLA Architectures · GR00T Series
← Back to CoRL 2026 survey · Home