CoRL 2026 MolmoAct2 - Heungwoo/research GitHub Wiki

CoRL 2026 — MolmoAct2: Action Reasoning Models for Real-World Deployment

Venue: CoRL 2026 (Austin, TX, Nov 9–12) · Allen Institute for AI (Ai2). Paper: arXiv 2605.02881. Representative of: fully-open action-reasoning VLA stack — a completely open (weights + data + tokenizer + code) reasoning policy that stays deployable across five embodiments while cutting reasoning latency. Companions: VLA Architectures · GR00T Series · CoRL 2026 survey.

MolmoAct2 post-training architecture with per-layer KV conditioning (figure from Fang et al., arXiv 2605.02881, © the authors)

1. Problem

Robot foundation models are hard to actually deploy: frontier systems (GPT-5, Gemini Robotics ER) are closed; open alternatives demand expensive hardware; reasoning-augmented policies add prohibitive per-step latency; and fine-tuned success rates stay below practical thresholds. The field lacks a fully open, efficient, deployable action-reasoning stack that works across multiple embodiments — which is what MolmoAct2 targets, releasing weights, data, tokenizer, and training code.

2. Method

  • Embodied-reasoning VLM backbone (MolmoER / Molmo2-ER): a Molmo-family VLM specialized for spatial/embodied reasoning on a ~3.3M-sample corpus with a "specialize-then-rehearse" recipe (embodied specialization, then joint refinement balancing embodied and general VQA).
  • Open action tokenizer (OpenFAST / MolmoAct2-FAST): open-weight, open-data tokenizer that maps ~1-second continuous trajectories to discrete tokens via a frequency-domain transform + BPE, trained on ~1M action sequences spanning five embodiments (YAM, SO-100/101, DROID Franka, and smaller sources).
  • Adaptive-depth reasoning (MolmoThink): re-predicts depth tokens only for scene regions that change between timesteps (quantized depth grid gated by RGB-patch cosine similarity), preserving geometric grounding while cutting latency.
  • Per-layer KV conditioning: a flow-matching continuous-action expert is grafted onto the discrete-token VLM by conditioning each expert layer on the corresponding VLM layer's keys/values via learned adapters — no backbone modification. Trained in stages: discrete-token pre-training, flow-expert post-training, then embodiment-specific fine-tuning (YAM, DROID, SO-100/101, LIBERO).

3. Results

  • Embodied reasoning: the ER backbone reaches ~63.8% average across 13 embodied-reasoning benchmarks, above GPT-5 (~57.9%) and Gemini Robotics ER-1.5.
  • Manipulation: on DROID/MolmoSpace tasks the MolmoAct2-DROID policy averages ~37.7% vs π₀.₅-DROID ~34.5%, with larger gains on pick and pick-and-place subtasks.
  • Data release: ~720 hours of teleoperated bimanual (YAM) trajectories, described as the largest open bimanual dataset to date; also evaluated on LIBERO, real-world YAM household tasks, and zero-shot SO-100/101 deployment.

4. Why it matters

  • Pushes "reasoning VLA" from closed demos toward a reproducible, end-to-end open release (backbone + tokenizer + data + code), lowering the barrier for labs without frontier-scale infrastructure.
  • Shows adaptive-depth reasoning can keep the accuracy benefit of spatial grounding while removing most of its latency cost — a concrete answer to the reasoning-vs-deployability tension.
  • The five-embodiment tokenizer and large open bimanual dataset are reusable assets independent of the specific policy.

Limitations (reviewer): exact numbers vary between the abstract and HTML draft and were read from an unrefereed preprint, so treat specific percentages as provisional; real-world evaluations are on a modest task suite (household YAM tasks) rather than at large scale; success rates, while state-of-the-art here, are still short of reliability thresholds for unsupervised deployment.

5. Links

← Back to CoRL 2026 survey · Home