RoboBrain 2 - Heungwoo/research GitHub Wiki

RoboBrain 2.0 — Embodied Reasoning Foundation Model (BAAI)

Venue: BAAI technical report · arXiv: 2507.02029 · Date: July 2025 Category: Embodied reasoning / planning VLM (System-2) Trend tag: Embodied brain (reasoning, not action head)

Approach diagram

flowchart LR
  IMG[Multi-image + long video<br/>high-res frames] --> VLM
  LANG[Language instruction] --> VLM
  SG[Structured scene graph] --> VLM
  subgraph VLM[RoboBrain 2.0 VLM]
    VE[Vision encoder ~689M<br/>+ MLP projector] --> LLM[Decoder LLM 7B / 32B<br/>init from Qwen2.5-VL<br/>long chain-of-thought]
  end
  VLM --> AFF[Affordance points /<br/>spatial referring coords]
  VLM --> TRAJ[Trajectory forecasts]
  VLM --> PLAN[Long-horizon &<br/>multi-agent plans]
  AFF --> POL[Low-level action policy<br/>e.g. RoboOS-orchestrated executor]
  TRAJ --> POL
  PLAN --> POL
Loading

Problem

Embodied agents need a System-2 reasoning layer that unifies perception, spatial grounding, and planning before any motor command is issued. General-purpose VLMs handle open-ended Q&A but are weak at the physical sub-skills a robot brain requires: predicting affordances, resolving precise spatial references, forecasting trajectories, maintaining a scene graph over time, and decomposing long-horizon / multi-agent tasks. RoboBrain 2.0 targets this gap as a dedicated embodied vision-language foundation model — explicitly a reasoning/planning brain, not a low-level action policy that emits joint commands.

Method

  • Heterogeneous architecture. A lightweight vision encoder (~689M params) with dynamic-resolution processing + an MLP projector, feeding a decoder-only LLM (7B or 32B) initialized from Qwen2.5-VL. The LLM performs long chain-of-thought and emits structured plans, spatial relations, and both relative and absolute coordinates.
  • Inputs. Multi-image, long-video, and high-resolution visual streams plus language and structured scene-graph context.
  • Unified capability set. Spatial understanding — affordance prediction, spatial referring, trajectory forecasting (pointing/coordinate outputs); temporal decision-making — closed-loop interaction, multi-agent long-horizon planning, and scene-graph updating.
  • Released variants. The report headlines a 7B and a 32B model; the public BAAI checkpoint collection additionally ships a lighter 3B variant (BAAI/RoboBrain2.0-3B / -7B / -32B).

Results

All numbers below are from the technical report (32B / 7B).

Spatial reasoning: BLINK (overall) 83.63 / 83.95 · CV-Bench 83.92 / 85.75 · RoboSpatial 72.43 / 54.23 · RefSpatial-Bench 54.00 / 32.50 · Where2Place 73.59 / 63.59 · SAT 86.67 / 75.33 · ShareRobot-Affordance 35.28 / 28.05.

Temporal reasoning: Multi-Robot Planning 80.33 / 81.50 · EgoPlan2 57.23 / 33.23 · RoboBench Planning 68.33 / 72.16.

The report states the 32B variant achieves leading performance across spatial and temporal benchmarks, surpassing prior open-source and proprietary models.

Significance

RoboBrain 2.0 is the embodied-reasoning brain that the planner/reasoner pages in this wiki orbit. In the System 0/1/2 review taxonomy it sits squarely in System 2 — deliberative perception/planning that hands affordances, trajectories, and sub-goals to a faster low-level executor (its sibling RoboOS multi-robot coordination layer fills the orchestration role). It is a natural Planner backbone alongside Vlaser, which likewise builds an embodied-reasoning VLM from a general VLM, and it is the embodied-VLM baseline that RoboInter compares its Planner against (RoboInter reports large RoboRefIt / RoboVQA gains over the RoboBrain-2.0 3B/7B checkpoints). The contrast is instructive: RoboBrain emphasizes a broad, general-purpose embodied reasoning model, while data-product efforts like RoboInter target dense intermediate-representation supervision — two complementary routes to the same plan-then-execute stack.

Links

Related pages

← Back to Home

⚠️ **GitHub.com Fallback** ⚠️