NeurIPS 2025 Robo2VLM 1 - Heungwoo/research GitHub Wiki

Robo2VLM-1 — VQA Dataset from Large-Scale Robot Manipulation Trajectories

Venue: NeurIPS 2025 (Spotlight, Datasets & Benchmarks track) · Authors: Kaiyuan Chen, Shuangyu Xie, Zehan Ma, Ken Goldberg (UC Berkeley), Pannag R. Sanketi (Google DeepMind) · arXiv: 2505.15517 Category: Data & Benchmarks

Approach diagram

flowchart LR
  Data[Open X-Embodiment<br/>176,139 real robot trajectories<br/>3,396 tasks / 463 scenes] --> Extract[Extract non-visual ground-truth:<br/>EE pose · gripper aperture · force sensing]
  Extract --> Phase[Segment into manipulation phases]
  Phase --> Q[Auto-generate VQA questions]
  Q --> Types[11 question types across spatial /<br/>goal-conditioned / interaction reasoning]
  Types --> Dataset[Robo2VLM-1:<br/>684,710 VQA pairs]
Loading

Problem

VLM benchmarks like MMMU measure reasoning on static images. But a VLA's real test is reasoning about actions, forces, contacts, and goals — concepts that don't show up well in image-only QA. ManipBench was a first step (12,617 questions on manipulation reasoning); Robo2VLM-1 scales it up using robot proprioception as non-visual ground-truth.

Method

Take 176,139 real robot manipulation trajectories sourced from Open X-Embodiment (13 constituent datasets including DROID and Fractal), spanning 463 scenes / 3,396 tasks. Each trajectory is segmented into manipulation phases; from each phase the pipeline extracts non-visual / non-descriptive ground truth:

  • End-effector pose
  • Gripper aperture
  • Force sensing

Combined with scene/interaction understanding (3D properties of robot, goal, target object), these label auto-generated VQA queries spanning 11 question categories across three reasoning groups:

  • Spatial reasoning: Object State, Spatial Relationship, Scene Understanding, Multiple View
  • Goal-conditioned reasoning: Task State-success, Task State-goal, Action Understanding, Interaction Phase, Trajectory Understanding
  • Interaction reasoning: Task State-grasp, Robot State

Output: 684,710 VQA pairs grounded in robot proprioception rather than human annotation.

Results

  • Current open-source VLMs perform poorly on the benchmark: best zero-shot is Qwen 2.5-VL-72B at 37.76%, and Chain-of-Thought prompting tops out around 41.30% (Qwen 2.5-VL-32B) — far from saturated, confirming manipulation reasoning is hard for off-the-shelf VLMs. (Evaluation covers open models: LLaVA, Llama 3.2, Qwen 2.5-VL.)
  • Fine-tuning LLaVA-1.6 on 10k–50k Robo2VLM-1 items improves accuracy across nearly all categories, showing the dataset can both benchmark and train.
  • Enables principled VLM evaluation on manipulation reasoning, complementing the ICLR 2026 VLM4VLA finding that VLM scores don't predict manipulation success.

Significance

NeurIPS 2025 Spotlight. The first large-scale VQA dataset grounded in robot proprioception — moves VLM evaluation from "does it understand the scene?" to "does it understand what the robot can do?"

Seeds the ICLR 2026 dataset cluster (RoboCasa365, RoboArena ∞, EgoDex) and feeds directly into VLM4VLA's analysis.

Links

Related pages

← Back to NeurIPS-2025

⚠️ **GitHub.com Fallback** ⚠️