NeurIPS 2025 Robo2VLM 1 - Heungwoo/research GitHub Wiki
Venue: NeurIPS 2025 (Spotlight, Datasets & Benchmarks track) · Authors: Kaiyuan Chen, Shuangyu Xie, Zehan Ma, Ken Goldberg (UC Berkeley), Pannag R. Sanketi (Google DeepMind) · arXiv: 2505.15517 Category: Data & Benchmarks
flowchart LR
Data[Open X-Embodiment<br/>176,139 real robot trajectories<br/>3,396 tasks / 463 scenes] --> Extract[Extract non-visual ground-truth:<br/>EE pose · gripper aperture · force sensing]
Extract --> Phase[Segment into manipulation phases]
Phase --> Q[Auto-generate VQA questions]
Q --> Types[11 question types across spatial /<br/>goal-conditioned / interaction reasoning]
Types --> Dataset[Robo2VLM-1:<br/>684,710 VQA pairs]
VLM benchmarks like MMMU measure reasoning on static images. But a VLA's real test is reasoning about actions, forces, contacts, and goals — concepts that don't show up well in image-only QA. ManipBench was a first step (12,617 questions on manipulation reasoning); Robo2VLM-1 scales it up using robot proprioception as non-visual ground-truth.
Take 176,139 real robot manipulation trajectories sourced from Open X-Embodiment (13 constituent datasets including DROID and Fractal), spanning 463 scenes / 3,396 tasks. Each trajectory is segmented into manipulation phases; from each phase the pipeline extracts non-visual / non-descriptive ground truth:
- End-effector pose
- Gripper aperture
- Force sensing
Combined with scene/interaction understanding (3D properties of robot, goal, target object), these label auto-generated VQA queries spanning 11 question categories across three reasoning groups:
- Spatial reasoning: Object State, Spatial Relationship, Scene Understanding, Multiple View
- Goal-conditioned reasoning: Task State-success, Task State-goal, Action Understanding, Interaction Phase, Trajectory Understanding
- Interaction reasoning: Task State-grasp, Robot State
Output: 684,710 VQA pairs grounded in robot proprioception rather than human annotation.
- Current open-source VLMs perform poorly on the benchmark: best zero-shot is Qwen 2.5-VL-72B at 37.76%, and Chain-of-Thought prompting tops out around 41.30% (Qwen 2.5-VL-32B) — far from saturated, confirming manipulation reasoning is hard for off-the-shelf VLMs. (Evaluation covers open models: LLaVA, Llama 3.2, Qwen 2.5-VL.)
- Fine-tuning LLaVA-1.6 on 10k–50k Robo2VLM-1 items improves accuracy across nearly all categories, showing the dataset can both benchmark and train.
- Enables principled VLM evaluation on manipulation reasoning, complementing the ICLR 2026 VLM4VLA finding that VLM scores don't predict manipulation success.
NeurIPS 2025 Spotlight. The first large-scale VQA dataset grounded in robot proprioception — moves VLM evaluation from "does it understand the scene?" to "does it understand what the robot can do?"
Seeds the ICLR 2026 dataset cluster (RoboCasa365, RoboArena ∞, EgoDex) and feeds directly into VLM4VLA's analysis.
- arXiv: https://arxiv.org/abs/2505.15517
- NeurIPS virtual page: https://neurips.cc/virtual/2025/loc/san-diego/poster/121678
- Project: https://berkeleyautomation.github.io/robo2vlm/
- ManipBench (CoRL 2025 ancestor)
- VLM4VLA (ICLR 2026 analysis paper)
- Review: VLM4VLA
← Back to NeurIPS-2025