RSS 2026 RoboVista - Heungwoo/research GitHub Wiki

RoboVista: Evaluating Vision Language Models for Diverse Robot Applications

Venue: RSS 2026 (Sydney, Jul 13โ€“17) ยท Session: Datasets and Benchmarks ยท paper #95 Authors: Shuangyu Xie, Kaiyuan Chen, Ziyang Chen, Simeon Adebola, Yixuan Huang, Zehan Ma, Tianshuang Qiu, Wentao Yuan, Dhruv Shah, Pannag R. Sanketi, Ken Goldberg (UC Berkeley, Princeton, Google DeepMind) arXiv: 2607.04610 ยท program page

Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

RoboVista benchmark overview (Figure 1 of arXiv 2607.04610, ยฉ the authors)

Figure 1's three panels show what the benchmark stresses: varied embodiments (surgical dVRK, quadruped, agricultural, industrial, domestic robots), deformable objects and complex cluttered scenes, and long-horizon contexts (failure reasoning, task sequencing, plant-growth understanding).

Problem

VLMs are candidates for general-purpose robotic reasoning, but real robot applications are fundamentally modular โ€” task sequencing, action planning, recovery โ€” and these decision points are hard to capture with benchmarks built from teleoperated end-to-end datasets. Domains like agriculture, surgery, and industrial robotics also lack the massive trajectory corpora that prior VQA-generation efforts (Robo2VLM, RoboBrain) rely on.

Method

RQA is a modular evaluation framework that abstracts any functional block of a robot pipeline as a tuple (Embodiment, State space, Task-output space, Constraints) across four layers โ€” perception, high-level decision making, motion/action estimation, and failure recovery โ€” and maps it to robot-centric VQA instances (visual input, question, answer, 5 candidates, expert rationale). RoboVista instantiates this: 474 expert-annotated multiple-choice questions (318 perception, 156 planning) over 39 task types in six domains (agriculture, driving, domestic, industrial, surgical, and open datasets DROID / Open X-Embodiment / AgiBot), sourced from 18 peer-reviewed publications; all questions are written and multi-pass-verified by robotics-expert annotators (more than half with robotics PhDs). Case studies include Dex-Net 4.0 ambidextrous bin picking and dVRK surgical knot tying.

Results

Zero-shot: best overall is Gemini 2.5 Pro at 56.5%, with GPT-4o 49.6%, GPT-5 48.1%, and the best open model Qwen3-235B-A22B at 51.3% (random 20%, text-only 25.1%) โ€” industrial (as low as ~31โ€“48%) and agricultural domains are hardest. Chain-of-thought hurts perception questions (drops up to ~12%) while helping planning for several models (+8.3 for Qwen3-VL-32B); in-context learning consistently reduces accuracy (โˆ’2.8 to โˆ’6.5%) and raises calibration error up to +9.7, indicating hallucination. Failure analysis: most errors are visual (misidentification 30.2% for Qwen2.5-VL-7B, dropping to 20.3% at 235B scale, where spatial errors persist). Physical validation: RoboVista scores correlate strongly with bimanual gripper-alignment execution error (Pearson r = โˆ’0.78, Spearman ฯ = โˆ’0.93) and with progress in shared-autonomy dVRK knot tying.

Significance

A deliberately small, expert-curated, modular benchmark that probes the decision points VLMs would actually face inside robot pipelines โ€” and demonstrates the benchmark's predictive validity on physical robots. Directly relevant to Review-VLA-Evaluation.

โ† Back to RSS 2026 survey ยท RSS-2026-Papers ยท Home