CoRL 2025 ManipBench - Heungwoo/research GitHub Wiki
Venue: CoRL 2025 ยท Authors: Enyu Zhao, Vedant Raval, Hejia Zhang, Jiageng Mao, Zeyu Shangguan, Stefanos Nikolaidis, Yue Wang, Daniel Seita (University of Southern California) ยท arXiv: 2505.09698 Category: Data & Benchmarks Trend tag: Manipulation-specific VLM benchmark predicts real-world skill
flowchart LR
IMG[Manipulation scene image] --> Q[Multi-choice question<br/>pick-place / articulated /<br/>deformable / dynamic / tool use]
Q --> VLMs[33 VLMs ร 10 families]
VLMs --> SCORES[Per-model accuracy<br/>12,617 MCQs]
SCORES --> ANALYSIS[Best VLM Gemini-2.5-pro < humans<br/>scores correlate w/ real-world<br/>manipulation rโ0.89]
Previous VLM benchmarks (MMMU, MMBench, etc.) test high-level visual reasoning. They don't test the low-level manipulation understanding a VLA actually needs: contact outcomes, deformable behavior, object-object interaction, stability, tool affordances. As a result, a VLM that scores well on generic benchmarks isn't necessarily a good VLA backbone.
ManipBench is 12,617 multiple-choice questions built specifically around low-level manipulation scenarios, spanning five task categories: pick-and-place, articulated-object manipulation, deformable-object manipulation, dynamic manipulation, and tool use. Questions are drawn from multiple data sources (including Bridge and DROID). 33 representative VLMs across 10 model families were evaluated, including both closed-source (GPT-4, Gemini) and open-source (InternVL, Qwen-VL) models and several model-size variants.
VLM accuracy varies substantially across the five task categories, and a large gap to human understanding remains. The best model, Gemini-2.5-pro, scored highest on most question types (e.g. 0.916 on Bridge Type-1) but still trailed humans (0.880 on Bridge Type-1, 0.990 on DROID-articulated Type-1). The headline result is a strong positive correlation between ManipBench scores and real-world manipulation performance (Pearson's r = 0.889, p = 0.003; Spearman's ฯ = 0.850), i.e. the manipulation-specific benchmark is predictive of downstream robot action-selection.
ManipBench argues that general VLM benchmarks (MMMU, MMBench, etc.) test high-level reasoning and do not capture the low-level manipulation understanding a VLA needs โ so a manipulation-targeted benchmark is required. Crucially, ManipBench shows its own manipulation-specific score does track real-world success (r โ 0.89), establishing it as a useful proxy for VLM-as-backbone selection. This complements ICLR 2026's VLM4VLA analysis, which separately probes whether generic VLM benchmark scores predict downstream manipulation success.
โ Back to CoRL-2025