ICLR 2026 Seeing Across Views - Heungwoo/research GitHub Wiki

MV-RoboBench โ€” Multi-View Spatial Reasoning Benchmark

Venue: ICLR 2026 Category: Benchmark โ€” Spatial reasoning / Embodied VLM evaluation Trend tag: Spatial ยท 3D for VLA ยท Multi-view evaluation Affiliations: Tsinghua, PKU, Fudan, Microsoft Research Asia, HKUST, Zhejiang Univ. (Feng, Kang, Wang, Du, Yan, Shi, Yuan, Liang, Deng, Li, Yang, An, Zheng, Wang, Chen, Xu, Liang, Yang, Guo)

Approach diagram

flowchart LR
  Src[AgiWorld + BridgeV2<br/>980 episodes, multi-camera] --> Filter[Rule-based + GPT-4.1 filter<br/>candidate triage]
  Filter --> Anno[Human annotation<br/>5-choice MCQ, 8 task templates]
  Anno --> QA[1,708 multi-view QA items]
  QA --> Eval[Zero-shot eval<br/>blind / proprietary / reasoning / open-source / MoE / human]
  Eval --> Cot[CoT-style augmentations<br/>text descriptions, VGGT views, MoGe-2 depth]
  Eval --> Cor[Correlation analysis<br/>spatial vs robotic, single-view vs multi-view]
Loading

Problem

VLM spatial-reasoning benchmarks are mostly single-view (EmbSpatial-Bench, RoboSpatial, SpatialVLM, VSI-Bench, OmniSpatial, ShareRobot). Real robots use multi-camera setups. The two existing multi-view benchmarks โ€” All-Angles Bench and Ego3D-Bench โ€” focus on photographic alignment or egocentric navigation, not manipulation. ERQA and MMSI-Bench include partial multi-view but limited diversity and basic perception only. The paper's claim: strong single-view scores do not predict multi-view robotic competence, so a dedicated benchmark is needed.

Detailed Method

Construction pipeline (3 stages)

Stage 1 โ€” Data Collection. Source from AgiWorld (single-arm + dual-arm) and BridgeV2 (single-arm). Rule-based filtering for temporal separation, scene diversity, visual clarity. GPT-4.1 triages candidate image pairs against the 8 task definitions. Final human verification.

Stage 2 โ€” QA Generation. Trained annotators write 5-choice MCQs from task-specific templates. GPT-4.1 is never used to generate QA content โ€” only auxiliary triage. Distractors must be plausible but clearly distinguishable. After automatic sampling and manual inspection, more than 3,000 high-quality image pairs were retained as the foundation pool (Appendix F).

Stage 3 โ€” Human-in-the-loop Quality Review. Iterative review checks alignment, sentence structure, image-pair quality, distractor balance, and answer-distribution randomization. Final pool: 1,708 QA items from 980 episodes.

8 subtasks in 2 categories

Spatial understanding (4 subtasks):

  1. Cross-View Matching โ€” identify the same object across views (e.g., red bbox in right-gripper view โ†’ which colored bbox in left-gripper / head view).
  2. Distance Judgement โ€” relative distances between objects from multi-view input.
  3. Viewpoint Identification โ€” reason about viewpoint transformation (which gripper view matches a given head-camera image).
  4. 3D Spatial Consistency โ€” assign coordinate triplets to objects given a head-view-anchored 3D cube.

Robotic execution (4 subtasks): 5. Action Planning โ€” choose a multi-step manipulation sequence in a head-anchored frame. 6. Step Execution โ€” verify whether a specific next-step movement is correct. 7. Trajectory Selection โ€” pick the colored line most likely to complete a task. 8. Affordance Recognition โ€” pick the grasp candidate most likely to succeed.

CoT-inspired enhancements (Section 2.3)

  • Textual CoT (w cot / w text): GPT-4.1-generated scene descriptions appended to input.
  • Visual CoT (w vggt): synthesized novel views via VGGT (Wang et al., 2025a).
  • Structural CoT (w depth): depth priors via MoGe-2 (Wang et al., 2025b).

Comprehensive Results

Main results โ€” full per-model table (Table 2, % accuracy)

Method Cross-View Match Distance Judge Viewpoint ID 3D Spatial Consist. Action Plan Step Exec Trajectory Sel Affordance Rec Avg
Blind
Random Choice โ€“ 17.80 19.40 20.00 19.07 19.41 21.54 20.65 19.71
GPT-3.5-turbo โ€“ 15.50 22.39 20.31 12.25 21.57 18.38 23.00 18.52
GPT-4-turbo โ€“ 19.00 13.43 19.92 7.84 41.67 31.20 20.00 22.91
Proprietary (non-reasoning)
GPT-4o-mini 24.00 22.89 23.44 11.76 24.51 28.21 20.50 23.44 22.52
GPT-4o 24.50 37.31 19.92 6.37 33.33 33.76 33.00 20.10 27.59
GPT-4.1-nano 17.50 25.37 18.75 14.71 22.55 22.22 20.00 17.22 20.85
GPT-4.1-mini 28.50 33.83 25.00 7.84 26.47 21.79 32.00 18.18 23.98
GPT-4.1 26.00 43.28 32.03 6.37 29.90 31.62 41.50 28.23 30.90
Claude-3.5 17.50 27.86 20.31 8.82 34.80 20.09 33.00 27.27 23.71
Claude-3.7 18.00 35.32 20.31 6.86 36.76 29.06 34.50 22.97 25.47
Gemini-2.0-flash 28.00 32.84 21.48 7.35 32.84 29.91 52.50 20.57 28.94
Gemini-2.5-flash 26.50 37.31 27.34 6.37 34.80 30.34 42.00 19.14 27.23
Proprietary (reasoning)
o4-mini 21.50 48.26 26.17 65.69 74.51 63.25 44.00 25.36 46.47
GPT-5-chat 30.00 42.79 31.64 4.90 36.76 40.17 38.00 27.75 31.63
GPT-5-nano 21.50 33.33 17.58 56.86 39.71 35.47 31.00 26.32 32.75
GPT-5-mini 22.00 49.25 25.78 72.55 66.18 48.72 47.00 27.75 38.28
GPT-5 29.00 55.22 44.14 82.35 79.41 68.38 54.50 39.23 56.41
Claude-3.7-think 24.40 35.04 36.00 52.45 21.50 37.81 21.08 23.05 31.67
Gemini-2.5-pro 39.50 56.22 38.28 49.02 65.20 50.85 65.50 31.58 49.52
Open-source dense
Gemma-3-4b 21.00 22.89 21.09 11.76 17.65 16.67 25.50 22.01 19.79
Gemma-3-12b 18.00 26.37 20.31 9.80 22.55 20.94 25.50 20.57 20.49
Gemma-3-27b 21.50 23.88 20.31 9.31 20.10 23.08 29.00 17.22 20.55
InternVL3-2b 16.50 15.42 20.70 20.59 17.16 20.94 21.00 19.14 18.93
InternVL3-8b 19.00 21.39 26.17 12.75 26.47 21.37 20.50 20.10 20.97
InternVL3-14b 19.50 22.39 24.61 10.78 23.53 23.50 24.00 23.44 21.47
InternVL3-38b 24.50 25.87 23.44 6.86 27.94 25.21 27.50 21.05 22.80
InternVL3-78b 19.00 28.86 23.83 11.76 29.90 29.06 26.50 21.05 23.25
Qwen2.5-vl-3b 17.50 21.89 22.66 17.65 17.16 17.95 22.00 25.84 20.37
Qwen2.5-vl-7b 20.50 20.40 20.70 8.82 22.55 26.07 24.50 22.49 20.84
Qwen2.5-vl-32b 20.50 25.87 25.39 10.78 24.51 19.66 30.50 22.49 22.48
Qwen2.5-vl-72b 20.50 34.83 27.34 4.90 28.43 27.35 29.00 24.88 24.29
Open-source MoE
Llama-4-Scout 20.50 22.39 23.83 7.35 25.49 28.21 23.00 18.18 22.12
Llama-4-Maverick 14.00 42.79 17.58 5.88 37.75 37.18 36.00 20.10 26.11
Human 95.02 94.03 92.19 93.66 86.34 89.74 87.56 89.05 91.04

Headline observations

  • GPT-5 (56.41) leads overall but is 34.6 points below human (91.04).
  • Gemini-2.5-pro (49.52) and o4-mini (46.47) are second and third โ€” all reasoning-optimized models.
  • Best non-reasoning proprietary: GPT-4.1 (30.90). Best open-source dense: Qwen2.5-vl-72B (24.29). Best open-source MoE: Llama-4-Maverick (26.11).
  • 3D Spatial Consistency is the most discriminating subtask: non-reasoning models cluster near random (~6โ€“14%) while reasoning models jump to 49โ€“82%. GPT-5 hits 82.35 here.
  • All models trail badly on Affordance Recognition (best non-human is GPT-5 at 39.23 vs human 89.05).
  • Single-view ablation (Appendix G, Table 7): restricting to the most informative third-person view (AgiWorld head camera / BridgeV2 view1) on the four degradable subtasks drops GPT-5's Distance Judgement by 18.90 points (multi-view 36.32 here, โˆ† +18.90), and GPT-4.1 by +10.44 โ€” confirming the benchmark genuinely tests multi-view fusion. Gains correlate with model capability: smaller models (Qwen2.5-vl-7B/32B) show negligible or negative โˆ†.

CoT-inspired enhancements (Table 3, ฮ” vs each model's baseline)

Model Variant Spatial ฮ”s Robotic ฮ”r
Qwen2.5-vl-7b w cot +0.58 โˆ’1.30
w text โˆ’0.70 +0.82
w vggt โˆ’1.40 โˆ’0.24
w depth +1.04 โˆ’0.48
Gemma-3-12B w cot +0.93 +2.96
w text โˆ’0.94 โˆ’0.47
w vggt โˆ’1.47 +0.11
w depth โˆ’0.18 +0.19
GPT-4.1 w cot โˆ’1.21 โˆ’0.25
w text +1.73 +1.81
w vggt โˆ’1.85 โˆ’1.58
w depth +3.25 +2.71

Synthetic novel views (w vggt) tend to hurt performance โ€” the authors note current view-synthesis methods perform poorly under narrow baselines, cluttered tabletops, and gripper-centric viewpoints. Depth priors help only when the backbone has capacity to exploit them (GPT-4.1 yes, Qwen2.5-vl-7B partial). Textual CoT shines for mid-capacity open-source models like Gemma-3-12B but barely moves saturated proprietary models.

Correlation analyses (Section 4)

  • Internal axis (Fig. 5). Spatial vs robotic accuracy: positive correlation, but only proprietary and reasoning-optimized models occupy the upward region. Most open-source models cluster near random on both axes.
  • External axis (Fig. 6). OmniSpatial (single-view) accuracy vs MV-RoboBench accuracy: strong single-view performance does not transfer reliably. Many models that score well on OmniSpatial remain near random on multi-view tasks.

Limitations / Discussion (as stated by authors)

The paper does not enumerate "limitations" in a single section but the discussion (Section 6) and appendices flag:

  1. Benchmark scale is modest โ€” 1.7k items vs 8.6k for Ego3D-Bench or 5k for VSI-Bench โ€” chosen for manual quality control over template generation.
  2. Source-domain bias โ€” derived only from AgiWorld + BridgeV2 (tabletop manipulation). Locomotion, mobile manipulation, and humanoid scenarios are out of scope.
  3. Multiple-choice format simplifies grading but cannot probe open-ended generation quality. Future work targets free-form spatial reasoning evaluation.
  4. CoT augmentations are off-the-shelf โ€” no fine-tuning of view-synthesis or depth models for robot scenes; current novel-view-synthesis baselines genuinely fail on narrow-baseline gripper views.
  5. Looking forward (Sec. 6). Authors call for (i) architectures with explicit geometric priors and cross-view consistency, (ii) training pipelines aligning perception with action grounding, and (iii) larger-scale multi-camera datasets reflecting real-world manipulation complexity.

Significance & Positioning

MV-RoboBench is the first benchmark coupling multi-view perception with robotic manipulation reasoning. It serves as the diagnostic counterpart to method papers like:

  • Spatial Forcing โ€” claims to inject spatial competence into VLAs via training.
  • SP-VLA โ€” spatially guided VLA inference.
  • FALCON โ€” converts spatial reasoning to action chunks.

For each of these, MV-RoboBench provides a way to ask: did the multi-view spatial competence actually transfer, rather than relying on single-image proxies like SpatialVLM or RoboSpatial. It complements (rather than replaces) ERQA and MMSI-Bench by being the first to put manipulation-style multi-view reasoning front and center.

The benchmark's main empirical result โ€” that reasoning-optimized models (o4-mini, GPT-5, Gemini-2.5-pro) win by 15โ€“30 points over non-reasoning peers and that 3D-Spatial-Consistency separates these tiers most cleanly โ€” is a strong signal that multi-view robotic reasoning rewards explicit chain-of-thought, not raw perception scaling. This validates the bet behind systems like RoboBrain, RoboRefer, and Embodied-R1 that pair VLAs with explicit reasoning passes.

Links

Related pages

โ† Back to ICLR-2026

โš ๏ธ **GitHub.com Fallback** โš ๏ธ