ICLR 2026 Seeing Across Views - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Category: Benchmark โ Spatial reasoning / Embodied VLM evaluation Trend tag: Spatial ยท 3D for VLA ยท Multi-view evaluation Affiliations: Tsinghua, PKU, Fudan, Microsoft Research Asia, HKUST, Zhejiang Univ. (Feng, Kang, Wang, Du, Yan, Shi, Yuan, Liang, Deng, Li, Yang, An, Zheng, Wang, Chen, Xu, Liang, Yang, Guo)
flowchart LR
Src[AgiWorld + BridgeV2<br/>980 episodes, multi-camera] --> Filter[Rule-based + GPT-4.1 filter<br/>candidate triage]
Filter --> Anno[Human annotation<br/>5-choice MCQ, 8 task templates]
Anno --> QA[1,708 multi-view QA items]
QA --> Eval[Zero-shot eval<br/>blind / proprietary / reasoning / open-source / MoE / human]
Eval --> Cot[CoT-style augmentations<br/>text descriptions, VGGT views, MoGe-2 depth]
Eval --> Cor[Correlation analysis<br/>spatial vs robotic, single-view vs multi-view]
VLM spatial-reasoning benchmarks are mostly single-view (EmbSpatial-Bench, RoboSpatial, SpatialVLM, VSI-Bench, OmniSpatial, ShareRobot). Real robots use multi-camera setups. The two existing multi-view benchmarks โ All-Angles Bench and Ego3D-Bench โ focus on photographic alignment or egocentric navigation, not manipulation. ERQA and MMSI-Bench include partial multi-view but limited diversity and basic perception only. The paper's claim: strong single-view scores do not predict multi-view robotic competence, so a dedicated benchmark is needed.
Stage 1 โ Data Collection. Source from AgiWorld (single-arm + dual-arm) and BridgeV2 (single-arm). Rule-based filtering for temporal separation, scene diversity, visual clarity. GPT-4.1 triages candidate image pairs against the 8 task definitions. Final human verification.
Stage 2 โ QA Generation. Trained annotators write 5-choice MCQs from task-specific templates. GPT-4.1 is never used to generate QA content โ only auxiliary triage. Distractors must be plausible but clearly distinguishable. After automatic sampling and manual inspection, more than 3,000 high-quality image pairs were retained as the foundation pool (Appendix F).
Stage 3 โ Human-in-the-loop Quality Review. Iterative review checks alignment, sentence structure, image-pair quality, distractor balance, and answer-distribution randomization. Final pool: 1,708 QA items from 980 episodes.
Spatial understanding (4 subtasks):
- Cross-View Matching โ identify the same object across views (e.g., red bbox in right-gripper view โ which colored bbox in left-gripper / head view).
- Distance Judgement โ relative distances between objects from multi-view input.
- Viewpoint Identification โ reason about viewpoint transformation (which gripper view matches a given head-camera image).
- 3D Spatial Consistency โ assign coordinate triplets to objects given a head-view-anchored 3D cube.
Robotic execution (4 subtasks): 5. Action Planning โ choose a multi-step manipulation sequence in a head-anchored frame. 6. Step Execution โ verify whether a specific next-step movement is correct. 7. Trajectory Selection โ pick the colored line most likely to complete a task. 8. Affordance Recognition โ pick the grasp candidate most likely to succeed.
-
Textual CoT (
w cot/w text): GPT-4.1-generated scene descriptions appended to input. -
Visual CoT (
w vggt): synthesized novel views via VGGT (Wang et al., 2025a). -
Structural CoT (
w depth): depth priors via MoGe-2 (Wang et al., 2025b).
| Method | Cross-View Match | Distance Judge | Viewpoint ID | 3D Spatial Consist. | Action Plan | Step Exec | Trajectory Sel | Affordance Rec | Avg |
|---|---|---|---|---|---|---|---|---|---|
| Blind | |||||||||
| Random Choice | โ | 17.80 | 19.40 | 20.00 | 19.07 | 19.41 | 21.54 | 20.65 | 19.71 |
| GPT-3.5-turbo | โ | 15.50 | 22.39 | 20.31 | 12.25 | 21.57 | 18.38 | 23.00 | 18.52 |
| GPT-4-turbo | โ | 19.00 | 13.43 | 19.92 | 7.84 | 41.67 | 31.20 | 20.00 | 22.91 |
| Proprietary (non-reasoning) | |||||||||
| GPT-4o-mini | 24.00 | 22.89 | 23.44 | 11.76 | 24.51 | 28.21 | 20.50 | 23.44 | 22.52 |
| GPT-4o | 24.50 | 37.31 | 19.92 | 6.37 | 33.33 | 33.76 | 33.00 | 20.10 | 27.59 |
| GPT-4.1-nano | 17.50 | 25.37 | 18.75 | 14.71 | 22.55 | 22.22 | 20.00 | 17.22 | 20.85 |
| GPT-4.1-mini | 28.50 | 33.83 | 25.00 | 7.84 | 26.47 | 21.79 | 32.00 | 18.18 | 23.98 |
| GPT-4.1 | 26.00 | 43.28 | 32.03 | 6.37 | 29.90 | 31.62 | 41.50 | 28.23 | 30.90 |
| Claude-3.5 | 17.50 | 27.86 | 20.31 | 8.82 | 34.80 | 20.09 | 33.00 | 27.27 | 23.71 |
| Claude-3.7 | 18.00 | 35.32 | 20.31 | 6.86 | 36.76 | 29.06 | 34.50 | 22.97 | 25.47 |
| Gemini-2.0-flash | 28.00 | 32.84 | 21.48 | 7.35 | 32.84 | 29.91 | 52.50 | 20.57 | 28.94 |
| Gemini-2.5-flash | 26.50 | 37.31 | 27.34 | 6.37 | 34.80 | 30.34 | 42.00 | 19.14 | 27.23 |
| Proprietary (reasoning) | |||||||||
| o4-mini | 21.50 | 48.26 | 26.17 | 65.69 | 74.51 | 63.25 | 44.00 | 25.36 | 46.47 |
| GPT-5-chat | 30.00 | 42.79 | 31.64 | 4.90 | 36.76 | 40.17 | 38.00 | 27.75 | 31.63 |
| GPT-5-nano | 21.50 | 33.33 | 17.58 | 56.86 | 39.71 | 35.47 | 31.00 | 26.32 | 32.75 |
| GPT-5-mini | 22.00 | 49.25 | 25.78 | 72.55 | 66.18 | 48.72 | 47.00 | 27.75 | 38.28 |
| GPT-5 | 29.00 | 55.22 | 44.14 | 82.35 | 79.41 | 68.38 | 54.50 | 39.23 | 56.41 |
| Claude-3.7-think | 24.40 | 35.04 | 36.00 | 52.45 | 21.50 | 37.81 | 21.08 | 23.05 | 31.67 |
| Gemini-2.5-pro | 39.50 | 56.22 | 38.28 | 49.02 | 65.20 | 50.85 | 65.50 | 31.58 | 49.52 |
| Open-source dense | |||||||||
| Gemma-3-4b | 21.00 | 22.89 | 21.09 | 11.76 | 17.65 | 16.67 | 25.50 | 22.01 | 19.79 |
| Gemma-3-12b | 18.00 | 26.37 | 20.31 | 9.80 | 22.55 | 20.94 | 25.50 | 20.57 | 20.49 |
| Gemma-3-27b | 21.50 | 23.88 | 20.31 | 9.31 | 20.10 | 23.08 | 29.00 | 17.22 | 20.55 |
| InternVL3-2b | 16.50 | 15.42 | 20.70 | 20.59 | 17.16 | 20.94 | 21.00 | 19.14 | 18.93 |
| InternVL3-8b | 19.00 | 21.39 | 26.17 | 12.75 | 26.47 | 21.37 | 20.50 | 20.10 | 20.97 |
| InternVL3-14b | 19.50 | 22.39 | 24.61 | 10.78 | 23.53 | 23.50 | 24.00 | 23.44 | 21.47 |
| InternVL3-38b | 24.50 | 25.87 | 23.44 | 6.86 | 27.94 | 25.21 | 27.50 | 21.05 | 22.80 |
| InternVL3-78b | 19.00 | 28.86 | 23.83 | 11.76 | 29.90 | 29.06 | 26.50 | 21.05 | 23.25 |
| Qwen2.5-vl-3b | 17.50 | 21.89 | 22.66 | 17.65 | 17.16 | 17.95 | 22.00 | 25.84 | 20.37 |
| Qwen2.5-vl-7b | 20.50 | 20.40 | 20.70 | 8.82 | 22.55 | 26.07 | 24.50 | 22.49 | 20.84 |
| Qwen2.5-vl-32b | 20.50 | 25.87 | 25.39 | 10.78 | 24.51 | 19.66 | 30.50 | 22.49 | 22.48 |
| Qwen2.5-vl-72b | 20.50 | 34.83 | 27.34 | 4.90 | 28.43 | 27.35 | 29.00 | 24.88 | 24.29 |
| Open-source MoE | |||||||||
| Llama-4-Scout | 20.50 | 22.39 | 23.83 | 7.35 | 25.49 | 28.21 | 23.00 | 18.18 | 22.12 |
| Llama-4-Maverick | 14.00 | 42.79 | 17.58 | 5.88 | 37.75 | 37.18 | 36.00 | 20.10 | 26.11 |
| Human | 95.02 | 94.03 | 92.19 | 93.66 | 86.34 | 89.74 | 87.56 | 89.05 | 91.04 |
- GPT-5 (56.41) leads overall but is 34.6 points below human (91.04).
- Gemini-2.5-pro (49.52) and o4-mini (46.47) are second and third โ all reasoning-optimized models.
- Best non-reasoning proprietary: GPT-4.1 (30.90). Best open-source dense: Qwen2.5-vl-72B (24.29). Best open-source MoE: Llama-4-Maverick (26.11).
- 3D Spatial Consistency is the most discriminating subtask: non-reasoning models cluster near random (~6โ14%) while reasoning models jump to 49โ82%. GPT-5 hits 82.35 here.
- All models trail badly on Affordance Recognition (best non-human is GPT-5 at 39.23 vs human 89.05).
- Single-view ablation (Appendix G, Table 7): restricting to the most informative third-person view (AgiWorld head camera / BridgeV2 view1) on the four degradable subtasks drops GPT-5's Distance Judgement by 18.90 points (multi-view 36.32 here, โ +18.90), and GPT-4.1 by +10.44 โ confirming the benchmark genuinely tests multi-view fusion. Gains correlate with model capability: smaller models (Qwen2.5-vl-7B/32B) show negligible or negative โ.
| Model | Variant | Spatial ฮs | Robotic ฮr |
|---|---|---|---|
| Qwen2.5-vl-7b | w cot | +0.58 | โ1.30 |
| w text | โ0.70 | +0.82 | |
| w vggt | โ1.40 | โ0.24 | |
| w depth | +1.04 | โ0.48 | |
| Gemma-3-12B | w cot | +0.93 | +2.96 |
| w text | โ0.94 | โ0.47 | |
| w vggt | โ1.47 | +0.11 | |
| w depth | โ0.18 | +0.19 | |
| GPT-4.1 | w cot | โ1.21 | โ0.25 |
| w text | +1.73 | +1.81 | |
| w vggt | โ1.85 | โ1.58 | |
| w depth | +3.25 | +2.71 |
Synthetic novel views (w vggt) tend to hurt performance โ the authors note current view-synthesis methods perform poorly under narrow baselines, cluttered tabletops, and gripper-centric viewpoints. Depth priors help only when the backbone has capacity to exploit them (GPT-4.1 yes, Qwen2.5-vl-7B partial). Textual CoT shines for mid-capacity open-source models like Gemma-3-12B but barely moves saturated proprietary models.
- Internal axis (Fig. 5). Spatial vs robotic accuracy: positive correlation, but only proprietary and reasoning-optimized models occupy the upward region. Most open-source models cluster near random on both axes.
- External axis (Fig. 6). OmniSpatial (single-view) accuracy vs MV-RoboBench accuracy: strong single-view performance does not transfer reliably. Many models that score well on OmniSpatial remain near random on multi-view tasks.
The paper does not enumerate "limitations" in a single section but the discussion (Section 6) and appendices flag:
- Benchmark scale is modest โ 1.7k items vs 8.6k for Ego3D-Bench or 5k for VSI-Bench โ chosen for manual quality control over template generation.
- Source-domain bias โ derived only from AgiWorld + BridgeV2 (tabletop manipulation). Locomotion, mobile manipulation, and humanoid scenarios are out of scope.
- Multiple-choice format simplifies grading but cannot probe open-ended generation quality. Future work targets free-form spatial reasoning evaluation.
- CoT augmentations are off-the-shelf โ no fine-tuning of view-synthesis or depth models for robot scenes; current novel-view-synthesis baselines genuinely fail on narrow-baseline gripper views.
- Looking forward (Sec. 6). Authors call for (i) architectures with explicit geometric priors and cross-view consistency, (ii) training pipelines aligning perception with action grounding, and (iii) larger-scale multi-camera datasets reflecting real-world manipulation complexity.
MV-RoboBench is the first benchmark coupling multi-view perception with robotic manipulation reasoning. It serves as the diagnostic counterpart to method papers like:
- Spatial Forcing โ claims to inject spatial competence into VLAs via training.
- SP-VLA โ spatially guided VLA inference.
- FALCON โ converts spatial reasoning to action chunks.
For each of these, MV-RoboBench provides a way to ask: did the multi-view spatial competence actually transfer, rather than relying on single-image proxies like SpatialVLM or RoboSpatial. It complements (rather than replaces) ERQA and MMSI-Bench by being the first to put manipulation-style multi-view reasoning front and center.
The benchmark's main empirical result โ that reasoning-optimized models (o4-mini, GPT-5, Gemini-2.5-pro) win by 15โ30 points over non-reasoning peers and that 3D-Spatial-Consistency separates these tiers most cleanly โ is a strong signal that multi-view robotic reasoning rewards explicit chain-of-thought, not raw perception scaling. This validates the bet behind systems like RoboBrain, RoboRefer, and Embodied-R1 that pair VLAs with explicit reasoning passes.
- OpenReview: https://openreview.net/forum?id=jXDZJAfRZB
- PDF: https://openreview.net/pdf?id=jXDZJAfRZB
- arXiv: https://arxiv.org/abs/2510.19400
- Code/data: https://github.com/microsoft/MV-RoboBench
โ Back to ICLR-2026