ICLR 2026 From Seeing To Doing - Heungwoo/research GitHub Wiki

From Seeing to Doing โ€” Spatial Reasoning Bridge for VLA

Venue: ICLR 2026 Category: Embodied reasoning ยท Spatial CoT Trend tag: Reasoning / decision bridge

Approach diagram

flowchart LR
  Img[Image + instruction] --> SR[Spatial reasoning step<br/>intermediate spatial representation]
  SR --> SCA[Self-consistency alignment<br/>coords โ†” visual signals]
  SCA --> Act[Manipulation policy / action]
Loading

Problem

General-purpose VLMs underlie most VLA models, but their zero-shot manipulation performance is fragile because embodied datasets are scarce and heterogeneous. The paper argues the missing link between "seeing" and "doing" is an intermediate spatial-reasoning stage that gives the policy fine-grained guidance.

Method

FSD (the evaluated model is FSD-13B) is a VLM trained to produce intermediate spatial representations โ€” a Spatial Relationship-Focused Visual Chain-of-Thought (Sr-CoT) that performs multi-step reasoning anchored by object coordinates and spatial relationships โ€” before action. Two ingredients: (1) a self-consistency mechanism that binds predicted spatial coordinates to specific visual signals, aligning understanding and generation during training, and (2) a weak-to-strong data construction pipeline that combines large-scale embodied datasets with common-sense data to build spatial-reasoning supervision from heterogeneous sources.

Results

  • 8 general spatial-reasoning / embodied-reference benchmarks: best overall ranking (avg rank 1.3 across CVBench, CRPE, SAT, BLINK, EmbSpatial; embodied reference RoboRefIt 56.7% vs RoboPoint 49.8% / GPT-4o 15.3%, Where2Place 45.8%).
  • VABench (the paper's own, more challenging benchmark): VABench-Point 61.82% accuracy (vs GPT-4o 9.30%); plus a VABench-VisualTrace split.
  • SimplerEnv: 40.6% average zero-shot success.
  • Real-world: 72% across 8 tasks.
  • Reported to outperform the strongest baseline by ~30% in robot settings.

Significance

Positions spatial reasoning as the explicit interface between perception and action โ€” an alternative to monolithic VLA fine-tuning and to pointing-only intermediates like Embodied-R1. Suggests that scaling reasoning supervision (not just trajectories) is a viable axis for VLA generalization.

Links

Related pages

โ† Back to ICLR-2026

โš ๏ธ **GitHub.com Fallback** โš ๏ธ