ICML 2026 Spatial Memory for Out of Vision Manipulation - Heungwoo/research GitHub Wiki

SOMA: Spatial Memory for Out-of-Vision Manipulation in VLA — Reasoning beyond the camera frustum

Venue: ICML 2026 (Poster) Category: VLA Architecture Traction (2026-06): 0 citations (arXiv)

The Out-of-Vision (OOV) limitation: existing VLAs are purely reactive and fail when targets leave the field of view (Figure 1 / x1 from the SOMA authors, 2026)

Problem

Most VLAs implicitly assume task-relevant objects are always visible, producing brittle, reactive behavior when targets fall outside the camera's field of view. Real manipulation — multi-step routines, dual-arm tasks, large workspaces — routinely requires acting on objects that are not currently in view. Reactive viewpoint adjustment alone is insufficient because spatial information must be preserved and reused across stages.

Method

SOMA equips VLAs with a persistent spatial memory built from multi-view observations acquired by a movable head camera, letting the policy reason beyond the current visual frustum. It has three components:

The SOMA framework: spatial memory construction, dynamic refinement, and contextual retrieval feeding the VLA (Figure 2 / x2 from the SOMA authors, 2026)

  • Spatial Memory Construction. A scanning video is subsampled, and each frame runs through a unified perception pipeline: VGGT for camera pose and coarse 3D geometry, YOLO for object localization/categorization, and DINOv3 for high-dimensional visual features. Per-instance appearance embeddings are fused with estimated 3D locations into an overview scene memory M₀ — a unified spatial-semantic map.
  • Dynamic Memory Refinement. As objects move, occlude, or appear, the memory is iteratively updated from the latest head-camera observation, with each memory token combining a semantic embedding and a positional embedding from the instance's estimated 3D location, keeping the representation globally consistent over time.
  • Contextual Memory Retrieval. VLM vision-language tokens act as queries in a cross-attention module over the scene memory (keys/values, after an alignment projection), injecting instruction-relevant spatial cues into the multimodal representation for grounded reasoning.

Results

SOMA is evaluated on five self-designed real-world OOV tasks (including multi-step and dual-arm scenarios) where targets are initially invisible, plus simulation. In the real world it achieves the highest success rates across all five tasks; a fixed-head variant fails whenever target or goal leaves view, and active-head baselines (StarVLA, SpatialVLA, GR00T-N1.5) degrade in multi-stage/dual-arm settings. Beyond success rate, SOMA induces qualitatively different behavior — vs. GR00T-N1.5, it cuts first-fixation time by 40–59% and shortens head-search path length, achieving faster localization and near one-shot grasping. In simulation, SOMA reaches 52.0% avg SR on GR-1 (300 demos), surpassing Diffusion Policy, GR00T, and StarVLA, with >15% gains on Dish Transfer and Tabletop Serving and strong sample efficiency at 30–100 demos. On SimplerEnv Visual Matching it attains the best 63.2% SR. Ablations show removing positional cues (49.3→45.1%), object semantics (→43.7%), or the dynamic update (→41.5%) each hurts, with the dynamic refinement being the most impactful.

Significance

SOMA argues that persistent structured spatial memory, not reactive viewpoint search, is the key to manipulation under partial observability. By aggregating multi-view geometry and semantics into a retrievable memory bank, it enables VLAs to manipulate objects outside the current view while also improving conventional fully-observable benchmarks — a meaningful step toward spatially-grounded, memory-augmented robot policies.

Links

← Back to ICML-2026