ICML 2026 Spatial Memory for Out of Vision Manipulation - Heungwoo/research GitHub Wiki
SOMA: Spatial Memory for Out-of-Vision Manipulation in VLA — Reasoning beyond the camera frustum
Venue: ICML 2026 (Poster) Category: VLA Architecture Traction (2026-06): 0 citations (arXiv)

Problem
Most VLAs implicitly assume task-relevant objects are always visible, producing brittle, reactive behavior when targets fall outside the camera's field of view. Real manipulation — multi-step routines, dual-arm tasks, large workspaces — routinely requires acting on objects that are not currently in view. Reactive viewpoint adjustment alone is insufficient because spatial information must be preserved and reused across stages.
Method
SOMA equips VLAs with a persistent spatial memory built from multi-view observations acquired by a movable head camera, letting the policy reason beyond the current visual frustum. It has three components:

- Spatial Memory Construction. A scanning video is subsampled, and each frame runs through a unified perception pipeline: VGGT for camera pose and coarse 3D geometry, YOLO for object localization/categorization, and DINOv3 for high-dimensional visual features. Per-instance appearance embeddings are fused with estimated 3D locations into an overview scene memory M₀ — a unified spatial-semantic map.
- Dynamic Memory Refinement. As objects move, occlude, or appear, the memory is iteratively updated from the latest head-camera observation, with each memory token combining a semantic embedding and a positional embedding from the instance's estimated 3D location, keeping the representation globally consistent over time.
- Contextual Memory Retrieval. VLM vision-language tokens act as queries in a cross-attention module over the scene memory (keys/values, after an alignment projection), injecting instruction-relevant spatial cues into the multimodal representation for grounded reasoning.
Results
SOMA is evaluated on five self-designed real-world OOV tasks (including multi-step and dual-arm scenarios) where targets are initially invisible, plus simulation. In the real world it achieves the highest success rates across all five tasks; a fixed-head variant fails whenever target or goal leaves view, and active-head baselines (StarVLA, SpatialVLA, GR00T-N1.5) degrade in multi-stage/dual-arm settings. Beyond success rate, SOMA induces qualitatively different behavior — vs. GR00T-N1.5, it cuts first-fixation time by 40–59% and shortens head-search path length, achieving faster localization and near one-shot grasping. In simulation, SOMA reaches 52.0% avg SR on GR-1 (300 demos), surpassing Diffusion Policy, GR00T, and StarVLA, with >15% gains on Dish Transfer and Tabletop Serving and strong sample efficiency at 30–100 demos. On SimplerEnv Visual Matching it attains the best 63.2% SR. Ablations show removing positional cues (49.3→45.1%), object semantics (→43.7%), or the dynamic update (→41.5%) each hurts, with the dynamic refinement being the most impactful.
Significance
SOMA argues that persistent structured spatial memory, not reactive viewpoint search, is the key to manipulation under partial observability. By aggregating multi-view geometry and semantics into a retrievable memory bank, it enables VLAs to manipulate objects outside the current view while also improving conventional fully-observable benchmarks — a meaningful step toward spatially-grounded, memory-augmented robot policies.
Links
- arXiv: 2605.22283
- ICML 2026: https://icml.cc/virtual/2026/poster/66214
← Back to ICML-2026