CVPR 2026 EchoVLA - Heungwoo/research GitHub Wiki

EchoVLA — Synergistic Declarative Memory for VLA-Driven Mobile Manipulation

Venue: CVPR 2026 (likely — pending CVF virtual-page confirmation) Category: Mobile Manipulation / Memory Trend tag: Trend 1 Affiliations: Sun Yat-sen University (Shenzhen) + Shanghai Jiao Tong University + Huawei Noah's Ark Lab

Approach diagram

flowchart LR
  EXP["episode experience"] --> SS["scene memory<br/>voxel map (PHC)"]
  EXP --> ET["episodic memory<br/>token buffer (hippocampus)"]
  OBS["current obs"] --> SS
  OBS --> ET
  SS -->|coarse cross-attn| CAT["concat → H_t"]
  ET -->|fine cross-attn| CAT
  CAT --> POL["per-part (base–arm)<br/>diffusion policy"]
  POL --> ACT["mobile manipulation action"]
Loading

Problem

Mobile manipulation requires both spatial memory (where things are in the environment) and episodic memory (what the robot just did and how it went). Most VLAs have neither; the few that do treat them as a single bank.

Method

A brain-inspired (declarative) memory with two complementary stores, mapped to distinct neural substrates:

  • Scene memory — a persistent, slowly-varying voxelized spatial-semantic map (parahippocampal-cortex analogue) holding stable 3D structure and object/room layout.
  • Episodic memory — a fixed-size FIFO token buffer of time-indexed multimodal state tokens (hippocampus analogue), preserving fine-grained recent task progress (e.g., whether a drawer was opened, an object grasped).

The two memories are stored, updated, and retrieved independently — not jointly queried. Retrieval is a coarse-to-fine hierarchy: top-k entries are selected by cosine similarity, then scene memory is read via coarse-grained cross-attention (query = current 3D voxel map) and episodic memory via fine-grained cross-attention (query = current state tokens). The two outputs are concatenated into a memory-augmented representation H_t that conditions the policy.

The policy is a per-part (base–arm) diffusion policy: separate denoising processes for the mobile-base and arm action subspaces, conditioned on H_t.

Results

On the RoboCasa simulator, EchoVLA reaches 0.52 SR on manipulation/navigation tasks (+0.20 over π0.5) and 0.31 SR on mobile manipulation tasks (+0.11 over π0.5). In a real-world 7m × 7m arena, it achieves the highest SR of 0.44, beating π0.5 (0.33) and Diffusion Policy (0.32).

Training is supported by MoMani, an automated benchmark that generates expert-level trajectories via MLLM-guided planning and feedback-driven refinement, supplemented with real-robot demonstrations.

Significance

EchoVLA's bet: mobile manipulation needs cortical-style memory, not LLM-style context-window memory. The split between scene (parahippocampal-cortex-style spatial) and episodic (hippocampus-style temporal) stores mirrors the cognitive-science distinction, and the coarse-to-fine retrieval hierarchy follows from their differing semantic granularity (slow spatial structure vs. fast time-indexed task progress). The empirical question is whether the split helps — and the RoboCasa/real-arena gains over π0.5 (+0.20 / +0.11 SR) are the paper's evidence that it does.

Links

Related pages

← Back to CVPR-2026

⚠️ **GitHub.com Fallback** ⚠️