CVPR 2026 EchoVLA - Heungwoo/research GitHub Wiki
Venue: CVPR 2026 (likely — pending CVF virtual-page confirmation) Category: Mobile Manipulation / Memory Trend tag: Trend 1 Affiliations: Sun Yat-sen University (Shenzhen) + Shanghai Jiao Tong University + Huawei Noah's Ark Lab
flowchart LR
EXP["episode experience"] --> SS["scene memory<br/>voxel map (PHC)"]
EXP --> ET["episodic memory<br/>token buffer (hippocampus)"]
OBS["current obs"] --> SS
OBS --> ET
SS -->|coarse cross-attn| CAT["concat → H_t"]
ET -->|fine cross-attn| CAT
CAT --> POL["per-part (base–arm)<br/>diffusion policy"]
POL --> ACT["mobile manipulation action"]
Mobile manipulation requires both spatial memory (where things are in the environment) and episodic memory (what the robot just did and how it went). Most VLAs have neither; the few that do treat them as a single bank.
A brain-inspired (declarative) memory with two complementary stores, mapped to distinct neural substrates:
- Scene memory — a persistent, slowly-varying voxelized spatial-semantic map (parahippocampal-cortex analogue) holding stable 3D structure and object/room layout.
- Episodic memory — a fixed-size FIFO token buffer of time-indexed multimodal state tokens (hippocampus analogue), preserving fine-grained recent task progress (e.g., whether a drawer was opened, an object grasped).
The two memories are stored, updated, and retrieved independently — not jointly queried. Retrieval is a coarse-to-fine hierarchy: top-k entries are selected by cosine similarity, then scene memory is read via coarse-grained cross-attention (query = current 3D voxel map) and episodic memory via fine-grained cross-attention (query = current state tokens). The two outputs are concatenated into a memory-augmented representation H_t that conditions the policy.
The policy is a per-part (base–arm) diffusion policy: separate denoising processes for the mobile-base and arm action subspaces, conditioned on H_t.
On the RoboCasa simulator, EchoVLA reaches 0.52 SR on manipulation/navigation tasks (+0.20 over π0.5) and 0.31 SR on mobile manipulation tasks (+0.11 over π0.5). In a real-world 7m × 7m arena, it achieves the highest SR of 0.44, beating π0.5 (0.33) and Diffusion Policy (0.32).
Training is supported by MoMani, an automated benchmark that generates expert-level trajectories via MLLM-guided planning and feedback-driven refinement, supplemented with real-robot demonstrations.
EchoVLA's bet: mobile manipulation needs cortical-style memory, not LLM-style context-window memory. The split between scene (parahippocampal-cortex-style spatial) and episodic (hippocampus-style temporal) stores mirrors the cognitive-science distinction, and the coarse-to-fine retrieval hierarchy follows from their differing semantic granularity (slow spatial structure vs. fast time-indexed task progress). The empirical question is whether the split helps — and the RoboCasa/real-arena gains over π0.5 (+0.20 / +0.11 SR) are the paper's evidence that it does.
- arXiv: 2511.18112
- VLA Memory · OptimusVLA (sister dual-memory paper)
- CVPR 2026 survey
← Back to CVPR-2026