Review VLA Memory - Heungwoo/research GitHub Wiki

In-Depth Review β€” Memory Architectures for VLAs

Compiled April 2026 Β· Focus: how recent VLA papers (Ο€-series + ICLR 2026 + CoRL 2025 + arXiv preprints 2024–2026) keep past observations and task progress around at inference time, and how those designs compare.

This is a cross-paper review page, not a per-paper summary. For single-paper deep-dives see Review-pi07, Review-VLM4VLA. For related topic landings see RL for VLA.


1. Why memory matters for VLAs

A naive VLA sees only the current frame (or a short stack). That's fine for "pick up the red block" β€” but it breaks on everything interesting:

  • Long-horizon tasks (fold this laundry pile, prep a full meal, clean a room) require memory of what has already been done to avoid undoing it.
  • Occlusion & re-identification β€” the object was on the counter, then hidden behind the fridge door; the policy needs to remember where it was.
  • Goal-state tracking under ambiguous language ("put it back where you found it").
  • Instruction following across sub-steps β€” Ο€0.5/Ο€0.7 coaching requires the model to remember the step-by-step plan.
  • Causal confusion with long raw contexts β€” more frames without structure often hurts because the model latches onto spurious correlations.

By 2026 "what's your memory architecture?" is a first-class design question for any serious VLA. A stateless VLA leaves huge performance on the table β€” HAMLET reports 29.2% β†’ 76.4% success on history-dependent tasks when memory is added on top of the same frozen GR00T N1.5 base.


2. The six categories β€” at a glance

flowchart TB
  A[Cat A: Compressed video tokens<br/>β€” temporal encoder β†’ fixed-size latent]
  B[Cat B: Cross-attention memory banks<br/>β€” moment tokens / recurrent queries]
  C[Cat C: Symbolic / cognitive memory<br/>β€” text-valued propositions]
  D[Cat D: Retrieval-based / episodic stores<br/>β€” external keyframe / embedding bank]
  E[Cat E: Raw-frame history<br/>β€” just extend the context window]
  F[Cat F: Planning-token hierarchies<br/>β€” subtask / subgoal as memory]

  CHOICE{Your VLA needs…}
  CHOICE -- seconds of context,<br/>no retrain --> A
  CHOICE -- history-awareness<br/>on a frozen base --> B
  CHOICE -- debuggable /<br/>auditable memory --> C
  CHOICE -- unbounded /<br/>lifelong recall --> D
  CHOICE -- a quick baseline --> E
  CHOICE -- hierarchical<br/>long-horizon plans --> F
Loading

Each category is detailed in Β§5 with exemplars, pros, cons, and when to pick it.


3. The papers in this review

Paper Venue/year Category Headline claim
MEM (Ο€0.6-MEM, Ο€0.7) arXiv 2603.03596 / 2026 A + C (hybrid) 15-min kitchen horizons via video + text memory
MemoryVLA ICLR 2026, 2508.19236 B + C Perceptual + cognitive banks; +26 pts long-horizon
HAMLET ICLR 2026, 2510.00695 B Plug-and-play moment tokens on a frozen VLA; 29.2 β†’ 76.4%
ContextVLA 2510.04246 / 2025 A Amortize history into a single context token
CronusVLA 2506.19816 / 2025 A Multi-frame feature chunks; 70.9% SimplerEnv
SAM2Act+ 2501.18564 / 2025 B (spatial) SAM2 features + spatial memory bank; 94.3% MemoryBench
MemER 2510.20328 / 2025 D Hierarchical keyframe retrieval; minute-scale recall
Long-VLA CoRL 2025, 2508.19958 (not really memory) Phase-aware attention masking for long horizons
MAP-VLA 2511.09516 / ICRA 2026 D Stage-level soft-prompt library, frozen VLA, +25% real
ReMem-VLA 2603.12942 / 2026 B (recurrent) Dual-level recurrent queries, 5 memory axes
EchoVLA 2511.18112 / 2025 C + D Scene memory + episodic memory for mobile manip
ExpReS-VLA ICRA 2026, 2511.06202 D (adaptation-time) Retrieval + replay for on-device adapt (31 s, 12 demos)
Keyframe-Chaining VLA 2603.01465 / 2026 D Auto keyframes interleaved as visual tokens; 92%
TraceVLA 2412.10345 / 2024 A (pixel-space) Visual trace overlay as history prompt
LoHoVLA 2506.00411 / 2025 F Interleave subtask + action tokens in one stream
EvoVLA 2511.16166 / 2025 F (RL) Gated long-horizon memory for RL shaping
RoboMemory 2508.01415 / 2025 C + D (agentic) Four parallel memories (spatial/temporal/episodic/semantic)
Ctrl-World ICLR 2026, 2510.10125 (world-model) Pose-conditioned memory retrieval for 20 s+ consistency

4. Deep-dives

4.1 MEM β€” Multi-Scale Embodied Memory (Physical Intelligence, 2026)

  • arXiv: 2603.03596 Β· PDF: https://www.pi.website/download/Mem.pdf
  • Mechanism: Bifurcated short-term video memory + long-term text memory. A ViT-based video encoder compresses recent frames into a fixed number of tokens (temporal + spatial compression); the VLM separately emits text-valued propositions about task progress as long-term memory.
  • Representation: Hybrid β€” compressed latent tokens (short-term) + natural-language propositions (long-term).
  • Frozen vs. trainable: Video encoder trainable; VLM + action expert jointly trained (same backbone as Ο€0.6-MEM β†’ Ο€0.7).
  • Pros: 15-min horizons (full kitchen cleanup); addresses causal confusion head-on; long-term text memory is inspectable; drops into the Ο€ series architecture cleanly.
  • Cons: Two encoders to tune; text summarization discards non-verbal cues (fine-grained contact, audio); cold-start on unseen task distributions.
  • Pipeline location: Short-term = encoder-level (visual tokens fed to VLM); long-term = prompt-level (text appended to instruction).
  • Seen in the wild: Ο€0.6-MEM (the MEM variant) β†’ Ο€0.7 (inherits the MEM video-history encoder).

4.2 MemoryVLA β€” Perceptual + Cognitive Banks (ICLR 2026)

  • arXiv: 2508.19236 Β· Project: https://shihao1895.github.io/MemoryVLA/
  • Mechanism: Two external memory banks β€” perceptual (compressed visual features per frame) and cognitive (VLM language-head tokens capturing semantic progress). Per-step consolidation (dedup, compression) and adaptive retrieval.
  • Representation: Two parallel banks; latent tokens in both (cognitive is less interpretable than MEM's raw text long-term).
  • Frozen vs. trainable: VLM backbone mostly frozen; bank read/write + diffusion action expert trained.
  • Pros: +26 pts on long-horizon real-world vs. context-only; human-memory-inspired split is useful; partially interpretable cognitive bank.
  • Cons: Bank management is a new hyperparameter; cognitive tokens are still latent; dual-bank retrieval scheduling is non-trivial.
  • Pipeline location: Cross-attention into the diffusion action expert at each step.
  • Summary page: MemoryVLA

4.3 HAMLET β€” Moment Tokens on a Frozen VLA (ICLR 2026)

  • arXiv: 2510.00695
  • Mechanism: Per-timestep moment tokens appended to VLM input, initialized with time-contrastive learning so they emphasize task-relevant dynamics. A small attention module aggregates them across timesteps into memory features; the base VLA (GR00T N1.5) is frozen.
  • Representation: Small set of learnable tokens, cross-attended into the (frozen) VLA.
  • Frozen vs. trainable: Base VLA frozen; only moment-token head + memory module trained.
  • Pros: Cheapest path to history-awareness on any existing VLA; dramatic empirical delta (29.2% β†’ 76.4% on history-dependent tasks); base model unchanged.
  • Cons: Bounded by token count; tokens are opaque; single-modality memory (no force / tactile).
  • Pipeline location: Cross-attention layer inside the frozen VLM.
  • Summary page: HAMLET

4.4 ContextVLA β€” Amortized Single-Token Context

  • arXiv: 2510.04246
  • Mechanism: Compresses the entire multi-frame history into a single amortized context token, prepended to VLM input.
  • Representation: Exactly one latent token for all past history.
  • Frozen vs. trainable: End-to-end; VLM jointly fine-tuned with the amortization head.
  • Pros: Minimal inference cost; matches multi-frame training benefit without the compute bill; modular.
  • Cons: Extreme bottleneck β€” a single token cannot hold minute-scale history; no event-specific retention.
  • Pipeline location: Encoder-level (prefix to the VLM).

4.5 CronusVLA β€” Feature Chunks Across Time

  • arXiv: 2506.19816
  • Mechanism: Two-stage β€” (1) pretrain VLM on embodied data; (2) adapt into multi-frame policy via learnable feature chunks that aggregate per-frame features.
  • Representation: Multiple learned temporal chunks (more capacity than ContextVLA's single token).
  • Frozen vs. trainable: Backbone jointly fine-tuned.
  • Pros: 70.9% SimplerEnv, +26.8% LIBERO; graceful with # frames; efficient vs. raw multi-frame.
  • Cons: Chunk boundaries fixed, not event-aware; seconds, not minutes.
  • Pipeline location: Encoder-level.

4.6 SAM2Act / SAM2Act+ β€” Visual Foundation Model + Spatial Bank

  • arXiv: 2501.18564
  • Mechanism: Transformer manipulation policy on SAM2 visual features; SAM2Act+ adds an explicit spatial memory bank with cross-attention retrieval. Introduces MemoryBench benchmark.
  • Representation: Spatial-feature bank keyed by current-observation queries.
  • Frozen vs. trainable: SAM2 frozen; bank + policy trained.
  • Pros: 86.8% RLBench (18 tasks); 94.3% MemoryBench; zero-shot spatial memory.
  • Cons: Spatial only (no semantic progress memory); not a full VLA (no language reasoning chain); single-view bias.
  • Pipeline location: Cross-attention bank inside policy network.

4.7 MemER β€” Hierarchical Keyframe Retrieval

  • arXiv: 2510.20328
  • Mechanism: High-level VLM (Qwen2.5-VL) selects/tracks relevant past keyframes from experience; low-level policy (Ο€0.5) acts conditioned on curated keyframes + recent frames.
  • Representation: External episodic store of keyframes, injected as prompt-level visual/text tokens into the low-level controller.
  • Frozen vs. trainable: Low-level Ο€0.5 pretrained; high-level Qwen2.5-VL trained to nominate keyframes; online consolidation.
  • Pros: Genuinely minute-scale recall without ballooning context; amortizes compute by keeping only keyframes; inherits Ο€0.5 quality.
  • Cons: Retrieval quality depends on high-level VLM; failures are opaque (why this frame?); two models to serve.
  • Pipeline location: Retrieval bank + prompt-level injection.

4.8 Long-VLA β€” Phase-Aware Attention Masking (CoRL 2025)

  • arXiv: 2508.19958
  • Mechanism: End-to-end long-horizon VLA using phase-aware input masking β€” binary segmentation of each sub-task into "moving" vs "interaction" phases, applied as attention shaping inside the VLA transformer. Not a memory module per se.
  • Representation: N/A β€” attention mask, not stored memory.
  • Frozen vs. trainable: End-to-end.
  • Pros: Gains on L-CALVIN; simple drop-in; no bank to maintain.
  • Cons: Binary phase segmentation is heuristic; no explicit history beyond context window; "long" = tens of seconds.

4.9 MAP-VLA β€” Memory-Augmented Prompting (2025)

  • arXiv: 2511.09516
  • Mechanism: Build a library of stage-level soft prompts from historical demonstrations; at inference, retrieve by trajectory similarity and inject into a frozen VLA.
  • Representation: Learnable soft prompts (one per task stage).
  • Frozen vs. trainable: VLA frozen; only soft prompts trained via prompt tuning.
  • Pros: True plug-and-play; +7% sim / +25% real on long-horizon.
  • Cons: Prompt library curated offline; trajectory-similarity retrieval can be noisy; no online update.
  • Pipeline location: Prompt-level retrieval bank.

4.10 ReMem-VLA β€” Dual-Level Recurrent Queries (2026)

  • arXiv: 2603.12942
  • Mechanism: Frame-level recurrent queries (short-term) + chunk-level recurrent queries (long-term); auxiliary Past-Observation-Prediction loss to keep queries informative.
  • Representation: Recurrent latent state only β€” no external bank.
  • Frozen vs. trainable: End-to-end.
  • Pros: Covers 5 memory axes (spatial, sequential, episodic, temporal, visual); no extra inference cost.
  • Cons: Recurrent latents are opaque; training stability over long rollouts is tricky.

4.11 EchoVLA β€” Scene + Episodic Memory for Mobile Manip (2025)

  • arXiv: 2511.18112
  • Mechanism: Spatial-semantic scene map + episodic task traces, fused via coarse + fine-grained attention; drives base + arm diffusion policies.
  • Representation: Hybrid β€” graph-structured scene map (semi-symbolic) + episodic trace embeddings.
  • Pros: Explicitly targets mobile manipulation; interpretable scene map; 0.52 / 0.31 sim SR vs Ο€0.5 baseline.
  • Cons: Scene-memory build-up is task-specific; heavy pipeline.

4.12 ExpReS-VLA β€” Experience Replay + Retrieval (ICRA 2026)

  • arXiv: 2511.06202
  • Mechanism: 97%-compressed embeddings from OpenVLA's frozen vision backbone; retrieve top-k by cosine similarity at adaptation time; Thresholded Hybrid Contrastive Loss uses failures as signal.
  • Representation: Compressed visual-embedding buffer (retrieval store).
  • Pros: On-device adaptation in 31 s from 12 demos; combats catastrophic forgetting.
  • Cons: Adaptation-time memory, not inference-time β€” doesn't directly help long horizons in a single rollout.

4.13 Keyframe-Chaining VLA (2026)

  • arXiv: 2603.01465
  • Mechanism: Discriminative embedding space + auto keyframe selection at state transitions + progress-aware retrieval; keyframes interleaved as visual tokens in VLA input.
  • Representation: Sparse semantic-history keyframes as extra visual tokens.
  • Pros: 92% SR on memory-intensive tasks vs 57% baseline; explicit non-Markovian design.
  • Cons: Keyframe selector generalization is the weak link; minute-scale not tested.

4.14 TraceVLA β€” Visual Trace Prompting (2024)

  • arXiv: 2412.10345
  • Mechanism: Point-tracked past trajectory overlaid on the current observation image as a pixel-space visual prompt.
  • Representation: Non-token β€” past trajectory encoded visually.
  • Pros: +10% SimplerEnv, +3.5Γ— real vs OpenVLA; almost free at inference.
  • Cons: 2D trace is limited; doesn't capture semantic events; short-horizon.

4.15 LoHoVLA β€” Unified Subtask + Action Tokens (2025)

  • arXiv: 2506.00411
  • Mechanism: Single VLM backbone jointly generates language subtask tokens and action tokens; hierarchical closed-loop control corrects both planning and control errors.
  • Representation: Interleaved language subtask tokens serve as planning memory substrate.
  • Pros: Unified training objective; subtask tokens double as memory; closed-loop correction.
  • Cons: "Long-horizon" here is still Ravens-scale; 20 tasks; no explicit past-observation store.

4.16 EvoVLA β€” Long-Horizon Memory for RL Shaping (2025)

  • arXiv: 2511.16166
  • Mechanism: Selective context retention + gated fusion over past stages to stabilize intrinsic reward shaping during RL rollouts; combats "stage hallucination."
  • Pros: Useful for long RL rollouts where vanilla policy-gradient variance blows up.
  • Cons: RL-loop-specific; limited real-robot eval.

4.17 RoboMemory β€” Four Parallel Memories (2025, agentic)

  • arXiv: 2508.01415
  • Mechanism: Four memories β€” spatial (knowledge graph), temporal (sequential), episodic (interactions), semantic (concepts) β€” orchestrated by a planner + critic. Prompt-level agentic (Qwen2.5-VL-72B).
  • Pros: +26.5% on EmbodiedBench; beats Claude-3.5-Sonnet on embodied tasks; lifelong-learning story.
  • Cons: Agentic pipeline, not end-to-end VLA; latency.

4.18 Ctrl-World β€” Pose-Conditioned Memory Retrieval (ICLR 2026)

  • arXiv: 2510.10125
  • Mechanism: Generative world model with pose-conditioned memory retrieval keeps 20 s+ spatial/temporal consistency; trained on ~95 k trajectories.
  • Representation: Implicit memory in world-model features, keyed by pose.
  • Pros: Long-horizon consistency for planning / imagined rollouts.
  • Cons: It's a world model, not a VLA β€” memory serves prediction, not action directly. Use in tandem with a policy.
  • Summary page: Ctrl-World

5. Categories compared β€” pros, cons, when to pick

Category A β€” Compressed video tokens (temporal encoders)

  • Mechanism: Pass past frames through a dedicated video/temporal encoder β†’ fixed-size latent tokens prepended to the VLM.
  • Exemplars: MEM (short-term side), ContextVLA, CronusVLA, TraceVLA (pixel-space variant).
  • Pros: βœ… Cheap at inference Β· βœ… backbone unchanged Β· βœ… strong on seconds-to-minutes.
  • Cons: ❌ Hard capacity ceiling Β· ❌ no explicit event retention Β· ❌ opaque.
  • Pick when: You want a few seconds of history and don't want to retrain the VLM.

Category B β€” Cross-attention memory banks (moment tokens / recurrent queries)

  • Mechanism: Maintain a small set of learnable tokens/queries that aggregate past evidence; policy cross-attends.
  • Exemplars: HAMLET (moment tokens), ReMem-VLA (dual-level recurrent), SAM2Act+ (spatial bank), MemoryVLA (perceptual bank side).
  • Pros: βœ… Plug-and-play onto frozen VLAs (HAMLET is canonical) Β· βœ… bounded state Β· βœ… good middle ground.
  • Cons: ❌ Token count caps capacity Β· ❌ opaque latents Β· ❌ no auditability.
  • Pick when: You want history-awareness without touching the base VLA.

Category C β€” Symbolic / cognitive memory (text-valued, interpretable)

  • Mechanism: VLM language head emits text-valued propositions about past state / task progress; stored and re-injected.
  • Exemplars: MEM (long-term text), MemoryVLA (cognitive bank β€” partial), RoboMemory (semantic memory layer), EchoVLA (scene graph).
  • Pros: βœ… Debuggable ("what did the policy know?") Β· βœ… natural LLM-tooling fit Β· βœ… robust across horizons.
  • Cons: ❌ Loses non-verbal cues (contact, force, texture) Β· ❌ language-head hallucination risk Β· ❌ extra VLM calls.
  • Pick when: Minute-scale tasks and you need introspection or auditability.

Category D β€” Retrieval-based / episodic stores

  • Mechanism: External store of past embeddings / keyframes / demos; retrieve at inference or adaptation time.
  • Exemplars: MemER, MAP-VLA, ExpReS-VLA, Keyframe-Chaining VLA, EchoVLA (episodic traces). Navigation-side precedent: ReMEmbR (arXiv 2409.13682, NVIDIA/USC) β€” VILA captions in a Milvus vector DB with LLM-agent retrieval over a robot's spatio-temporal history (NaVQA benchmark); retrieval-memory done right, but for nav-QA rather than manipulation control.
  • Pros: βœ… Unbounded memory Β· βœ… supports lifelong learning Β· βœ… amortizes fine-tuning via retrieval.
  • Cons: ❌ Retrieval-quality bottleneck Β· ❌ two-stage inference Β· ❌ index management at scale.
  • Pick when: Long-horizon or distribution shifts over deployment β€” anywhere an LLM would call for RAG.

Category E β€” Raw-frame history (context-window extension)

  • Mechanism: Concatenate more past frames into the VLM context. No compression, no retrieval.
  • Exemplars: Naive multi-frame OpenVLA baselines; pre-MEM Ο€-series.
  • Pros: βœ… Zero new machinery Β· βœ… benefits from any long-context VLM improvement.
  • Cons: ❌ Quadratic attention cost Β· ❌ causal confusion Β· ❌ memorization of specific histories.
  • Pick when: Baselining or very short histories.

Category F β€” Planning-token / subgoal-memory hierarchies

  • Mechanism: Generate language or latent subgoal tokens; use them as the memory substrate bridging steps.
  • Exemplars: LoHoVLA (subtask + action in one stream), UniVLA (latent action tokens), Ο€0.5 / Ο€0.6 / Ο€0.7 (subtask text + subgoal images), EvoVLA (gated stages). Long-VLA lives nearby via phase masking.
  • Pros: βœ… Memory is actionable (the plan IS the memory) Β· βœ… plays well with LLM-style reasoning Β· βœ… supports error correction.
  • Cons: ❌ Only as good as the subgoal abstraction Β· ❌ doesn't remember fine-grained observations Β· ❌ brittle if subgoals are wrong.
  • Pick when: Genuinely long-horizon tasks with clear sub-structure (cooking, laundry, assembly).

6. Cross-axis comparison

Axis Best in class
Cheapest plug-in onto a frozen VLA HAMLET (B), MAP-VLA (D)
Most interpretable / debuggable MEM long-term text (C), keyframe-based (D)
Longest real horizon today MEM β‰ˆ 15 min (A+C hybrid); MemER minute-scale via (D)
Best for spatial memory specifically SAM2Act+ (B), EchoVLA scene map (D)
Best for on-device adaptation / lifelong ExpReS-VLA (D)
Best for hierarchical compositional tasks Ο€0.7 (F + C), LoHoVLA (F), UniVLA (F)
Headline delta vs. stateless base HAMLET: 29.2 β†’ 76.4% on history-dep tasks (GR00T N1.5)

7. Research trends (2023 β†’ 2026)

  1. From "no memory" (2023) β†’ "fixed context window" (2024) β†’ "architected memory" (2026). OpenVLA and first-gen GR00T ran on single-frame observations. By late 2025 every serious paper has at least a history module; by 2026 "what's your memory design?" is a standard section. HAMLET's 29.2% β†’ 76.4% delta on the same GR00T N1.5 base shows how much performance was being left on the table.

  2. Hybrid multi-scale is winning over single-scale. MEM (video + text), MemoryVLA (perceptual + cognitive), ReMem-VLA (frame + chunk), EchoVLA (scene + episodic), RoboMemory (four types). The field converged on "short-term dense + long-term symbolic" after trying each separately.

  3. Plug-and-play / frozen-backbone memory modules are ascendant. HAMLET, MAP-VLA, MemoryVLA (mostly), ExpReS-VLA. The economics favor adding memory as a cheap adapter on top of an expensive pretrained VLA β€” mirroring the LLM field's LoRA-over-full-fine-tune move.

  4. Retrieval is eating robotic memory, just as it ate LLM context. MemER, MAP-VLA, ExpReS-VLA, Keyframe-Chaining, EchoVLA β€” all use retrieval. The VLA world is reinventing RAG with embodied keys (poses, keyframes, trajectories).

  5. Prompt expansion as the integration surface. Ο€0.7 is the clearest instance: instead of new memory hardware, dump subtask text, subgoal images, episode metadata, and MEM summaries into the prompt. Classifier-free guidance does the rest. Expect 2026–2027 to push this harder β€” multimodal prompt = multimodal memory.

ICRA 2026 developments

The ICRA 2026 cohort is striking for where it puts memory: not in the architecture-and-benchmark frame of the ML venues, but squarely in a "memory as a cheap adapter on a frozen, deployed VLA" posture (see the ICRA 2026 Survey, which files these under a dedicated "Memory & retrieval" cluster). All three of its memory contributions are retrieval-flavored, which is the strongest single-venue evidence yet for trend #4.

  • MAP-VLA (Cat D, already a row above) is the canonical instance: a library of stage-level soft prompts built offline from demos, retrieved by β„“β‚‚ trajectory-similarity over a sliding window restricted to neighboring stages, and injected by element-wise addition onto the base prompt of a frozen Ο€0. The economics are explicitly RAG-over-frozen-model + LoRA-over-full-finetune; the tell is that the real-robot gain (+25%) dwarfs the sim gain (+7%), i.e. stage memory matters most exactly where long-horizon distribution shift bites.
  • ExpReS-VLA (Cat D, already a row) reinforces the same retrieval economics on the adaptation axis rather than the inference axis β€” replay + cosine-similarity retrieval over compressed embeddings to specialize a generalist (OpenVLA) to a fixed deployment without full retraining.
  • TrackVLA++ (2510.07134) extends the pattern beyond manipulation: a Target-Identification Memory (paired with a Polar-CoT spatial reasoning token) lets an embodied visual-tracking VLA re-identify its target across occlusion and distractors β€” the navigation/tracking analogue of the re-identification problem in Β§1, and another vote for explicit per-target episodic state over a wider context window.

Read together, ICRA 2026 says the deployment community has already picked a side: lightweight retrieval/prompting over a frozen backbone beats retraining, and retrieval keys are diversifying (trajectory windows, compressed embeddings, tracked-target identity) exactly as predicted by the "embodied RAG" thesis.


8. Open questions

  • Selective write is unsolved. Most papers either "retain everything compressed" or "learn a selector." Robot memory actually needs attention at encoding time (what's worth remembering?), not just at retrieval time. HAMLET's time-contrastive init and MemER's keyframe nomination are first steps; neither is robust under distribution shift.
  • Minute-scale β†’ hour-scale. MEM reaches 15 min. Nobody shows hour-scale continuous tasks yet (full meal prep end-to-end). The memory compression curve beyond 15 min is unknown.
  • Episodic vs. parametric β€” undecided. ExpReS-VLA shows retrieval + frozen backbone can adapt in 31 s from 12 demos β€” striking. But Ο€*0.6/RECAP shows parametric wins with enough data. The answer is almost certainly episodic for tail distributions, parametric for the head β€” same story as LLMs.
  • VLA memory ↔ LLM RAG convergence. Directionally yes (MemER β‰ˆ RAG for robots), but the modality differs β€” robot memory needs action-outcomes and spatial state, not just text. Expect convergence on a shared interface (retrieval + structured re-injection) with robotics-specific keys (pose, proprio, contact).
  • Interpretability vs. performance. Text-valued memory (MEM long-term, RoboMemory semantic) is debuggable but lossy. Latent memory (moment tokens, recurrent queries) is performant but opaque. A standardized "memory probe" that dumps a policy's state as natural language would unblock safety work β€” nobody has this yet.
  • No standard benchmark. SAM2Act has MemoryBench; Long-VLA has L-CALVIN; MemER has a real-robot suite; HAMLET has history-dependent tasks. Cross-paper comparison is near impossible. A canonical "VLA-MemoryBench" covering minute-scale, occlusion, spatial recall, and episodic transfer would be extremely useful.
  • World-model memory vs. control memory. Ctrl-World has pose-conditioned memory retrieval for imagination. Whether imagination-memory and control-memory should share a substrate (cf. Ο€0.7's subgoal images from BAGEL-14B) is an open design question.
  • Memory + RL is underexplored. EvoVLA is the main attempt. If memory reduces long-rollout variance, RL post-training (Ο€*0.6 / RECAP style) could scale further.

9. Practical decision guide

If you're shipping a VLA and wondering where to spend a week:

  1. No history module yet? β†’ Start with HAMLET-style moment tokens. Plug-and-play, frozen base, largest per-effort delta.
  2. Tasks longer than ~30 s? β†’ Add MEM-style video + text memory or MemER-style keyframe retrieval. Don't stack raw frames.
  3. Need debuggability / safety probes? β†’ Include a text-valued long-term memory (MEM-style) even if you also keep latent banks.
  4. Distribution shifts over deployment (new homes, new users)? β†’ Add a retrieval store (MAP-VLA or ExpReS-VLA) rather than fine-tune in place.
  5. Task has obvious hierarchical structure (cooking, laundry)? β†’ Use planning-token / subgoal memory (LoHoVLA-style or Ο€0.7-style prompt expansion).
  6. Contact-rich manipulation? β†’ Remember that all memory modules above are visual. Torque / tactile memory is still an open gap (see TA-VLA for torque conditioning but not memory).

10. Sources

Individual arXiv links are inline in the table (Β§3). Wiki cross-links:


πŸ—“ State of the Field (updated Aug 2026)

Verdict: the frame shifted from bigger in-policy memory to selective history and agentic externalized memory β€” with navigation, not manipulation, holding the working demo.

πŸ“ˆ Trend

RSS 2026 moved the frame twice: (1) selective history beats full context β€” key-history-frame retrieval (RSS #201) and memory-retrieval studies (RSS #10) outperform naively long windows, and RobotManip's stochastic context sampling shows why (recency shortcuts collapse the mechanism); (2) agentic decomposition arrived β€” Qwen-RobotNav's planner/executor split with a persistent evidence notebook (survives context compression, auditable belief revisions) set EQA SOTA with 77% fewer steps.

βš–οΈ Approaches & trade-offs

Approach Pros Cons
In-policy memory modules (the 6 architectures above) End-to-end, fast Context cost; recency-shortcut risk
Retrieval over episodic history Scales with episode length Bounded by retrieval quality
Agentic notebook memory (planner + executor) Auditable, compressible, composable Planner latency; interface design becomes the research object
Hierarchical/spatial memory (ICML 2026) HiMe: Executor/Sentry/Planner tiers with add/update/delete knowledge ops; SOMA: persistent spatial-semantic memory enables out-of-view manipulation Young; benchmark coverage thin
Structured vs parametric memory (IROS 2026 πŸ†•) dedicated "Building Memory into VL Robot Agents" session; GaussMemory (3D-Gaussian scene memory) & PROMPT (probabilistic memory trees) = structured; TempoFit (layer-wise temporal KV memory for long-horizon VLA) & RoboSSM (state-space) = parametric/context β€” the field bifurcates Structured: build/maintain cost; parametric: opaque, bounded by context mechanism

⚠️ Limitations & open problems

  • Long-horizon manipulation autonomy record is 11 stages / 2.5 min (ViTacFormer) β€” minutes, not hours.
  • Drift and mid-task recovery are barely measured (RoboMME, now an ICML 2026 Oral, is the main probe; CAPS reframes long-horizon instruction drift as a correctable sampling error β€” training-free SNR-triggered MCMC).
  • No manipulation-side equivalent of RobotNav's agentic demo exists yet.

← Back to Home

⚠️ **GitHub.com Fallback** ⚠️