Review VLA Memory - Heungwoo/research GitHub Wiki
Compiled April 2026 Β· Focus: how recent VLA papers (Ο-series + ICLR 2026 + CoRL 2025 + arXiv preprints 2024β2026) keep past observations and task progress around at inference time, and how those designs compare.
This is a cross-paper review page, not a per-paper summary. For single-paper deep-dives see Review-pi07, Review-VLM4VLA. For related topic landings see RL for VLA.
A naive VLA sees only the current frame (or a short stack). That's fine for "pick up the red block" β but it breaks on everything interesting:
- Long-horizon tasks (fold this laundry pile, prep a full meal, clean a room) require memory of what has already been done to avoid undoing it.
- Occlusion & re-identification β the object was on the counter, then hidden behind the fridge door; the policy needs to remember where it was.
- Goal-state tracking under ambiguous language ("put it back where you found it").
- Instruction following across sub-steps β Ο0.5/Ο0.7 coaching requires the model to remember the step-by-step plan.
- Causal confusion with long raw contexts β more frames without structure often hurts because the model latches onto spurious correlations.
By 2026 "what's your memory architecture?" is a first-class design question for any serious VLA. A stateless VLA leaves huge performance on the table β HAMLET reports 29.2% β 76.4% success on history-dependent tasks when memory is added on top of the same frozen GR00T N1.5 base.
flowchart TB
A[Cat A: Compressed video tokens<br/>β temporal encoder β fixed-size latent]
B[Cat B: Cross-attention memory banks<br/>β moment tokens / recurrent queries]
C[Cat C: Symbolic / cognitive memory<br/>β text-valued propositions]
D[Cat D: Retrieval-based / episodic stores<br/>β external keyframe / embedding bank]
E[Cat E: Raw-frame history<br/>β just extend the context window]
F[Cat F: Planning-token hierarchies<br/>β subtask / subgoal as memory]
CHOICE{Your VLA needsβ¦}
CHOICE -- seconds of context,<br/>no retrain --> A
CHOICE -- history-awareness<br/>on a frozen base --> B
CHOICE -- debuggable /<br/>auditable memory --> C
CHOICE -- unbounded /<br/>lifelong recall --> D
CHOICE -- a quick baseline --> E
CHOICE -- hierarchical<br/>long-horizon plans --> F
Each category is detailed in Β§5 with exemplars, pros, cons, and when to pick it.
| Paper | Venue/year | Category | Headline claim |
|---|---|---|---|
| MEM (Ο0.6-MEM, Ο0.7) | arXiv 2603.03596 / 2026 | A + C (hybrid) | 15-min kitchen horizons via video + text memory |
| MemoryVLA | ICLR 2026, 2508.19236 | B + C | Perceptual + cognitive banks; +26 pts long-horizon |
| HAMLET | ICLR 2026, 2510.00695 | B | Plug-and-play moment tokens on a frozen VLA; 29.2 β 76.4% |
| ContextVLA | 2510.04246 / 2025 | A | Amortize history into a single context token |
| CronusVLA | 2506.19816 / 2025 | A | Multi-frame feature chunks; 70.9% SimplerEnv |
| SAM2Act+ | 2501.18564 / 2025 | B (spatial) | SAM2 features + spatial memory bank; 94.3% MemoryBench |
| MemER | 2510.20328 / 2025 | D | Hierarchical keyframe retrieval; minute-scale recall |
| Long-VLA | CoRL 2025, 2508.19958 | (not really memory) | Phase-aware attention masking for long horizons |
| MAP-VLA | 2511.09516 / ICRA 2026 | D | Stage-level soft-prompt library, frozen VLA, +25% real |
| ReMem-VLA | 2603.12942 / 2026 | B (recurrent) | Dual-level recurrent queries, 5 memory axes |
| EchoVLA | 2511.18112 / 2025 | C + D | Scene memory + episodic memory for mobile manip |
| ExpReS-VLA | ICRA 2026, 2511.06202 | D (adaptation-time) | Retrieval + replay for on-device adapt (31 s, 12 demos) |
| Keyframe-Chaining VLA | 2603.01465 / 2026 | D | Auto keyframes interleaved as visual tokens; 92% |
| TraceVLA | 2412.10345 / 2024 | A (pixel-space) | Visual trace overlay as history prompt |
| LoHoVLA | 2506.00411 / 2025 | F | Interleave subtask + action tokens in one stream |
| EvoVLA | 2511.16166 / 2025 | F (RL) | Gated long-horizon memory for RL shaping |
| RoboMemory | 2508.01415 / 2025 | C + D (agentic) | Four parallel memories (spatial/temporal/episodic/semantic) |
| Ctrl-World | ICLR 2026, 2510.10125 | (world-model) | Pose-conditioned memory retrieval for 20 s+ consistency |
- arXiv: 2603.03596 Β· PDF: https://www.pi.website/download/Mem.pdf
- Mechanism: Bifurcated short-term video memory + long-term text memory. A ViT-based video encoder compresses recent frames into a fixed number of tokens (temporal + spatial compression); the VLM separately emits text-valued propositions about task progress as long-term memory.
- Representation: Hybrid β compressed latent tokens (short-term) + natural-language propositions (long-term).
- Frozen vs. trainable: Video encoder trainable; VLM + action expert jointly trained (same backbone as Ο0.6-MEM β Ο0.7).
- Pros: 15-min horizons (full kitchen cleanup); addresses causal confusion head-on; long-term text memory is inspectable; drops into the Ο series architecture cleanly.
- Cons: Two encoders to tune; text summarization discards non-verbal cues (fine-grained contact, audio); cold-start on unseen task distributions.
- Pipeline location: Short-term = encoder-level (visual tokens fed to VLM); long-term = prompt-level (text appended to instruction).
- Seen in the wild: Ο0.6-MEM (the MEM variant) β Ο0.7 (inherits the MEM video-history encoder).
- arXiv: 2508.19236 Β· Project: https://shihao1895.github.io/MemoryVLA/
- Mechanism: Two external memory banks β perceptual (compressed visual features per frame) and cognitive (VLM language-head tokens capturing semantic progress). Per-step consolidation (dedup, compression) and adaptive retrieval.
- Representation: Two parallel banks; latent tokens in both (cognitive is less interpretable than MEM's raw text long-term).
- Frozen vs. trainable: VLM backbone mostly frozen; bank read/write + diffusion action expert trained.
- Pros: +26 pts on long-horizon real-world vs. context-only; human-memory-inspired split is useful; partially interpretable cognitive bank.
- Cons: Bank management is a new hyperparameter; cognitive tokens are still latent; dual-bank retrieval scheduling is non-trivial.
- Pipeline location: Cross-attention into the diffusion action expert at each step.
- Summary page: MemoryVLA
- arXiv: 2510.00695
- Mechanism: Per-timestep moment tokens appended to VLM input, initialized with time-contrastive learning so they emphasize task-relevant dynamics. A small attention module aggregates them across timesteps into memory features; the base VLA (GR00T N1.5) is frozen.
- Representation: Small set of learnable tokens, cross-attended into the (frozen) VLA.
- Frozen vs. trainable: Base VLA frozen; only moment-token head + memory module trained.
- Pros: Cheapest path to history-awareness on any existing VLA; dramatic empirical delta (29.2% β 76.4% on history-dependent tasks); base model unchanged.
- Cons: Bounded by token count; tokens are opaque; single-modality memory (no force / tactile).
- Pipeline location: Cross-attention layer inside the frozen VLM.
- Summary page: HAMLET
- arXiv: 2510.04246
- Mechanism: Compresses the entire multi-frame history into a single amortized context token, prepended to VLM input.
- Representation: Exactly one latent token for all past history.
- Frozen vs. trainable: End-to-end; VLM jointly fine-tuned with the amortization head.
- Pros: Minimal inference cost; matches multi-frame training benefit without the compute bill; modular.
- Cons: Extreme bottleneck β a single token cannot hold minute-scale history; no event-specific retention.
- Pipeline location: Encoder-level (prefix to the VLM).
- arXiv: 2506.19816
- Mechanism: Two-stage β (1) pretrain VLM on embodied data; (2) adapt into multi-frame policy via learnable feature chunks that aggregate per-frame features.
- Representation: Multiple learned temporal chunks (more capacity than ContextVLA's single token).
- Frozen vs. trainable: Backbone jointly fine-tuned.
- Pros: 70.9% SimplerEnv, +26.8% LIBERO; graceful with # frames; efficient vs. raw multi-frame.
- Cons: Chunk boundaries fixed, not event-aware; seconds, not minutes.
- Pipeline location: Encoder-level.
- arXiv: 2501.18564
- Mechanism: Transformer manipulation policy on SAM2 visual features; SAM2Act+ adds an explicit spatial memory bank with cross-attention retrieval. Introduces MemoryBench benchmark.
- Representation: Spatial-feature bank keyed by current-observation queries.
- Frozen vs. trainable: SAM2 frozen; bank + policy trained.
- Pros: 86.8% RLBench (18 tasks); 94.3% MemoryBench; zero-shot spatial memory.
- Cons: Spatial only (no semantic progress memory); not a full VLA (no language reasoning chain); single-view bias.
- Pipeline location: Cross-attention bank inside policy network.
- arXiv: 2510.20328
- Mechanism: High-level VLM (Qwen2.5-VL) selects/tracks relevant past keyframes from experience; low-level policy (Ο0.5) acts conditioned on curated keyframes + recent frames.
- Representation: External episodic store of keyframes, injected as prompt-level visual/text tokens into the low-level controller.
- Frozen vs. trainable: Low-level Ο0.5 pretrained; high-level Qwen2.5-VL trained to nominate keyframes; online consolidation.
- Pros: Genuinely minute-scale recall without ballooning context; amortizes compute by keeping only keyframes; inherits Ο0.5 quality.
- Cons: Retrieval quality depends on high-level VLM; failures are opaque (why this frame?); two models to serve.
- Pipeline location: Retrieval bank + prompt-level injection.
- arXiv: 2508.19958
- Mechanism: End-to-end long-horizon VLA using phase-aware input masking β binary segmentation of each sub-task into "moving" vs "interaction" phases, applied as attention shaping inside the VLA transformer. Not a memory module per se.
- Representation: N/A β attention mask, not stored memory.
- Frozen vs. trainable: End-to-end.
- Pros: Gains on L-CALVIN; simple drop-in; no bank to maintain.
- Cons: Binary phase segmentation is heuristic; no explicit history beyond context window; "long" = tens of seconds.
- arXiv: 2511.09516
- Mechanism: Build a library of stage-level soft prompts from historical demonstrations; at inference, retrieve by trajectory similarity and inject into a frozen VLA.
- Representation: Learnable soft prompts (one per task stage).
- Frozen vs. trainable: VLA frozen; only soft prompts trained via prompt tuning.
- Pros: True plug-and-play; +7% sim / +25% real on long-horizon.
- Cons: Prompt library curated offline; trajectory-similarity retrieval can be noisy; no online update.
- Pipeline location: Prompt-level retrieval bank.
- arXiv: 2603.12942
- Mechanism: Frame-level recurrent queries (short-term) + chunk-level recurrent queries (long-term); auxiliary Past-Observation-Prediction loss to keep queries informative.
- Representation: Recurrent latent state only β no external bank.
- Frozen vs. trainable: End-to-end.
- Pros: Covers 5 memory axes (spatial, sequential, episodic, temporal, visual); no extra inference cost.
- Cons: Recurrent latents are opaque; training stability over long rollouts is tricky.
- arXiv: 2511.18112
- Mechanism: Spatial-semantic scene map + episodic task traces, fused via coarse + fine-grained attention; drives base + arm diffusion policies.
- Representation: Hybrid β graph-structured scene map (semi-symbolic) + episodic trace embeddings.
- Pros: Explicitly targets mobile manipulation; interpretable scene map; 0.52 / 0.31 sim SR vs Ο0.5 baseline.
- Cons: Scene-memory build-up is task-specific; heavy pipeline.
- arXiv: 2511.06202
- Mechanism: 97%-compressed embeddings from OpenVLA's frozen vision backbone; retrieve top-k by cosine similarity at adaptation time; Thresholded Hybrid Contrastive Loss uses failures as signal.
- Representation: Compressed visual-embedding buffer (retrieval store).
- Pros: On-device adaptation in 31 s from 12 demos; combats catastrophic forgetting.
- Cons: Adaptation-time memory, not inference-time β doesn't directly help long horizons in a single rollout.
- arXiv: 2603.01465
- Mechanism: Discriminative embedding space + auto keyframe selection at state transitions + progress-aware retrieval; keyframes interleaved as visual tokens in VLA input.
- Representation: Sparse semantic-history keyframes as extra visual tokens.
- Pros: 92% SR on memory-intensive tasks vs 57% baseline; explicit non-Markovian design.
- Cons: Keyframe selector generalization is the weak link; minute-scale not tested.
- arXiv: 2412.10345
- Mechanism: Point-tracked past trajectory overlaid on the current observation image as a pixel-space visual prompt.
- Representation: Non-token β past trajectory encoded visually.
- Pros: +10% SimplerEnv, +3.5Γ real vs OpenVLA; almost free at inference.
- Cons: 2D trace is limited; doesn't capture semantic events; short-horizon.
- arXiv: 2506.00411
- Mechanism: Single VLM backbone jointly generates language subtask tokens and action tokens; hierarchical closed-loop control corrects both planning and control errors.
- Representation: Interleaved language subtask tokens serve as planning memory substrate.
- Pros: Unified training objective; subtask tokens double as memory; closed-loop correction.
- Cons: "Long-horizon" here is still Ravens-scale; 20 tasks; no explicit past-observation store.
- arXiv: 2511.16166
- Mechanism: Selective context retention + gated fusion over past stages to stabilize intrinsic reward shaping during RL rollouts; combats "stage hallucination."
- Pros: Useful for long RL rollouts where vanilla policy-gradient variance blows up.
- Cons: RL-loop-specific; limited real-robot eval.
- arXiv: 2508.01415
- Mechanism: Four memories β spatial (knowledge graph), temporal (sequential), episodic (interactions), semantic (concepts) β orchestrated by a planner + critic. Prompt-level agentic (Qwen2.5-VL-72B).
- Pros: +26.5% on EmbodiedBench; beats Claude-3.5-Sonnet on embodied tasks; lifelong-learning story.
- Cons: Agentic pipeline, not end-to-end VLA; latency.
- arXiv: 2510.10125
- Mechanism: Generative world model with pose-conditioned memory retrieval keeps 20 s+ spatial/temporal consistency; trained on ~95 k trajectories.
- Representation: Implicit memory in world-model features, keyed by pose.
- Pros: Long-horizon consistency for planning / imagined rollouts.
- Cons: It's a world model, not a VLA β memory serves prediction, not action directly. Use in tandem with a policy.
- Summary page: Ctrl-World
- Mechanism: Pass past frames through a dedicated video/temporal encoder β fixed-size latent tokens prepended to the VLM.
- Exemplars: MEM (short-term side), ContextVLA, CronusVLA, TraceVLA (pixel-space variant).
- Pros: β Cheap at inference Β· β backbone unchanged Β· β strong on seconds-to-minutes.
- Cons: β Hard capacity ceiling Β· β no explicit event retention Β· β opaque.
- Pick when: You want a few seconds of history and don't want to retrain the VLM.
- Mechanism: Maintain a small set of learnable tokens/queries that aggregate past evidence; policy cross-attends.
- Exemplars: HAMLET (moment tokens), ReMem-VLA (dual-level recurrent), SAM2Act+ (spatial bank), MemoryVLA (perceptual bank side).
- Pros: β Plug-and-play onto frozen VLAs (HAMLET is canonical) Β· β bounded state Β· β good middle ground.
- Cons: β Token count caps capacity Β· β opaque latents Β· β no auditability.
- Pick when: You want history-awareness without touching the base VLA.
- Mechanism: VLM language head emits text-valued propositions about past state / task progress; stored and re-injected.
- Exemplars: MEM (long-term text), MemoryVLA (cognitive bank β partial), RoboMemory (semantic memory layer), EchoVLA (scene graph).
- Pros: β Debuggable ("what did the policy know?") Β· β natural LLM-tooling fit Β· β robust across horizons.
- Cons: β Loses non-verbal cues (contact, force, texture) Β· β language-head hallucination risk Β· β extra VLM calls.
- Pick when: Minute-scale tasks and you need introspection or auditability.
- Mechanism: External store of past embeddings / keyframes / demos; retrieve at inference or adaptation time.
- Exemplars: MemER, MAP-VLA, ExpReS-VLA, Keyframe-Chaining VLA, EchoVLA (episodic traces). Navigation-side precedent: ReMEmbR (arXiv 2409.13682, NVIDIA/USC) β VILA captions in a Milvus vector DB with LLM-agent retrieval over a robot's spatio-temporal history (NaVQA benchmark); retrieval-memory done right, but for nav-QA rather than manipulation control.
- Pros: β Unbounded memory Β· β supports lifelong learning Β· β amortizes fine-tuning via retrieval.
- Cons: β Retrieval-quality bottleneck Β· β two-stage inference Β· β index management at scale.
- Pick when: Long-horizon or distribution shifts over deployment β anywhere an LLM would call for RAG.
- Mechanism: Concatenate more past frames into the VLM context. No compression, no retrieval.
- Exemplars: Naive multi-frame OpenVLA baselines; pre-MEM Ο-series.
- Pros: β Zero new machinery Β· β benefits from any long-context VLM improvement.
- Cons: β Quadratic attention cost Β· β causal confusion Β· β memorization of specific histories.
- Pick when: Baselining or very short histories.
- Mechanism: Generate language or latent subgoal tokens; use them as the memory substrate bridging steps.
- Exemplars: LoHoVLA (subtask + action in one stream), UniVLA (latent action tokens), Ο0.5 / Ο0.6 / Ο0.7 (subtask text + subgoal images), EvoVLA (gated stages). Long-VLA lives nearby via phase masking.
- Pros: β Memory is actionable (the plan IS the memory) Β· β plays well with LLM-style reasoning Β· β supports error correction.
- Cons: β Only as good as the subgoal abstraction Β· β doesn't remember fine-grained observations Β· β brittle if subgoals are wrong.
- Pick when: Genuinely long-horizon tasks with clear sub-structure (cooking, laundry, assembly).
| Axis | Best in class |
|---|---|
| Cheapest plug-in onto a frozen VLA | HAMLET (B), MAP-VLA (D) |
| Most interpretable / debuggable | MEM long-term text (C), keyframe-based (D) |
| Longest real horizon today | MEM β 15 min (A+C hybrid); MemER minute-scale via (D) |
| Best for spatial memory specifically | SAM2Act+ (B), EchoVLA scene map (D) |
| Best for on-device adaptation / lifelong | ExpReS-VLA (D) |
| Best for hierarchical compositional tasks | Ο0.7 (F + C), LoHoVLA (F), UniVLA (F) |
| Headline delta vs. stateless base | HAMLET: 29.2 β 76.4% on history-dep tasks (GR00T N1.5) |
-
From "no memory" (2023) β "fixed context window" (2024) β "architected memory" (2026). OpenVLA and first-gen GR00T ran on single-frame observations. By late 2025 every serious paper has at least a history module; by 2026 "what's your memory design?" is a standard section. HAMLET's 29.2% β 76.4% delta on the same GR00T N1.5 base shows how much performance was being left on the table.
-
Hybrid multi-scale is winning over single-scale. MEM (video + text), MemoryVLA (perceptual + cognitive), ReMem-VLA (frame + chunk), EchoVLA (scene + episodic), RoboMemory (four types). The field converged on "short-term dense + long-term symbolic" after trying each separately.
-
Plug-and-play / frozen-backbone memory modules are ascendant. HAMLET, MAP-VLA, MemoryVLA (mostly), ExpReS-VLA. The economics favor adding memory as a cheap adapter on top of an expensive pretrained VLA β mirroring the LLM field's LoRA-over-full-fine-tune move.
-
Retrieval is eating robotic memory, just as it ate LLM context. MemER, MAP-VLA, ExpReS-VLA, Keyframe-Chaining, EchoVLA β all use retrieval. The VLA world is reinventing RAG with embodied keys (poses, keyframes, trajectories).
-
Prompt expansion as the integration surface. Ο0.7 is the clearest instance: instead of new memory hardware, dump subtask text, subgoal images, episode metadata, and MEM summaries into the prompt. Classifier-free guidance does the rest. Expect 2026β2027 to push this harder β multimodal prompt = multimodal memory.
The ICRA 2026 cohort is striking for where it puts memory: not in the architecture-and-benchmark frame of the ML venues, but squarely in a "memory as a cheap adapter on a frozen, deployed VLA" posture (see the ICRA 2026 Survey, which files these under a dedicated "Memory & retrieval" cluster). All three of its memory contributions are retrieval-flavored, which is the strongest single-venue evidence yet for trend #4.
- MAP-VLA (Cat D, already a row above) is the canonical instance: a library of stage-level soft prompts built offline from demos, retrieved by ββ trajectory-similarity over a sliding window restricted to neighboring stages, and injected by element-wise addition onto the base prompt of a frozen Ο0. The economics are explicitly RAG-over-frozen-model + LoRA-over-full-finetune; the tell is that the real-robot gain (+25%) dwarfs the sim gain (+7%), i.e. stage memory matters most exactly where long-horizon distribution shift bites.
- ExpReS-VLA (Cat D, already a row) reinforces the same retrieval economics on the adaptation axis rather than the inference axis β replay + cosine-similarity retrieval over compressed embeddings to specialize a generalist (OpenVLA) to a fixed deployment without full retraining.
- TrackVLA++ (2510.07134) extends the pattern beyond manipulation: a Target-Identification Memory (paired with a Polar-CoT spatial reasoning token) lets an embodied visual-tracking VLA re-identify its target across occlusion and distractors β the navigation/tracking analogue of the re-identification problem in Β§1, and another vote for explicit per-target episodic state over a wider context window.
Read together, ICRA 2026 says the deployment community has already picked a side: lightweight retrieval/prompting over a frozen backbone beats retraining, and retrieval keys are diversifying (trajectory windows, compressed embeddings, tracked-target identity) exactly as predicted by the "embodied RAG" thesis.
- Selective write is unsolved. Most papers either "retain everything compressed" or "learn a selector." Robot memory actually needs attention at encoding time (what's worth remembering?), not just at retrieval time. HAMLET's time-contrastive init and MemER's keyframe nomination are first steps; neither is robust under distribution shift.
- Minute-scale β hour-scale. MEM reaches 15 min. Nobody shows hour-scale continuous tasks yet (full meal prep end-to-end). The memory compression curve beyond 15 min is unknown.
- Episodic vs. parametric β undecided. ExpReS-VLA shows retrieval + frozen backbone can adapt in 31 s from 12 demos β striking. But Ο*0.6/RECAP shows parametric wins with enough data. The answer is almost certainly episodic for tail distributions, parametric for the head β same story as LLMs.
- VLA memory β LLM RAG convergence. Directionally yes (MemER β RAG for robots), but the modality differs β robot memory needs action-outcomes and spatial state, not just text. Expect convergence on a shared interface (retrieval + structured re-injection) with robotics-specific keys (pose, proprio, contact).
- Interpretability vs. performance. Text-valued memory (MEM long-term, RoboMemory semantic) is debuggable but lossy. Latent memory (moment tokens, recurrent queries) is performant but opaque. A standardized "memory probe" that dumps a policy's state as natural language would unblock safety work β nobody has this yet.
- No standard benchmark. SAM2Act has MemoryBench; Long-VLA has L-CALVIN; MemER has a real-robot suite; HAMLET has history-dependent tasks. Cross-paper comparison is near impossible. A canonical "VLA-MemoryBench" covering minute-scale, occlusion, spatial recall, and episodic transfer would be extremely useful.
- World-model memory vs. control memory. Ctrl-World has pose-conditioned memory retrieval for imagination. Whether imagination-memory and control-memory should share a substrate (cf. Ο0.7's subgoal images from BAGEL-14B) is an open design question.
- Memory + RL is underexplored. EvoVLA is the main attempt. If memory reduces long-rollout variance, RL post-training (Ο*0.6 / RECAP style) could scale further.
If you're shipping a VLA and wondering where to spend a week:
- No history module yet? β Start with HAMLET-style moment tokens. Plug-and-play, frozen base, largest per-effort delta.
- Tasks longer than ~30 s? β Add MEM-style video + text memory or MemER-style keyframe retrieval. Don't stack raw frames.
- Need debuggability / safety probes? β Include a text-valued long-term memory (MEM-style) even if you also keep latent banks.
- Distribution shifts over deployment (new homes, new users)? β Add a retrieval store (MAP-VLA or ExpReS-VLA) rather than fine-tune in place.
- Task has obvious hierarchical structure (cooking, laundry)? β Use planning-token / subgoal memory (LoHoVLA-style or Ο0.7-style prompt expansion).
- Contact-rich manipulation? β Remember that all memory modules above are visual. Torque / tactile memory is still an open gap (see TA-VLA for torque conditioning but not memory).
Individual arXiv links are inline in the table (Β§3). Wiki cross-links:
- Ο0.6 Β· Ο0.7 Β· Ο series evolution
- MemoryVLA Β· HAMLET Β· Ctrl-World
- Review-pi07 (long-form Ο0.7 review β MEM integration discussed there)
- Survey: VLA & Manipulation (ICLR 2026)
- Survey: VLA & Manipulation (CoRL 2025)
Verdict: the frame shifted from bigger in-policy memory to selective history and agentic externalized memory β with navigation, not manipulation, holding the working demo.
RSS 2026 moved the frame twice: (1) selective history beats full context β key-history-frame retrieval (RSS #201) and memory-retrieval studies (RSS #10) outperform naively long windows, and RobotManip's stochastic context sampling shows why (recency shortcuts collapse the mechanism); (2) agentic decomposition arrived β Qwen-RobotNav's planner/executor split with a persistent evidence notebook (survives context compression, auditable belief revisions) set EQA SOTA with 77% fewer steps.
| Approach | Pros | Cons |
|---|---|---|
| In-policy memory modules (the 6 architectures above) | End-to-end, fast | Context cost; recency-shortcut risk |
| Retrieval over episodic history | Scales with episode length | Bounded by retrieval quality |
| Agentic notebook memory (planner + executor) | Auditable, compressible, composable | Planner latency; interface design becomes the research object |
| Hierarchical/spatial memory (ICML 2026) | HiMe: Executor/Sentry/Planner tiers with add/update/delete knowledge ops; SOMA: persistent spatial-semantic memory enables out-of-view manipulation | Young; benchmark coverage thin |
| Structured vs parametric memory (IROS 2026 π) | dedicated "Building Memory into VL Robot Agents" session; GaussMemory (3D-Gaussian scene memory) & PROMPT (probabilistic memory trees) = structured; TempoFit (layer-wise temporal KV memory for long-horizon VLA) & RoboSSM (state-space) = parametric/context β the field bifurcates | Structured: build/maintain cost; parametric: opaque, bounded by context mechanism |
- Long-horizon manipulation autonomy record is 11 stages / 2.5 min (ViTacFormer) β minutes, not hours.
- Drift and mid-task recovery are barely measured (RoboMME, now an ICML 2026 Oral, is the main probe; CAPS reframes long-horizon instruction drift as a correctable sampling error β training-free SNR-triggered MCMC).
- No manipulation-side equivalent of RobotNav's agentic demo exists yet.
β Back to Home