ICLR 2026 Memory Experience Retrieval - Heungwoo/research GitHub Wiki

MemER โ€” scaling robot memory via experience retrieval

Venue: ICLR 2026 ยท Authors: Ajay Sridhar, Jennifer Pan, Satvik Sharma, Chelsea Finn ยท arXiv:2510.20328 ยท Category: Data / memory / representation for manipulation ยท Trend tag: Memory / long horizon

Official title: Scaling up Memory for Robotic Control via Experience Retrieval (a.k.a. MemER).

Approach diagram

flowchart LR
  Hist[Long observation history] --> HL[High-level policy<br/>Qwen2.5-VL-7B-Instruct]
  HL --> KF[Select + track<br/>relevant keyframes]
  KF --> Instr[Text instruction<br/>keyframes + recent frames]
  Instr --> LL[Low-level policy<br/>ฯ€โ‚€.โ‚… VLA]
  LL --> Act[Action]
Loading

Problem

Most robot policies lack memory. Naively conditioning on long observation histories is computationally expensive and brittle under covariate shift, while indiscriminately subsampling history yields irrelevant or redundant context. The challenge is reasoning over minutes-long dependencies without drowning the policy in stale frames.

Method

MemER is a hierarchical policy that learns to retrieve experience:

  • The high-level policy is trained to select and track previous task-relevant keyframes from its own experience.
  • It then conditions on those selected keyframes plus the most recent frames to generate text instructions for a low-level policy.
  • The design drops in on top of existing VLA models, enabling efficient long-horizon reasoning.

In experiments the authors fine-tune Qwen2.5-VL-7B-Instruct as the high-level policy and ฯ€โ‚€.โ‚… as the low-level policy, using demonstrations supplemented with minimal language annotations.

Results

MemER outperforms prior methods on three real-world long-horizon manipulation tasks that require minutes of memory. (Per-task success numbers are reported on the project page / paper and omitted here pending confirmation.)

Significance

Instead of widening the context window, MemER learns which past frames matter and retrieves only those โ€” a learned-retrieval view of robot memory that is cheaper and more robust than dense history conditioning, and that layers cleanly onto off-the-shelf VLAs. Contrasts with bank-style latent memory in MemoryVLA and history-aware attention in HAMLET.

Links

Related pages

โ† Back to ICLR-2026

โš ๏ธ **GitHub.com Fallback** โš ๏ธ