ICLR 2026 HAMLET - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 (under review) ยท arXiv 2510.00695 Authors: Myungkyu Koo, Daewon Choi, Taeyoung Kim, Kyungmin Lee, Changyeon Kim, Younggyo Seo, Jinwoo Shin (KAIST ยท UC Berkeley ยท RLWRLD) Category: VLA Architecture โ Memory Trend tag: Memory / long horizon
flowchart LR
Obs[Per-step observation] --> VLM[Pretrained VLM backbone<br/>FROZEN]
VLM --> MT[Moment token<br/>TCL-initialized event summary]
MT --> Mem[Memory module<br/>causal-attention Transformer]
Past[Past moment tokens] --> Mem
Mem --> Cond[History condition]
VLM --> Cond
Cond --> Act[Action head] --> A[Action]
Most VLAs are trained on tasks short enough to fit the backbone's context window. Long-horizon tasks (fold this laundry pile; assemble this kit) require memory of events from many seconds ago. Naive trajectory concatenation either explodes context or causes the model to memorize specific trajectories.
Plug-in memory framework on top of a pretrained VLA. Two components:
- Moment tokens โ learnable tokens appended to the VLM input that compress the per-step vision-language representation into a compact event summary. They are initialized with time-contrastive learning (TCL): augmented views of a frame act as positives, distant timesteps as hard negatives, so the tokens capture temporally distinctive aspects and filter redundant static content.
- Memory module โ a lightweight Transformer that aggregates past moment tokens via causal self-attention into a temporally-informed condition concatenated with the current VLM representation before action prediction.
The VLM backbone stays frozen; only the moment tokens and the memory module are fine-tuned (a parameter-efficient adapter, not a full retrain โ and not zero-training).
- Real-world history-dependent tasks (on GR00T N1.5): 76.4% average success, +47.2 pp over the baseline โ the headline gain, since these tasks genuinely require memory.
- RoboCasa Kitchen (100-demo): 64.1% โ 66.4%.
- LIBERO: 95.6% โ 97.7%.
- SimplerEnv-Bridge (on CogACT): ~52.1% โ ~63.5%, showing the framework generalizes across backbones.
Backbones tested: GR00T N1.5/N1, CogACT, ฯโ / ฯโ-FAST, plus an OpenVLA autoregressive extension. Ablations: removing the memory module causes the largest drop; removing TCL initialization hurts consistently; a Transformer memory beats RNN/LSTM/GRU variants.
Plug-and-play design makes long-horizon capability a cheap add-on, not a reason to retrain the whole model. Introduces a principled distinction between context (raw window) and memory (compressed past-but-relevant info).
For a comparison with all major VLA memory architectures (MEM, MemoryVLA, MemER, SAM2Act+, etc.) โ categorized into 6 groups with pros/cons: VLA Memory Architectures Review.
- MemoryVLA (perceptual + cognitive memory bank)
โ Back to ICLR-2026