ICLR 2026 HAMLET - Heungwoo/research GitHub Wiki

HAMLET โ€” Switch Your VLA into a History-Aware Policy

Venue: ICLR 2026 (under review) ยท arXiv 2510.00695 Authors: Myungkyu Koo, Daewon Choi, Taeyoung Kim, Kyungmin Lee, Changyeon Kim, Younggyo Seo, Jinwoo Shin (KAIST ยท UC Berkeley ยท RLWRLD) Category: VLA Architecture โ€” Memory Trend tag: Memory / long horizon

Approach diagram

flowchart LR
  Obs[Per-step observation] --> VLM[Pretrained VLM backbone<br/>FROZEN]
  VLM --> MT[Moment token<br/>TCL-initialized event summary]
  MT --> Mem[Memory module<br/>causal-attention Transformer]
  Past[Past moment tokens] --> Mem
  Mem --> Cond[History condition]
  VLM --> Cond
  Cond --> Act[Action head] --> A[Action]
Loading

Problem

Most VLAs are trained on tasks short enough to fit the backbone's context window. Long-horizon tasks (fold this laundry pile; assemble this kit) require memory of events from many seconds ago. Naive trajectory concatenation either explodes context or causes the model to memorize specific trajectories.

Method

Plug-in memory framework on top of a pretrained VLA. Two components:

  • Moment tokens โ€” learnable tokens appended to the VLM input that compress the per-step vision-language representation into a compact event summary. They are initialized with time-contrastive learning (TCL): augmented views of a frame act as positives, distant timesteps as hard negatives, so the tokens capture temporally distinctive aspects and filter redundant static content.
  • Memory module โ€” a lightweight Transformer that aggregates past moment tokens via causal self-attention into a temporally-informed condition concatenated with the current VLM representation before action prediction.

The VLM backbone stays frozen; only the moment tokens and the memory module are fine-tuned (a parameter-efficient adapter, not a full retrain โ€” and not zero-training).

Results

  • Real-world history-dependent tasks (on GR00T N1.5): 76.4% average success, +47.2 pp over the baseline โ€” the headline gain, since these tasks genuinely require memory.
  • RoboCasa Kitchen (100-demo): 64.1% โ†’ 66.4%.
  • LIBERO: 95.6% โ†’ 97.7%.
  • SimplerEnv-Bridge (on CogACT): ~52.1% โ†’ ~63.5%, showing the framework generalizes across backbones.

Backbones tested: GR00T N1.5/N1, CogACT, ฯ€โ‚€ / ฯ€โ‚€-FAST, plus an OpenVLA autoregressive extension. Ablations: removing the memory module causes the largest drop; removing TCL initialization hurts consistently; a Transformer memory beats RNN/LSTM/GRU variants.

Significance

Plug-and-play design makes long-horizon capability a cheap add-on, not a reason to retrain the whole model. Introduces a principled distinction between context (raw window) and memory (compressed past-but-relevant info).

Links

๐Ÿ“– In-depth cross-paper review

For a comparison with all major VLA memory architectures (MEM, MemoryVLA, MemER, SAM2Act+, etc.) โ€” categorized into 6 groups with pros/cons: VLA Memory Architectures Review.

Related pages

  • MemoryVLA (perceptual + cognitive memory bank)

โ† Back to ICLR-2026

โš ๏ธ **GitHub.com Fallback** โš ๏ธ