ICRA 2026 MAP VLA - Heungwoo/research GitHub Wiki

MAP-VLA โ€” Memory-Augmented Prompting for VLA

Venue: ICRA 2026 ยท Authors: Runhao Li, Wenkai Guo, Zhenyu Wu, Changyuan Wang, Haoyuan Deng, Zhenyu Weng, Yap-Peng Tan, Ziwei Wang (NTU, BUPT, Tsinghua, SCUT, VinUniversity) ยท arXiv: 2511.09516 Category: VLA memory Trend tag: Memory-augmented prompting

Approach diagram

flowchart LR
  Demo[Historical demonstrations] --> Build[Build memory library<br/>per task-stage units]
  Build --> SP[Learnable soft prompts<br/>prompt tuning, frozen VLA]
  Obs[Live observation + recent trajectory] --> Match[Trajectory similarity matching<br/>โ„“โ‚‚ over sliding window]
  SP --> Match
  Match --> Sel[Retrieve stage-matched<br/>soft prompt ๐’ฑโ‚–]
  Sel --> Inj["Inject: ๐’ซโ‚– = ๐’ซ_base + ๐’ฑโ‚–<br/>(element-wise add)"]
  Inj --> VLA[Frozen VLA<br/>ฯ€โ‚€ flow-matching policy]
  VLA --> Act[Action chunk]
Loading

Problem

Long-horizon manipulation is non-Markovian: a frozen VLA conditioned only on the current observation forgets which sub-stage of a multi-step task it is in, causing stage confusion and drift. Fine-tuning the whole policy to add memory is expensive and risks degrading the pretrained backbone. MAP-VLA asks whether memory can be added as a cheap, plug-and-play module over a frozen VLA.

Method

MAP-VLA builds a memory library from historical demonstrations, where each memory unit captures a specific stage of a task and is encoded as a learnable soft prompt (a vector sequence trained via prompt tuning while the VLA stays frozen). At execution time it performs trajectory similarity matching: the robot's recent state segment is compared against demonstration windows by โ„“โ‚‚ distance with a sliding-window search restricted to neighboring task stages, retrieving the most relevant stage's soft prompt. The retrieved prompt is injected by element-wise addition onto the base prompt (๐’ซโ‚– = ๐’ซ_base + ๐’ฑโ‚–), dynamically steering action generation. The base policy is ฯ€โ‚€ (conditional flow-matching action expert); OpenVLA is also used as a comparison backbone. Because the VLA weights never change, the module is plug-and-play.

Results

  • Simulation (LIBERO-Long): ฯ€โ‚€ baseline โ‰ˆ 76.4% average success; MAP-VLA adds up to +7.0% absolute gain.
  • Real robot (long-horizon tasks): +25.0% absolute improvement over the frozen-VLA baseline.
  • Gains concentrate on long-horizon, multi-stage tasks, consistent with the stage-memory design; the module adds memory without retraining the backbone.

Significance

MAP-VLA is one of the cheapest memory mechanisms for VLAs: it stores experience as stage-level soft prompts and retrieves them by trajectory matching, mirroring the LLM field's RAG-over-frozen-model and LoRA-over-full-finetune economics. The much larger real-world gain (+25%) than sim (+7%) underscores that stage memory matters most under real long-horizon distribution shift. See the cross-paper comparison in the VLA Memory review (it is listed there) and architectural context in the VLA Architectures review.

Links

Related pages

โ† Back to ICRA-2026

โš ๏ธ **GitHub.com Fallback** โš ๏ธ