ICRA 2026 MAP VLA - Heungwoo/research GitHub Wiki
Venue: ICRA 2026 ยท Authors: Runhao Li, Wenkai Guo, Zhenyu Wu, Changyuan Wang, Haoyuan Deng, Zhenyu Weng, Yap-Peng Tan, Ziwei Wang (NTU, BUPT, Tsinghua, SCUT, VinUniversity) ยท arXiv: 2511.09516 Category: VLA memory Trend tag: Memory-augmented prompting
flowchart LR
Demo[Historical demonstrations] --> Build[Build memory library<br/>per task-stage units]
Build --> SP[Learnable soft prompts<br/>prompt tuning, frozen VLA]
Obs[Live observation + recent trajectory] --> Match[Trajectory similarity matching<br/>โโ over sliding window]
SP --> Match
Match --> Sel[Retrieve stage-matched<br/>soft prompt ๐ฑโ]
Sel --> Inj["Inject: ๐ซโ = ๐ซ_base + ๐ฑโ<br/>(element-wise add)"]
Inj --> VLA[Frozen VLA<br/>ฯโ flow-matching policy]
VLA --> Act[Action chunk]
Long-horizon manipulation is non-Markovian: a frozen VLA conditioned only on the current observation forgets which sub-stage of a multi-step task it is in, causing stage confusion and drift. Fine-tuning the whole policy to add memory is expensive and risks degrading the pretrained backbone. MAP-VLA asks whether memory can be added as a cheap, plug-and-play module over a frozen VLA.
MAP-VLA builds a memory library from historical demonstrations, where each memory unit captures a specific stage of a task and is encoded as a learnable soft prompt (a vector sequence trained via prompt tuning while the VLA stays frozen). At execution time it performs trajectory similarity matching: the robot's recent state segment is compared against demonstration windows by โโ distance with a sliding-window search restricted to neighboring task stages, retrieving the most relevant stage's soft prompt. The retrieved prompt is injected by element-wise addition onto the base prompt (๐ซโ = ๐ซ_base + ๐ฑโ), dynamically steering action generation. The base policy is ฯโ (conditional flow-matching action expert); OpenVLA is also used as a comparison backbone. Because the VLA weights never change, the module is plug-and-play.
- Simulation (LIBERO-Long): ฯโ baseline โ 76.4% average success; MAP-VLA adds up to +7.0% absolute gain.
- Real robot (long-horizon tasks): +25.0% absolute improvement over the frozen-VLA baseline.
- Gains concentrate on long-horizon, multi-stage tasks, consistent with the stage-memory design; the module adds memory without retraining the backbone.
MAP-VLA is one of the cheapest memory mechanisms for VLAs: it stores experience as stage-level soft prompts and retrieves them by trajectory matching, mirroring the LLM field's RAG-over-frozen-model and LoRA-over-full-finetune economics. The much larger real-world gain (+25%) than sim (+7%) underscores that stage memory matters most under real long-horizon distribution shift. See the cross-paper comparison in the VLA Memory review (it is listed there) and architectural context in the VLA Architectures review.
- arXiv:2511.09516
- ICRA 2026 listing
โ Back to ICRA-2026