RSS 2026 Memory Retrieval in Visuomotor Policies - Heungwoo/research GitHub Wiki
Memory Retrieval in Visuomotor Policies for Long-Horizon Robot Control
Venue: RSS 2026 (Sydney, Jul 13–17) · Session: World Models & Memory · paper #10 Authors: Rutav Shah, Yisu Li, Femi Bello, Yuke Zhu, Roberto Martín-Martín arXiv: 2606.25136 · program page
Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Figure 1 shows the motivating long-horizon household task: a mobile manipulator is told to "Transfer ALL the chip packets to kitchen drawer and close it." The six panels trace the rollout — pick up a chip packet, place it in a drawer, navigate to the living room, pick up another packet, place it, close the drawer — and at the decision point ("Close the drawer or go back to living room?") the policy must retrieve from memory how many packets remain, information no longer visible in the current observation.
Problem
Household robots operating under partial observability must recall diverse past information — object locations, object relations, numerical counts, event times — to complete long-horizon tasks. Hand-designed or heuristic memory retrieval relies on task-specific assumptions that do not generalize, while naively adding attention-based (long-context transformer) retrieval to imitation learning creates two failure modes: spurious correlations between past observations and actions, and compounding prediction errors in memory that cause model drift and cascading failures in closed-loop control.
Method
HALO (History-Aware visuomotor policy for LOng-horizon imitation learning) equips a transformer visuomotor policy with a learned attention-based retrieval module over encoded history M_t of past observations and actions, plus dual heads for actions and text. Two components tame the retrieval. (1) VLM-prior distillation via video question answering: a multi-stage pipeline (trajectory-to-text summarization with grounded vision models, LLM question/answer synthesis anchored to randomly sampled frame indices, LLM verification that filters out about 20% of pairs) generates task-relevant VQA pairs, and the policy is co-trained with loss L = L_IL + λ·L_VQA so retrieval attends to task-relevant history. (2) Top-k attention sparsification: retrieval is restricted to the k most relevant memory entries (straight-through estimator for gradients), limiting the influence of accumulated errors. The arXiv version reports retrieval over up to eight minutes of past experience (the conference program abstract cites two minutes).
Results
On the ReMemBench simulation benchmark (50 rollouts/task; spatial, relational, numerical and event-time tasks), HALO averages 0.41 success vs 0.22 for a standard transformer, 0.29 for hand-designed features (MemER-style), 0.20 for SAM2Act++, 0.18 for ReMemBer, 0.34 for Scene Memory Transformer, and 0.29 for Token Merging — the paper summarizes this as +7% absolute over diverse tasks, beating hand-designed rules (−12%) and task-specific features (−21%). Policies using VLM priors alone reach only 18% absolute success vs 41% for HALO; top-k sparsification alone contributes +9% absolute. On five real-world tasks across a fixed-base and a mobile manipulator (20 rollouts/task), HALO averages 0.55 vs 0.36 for the standard transformer.
Significance
A clean demonstration that learned, data-driven memory retrieval — steered by VLM-generated VQA supervision and stabilized by sparse attention — beats both hand-crafted memory rules and compressed-history baselines for long-horizon control. Directly relevant to the wiki's Review-VLA-Memory and Review-World-Models threads on how policies should store and access history.
← Back to RSS 2026 survey · RSS-2026-Papers · Home