ICML 2026 HiMe - Heungwoo/research GitHub Wiki
HiMe: Hierarchical Embodied Memory for Long-Horizon Vision-Language-Action Control — Decoupling fast execution, working memory, and slow planning
Venue: ICML 2026 (Poster) Category: VLA Architecture Affiliations: Li Ji, Siyin Wang, Pengfang Qian, Xiaopeng Yu, Yihai Tian, Zhaoye Fei, Jingjing Gong, Xipeng Qiu
Problem
Current Vision-Language-Action (VLA) models excel at robotic manipulation but often struggle with non-Markovian tasks — those requiring long-term memory and reasoning — because they rely almost entirely on immediate observations. The authors identify a core tension underlying this gap: high-performance reasoning systems are computationally slow, while faster alternatives lack reasoning depth. A long-horizon controller therefore needs both real-time reactivity and the ability to remember and reason over extended histories, two requirements that pull system design in opposite directions.
Method
HiMe is a hierarchical embodied memory framework that decouples control into three components operating at different timescales:
- Executor — a high-frequency module responsible for low-level execution / real-time action.
- Sentry — maintains working memory, bridging immediate observation and longer-term strategy.
- Planner — handles long-term strategy and slow, deliberate reasoning.
Underpinning these is a dynamic knowledge system built on cross-modal semantic schemas with active management mechanisms. The system supports explicit Add, Update, and Delete operations over its stored knowledge, so the robot can keep its memory adaptive and current as a task unfolds rather than carrying stale or irrelevant context.
This hierarchical design is what lets HiMe balance the conflict between real-time execution and slow-thinking planning: the fast Executor never has to wait on the slow Planner, while the Sentry's working memory and the active knowledge system keep the two layers coherent.
flowchart TB
P[Planner: long-term strategy / slow thinking] --> S[Sentry: working memory]
S --> E[Executor: high-frequency execution]
KB[(Dynamic knowledge system\ncross-modal semantic schemas)]
P <--> KB
S <--> KB
KB -. Add / Update / Delete .-> KB
O[Observations] --> E
E --> A[Actions]
Results
Experiments show HiMe improves over traditional / flat memory approaches on long-horizon, memory-dependent tasks. Notably, the framework demonstrates a novel ability to self-correct its internal knowledge based on human preferences — refining what it has stored in response to human feedback. (The ICML abstract page does not list quantitative benchmark numbers; these will be added once the full paper is available.)
Significance
HiMe targets a known weakness of VLA models — their Markovian, observation-bound nature — by giving manipulation policies an explicit, editable memory hierarchy. Separating a fast Executor from a slow Planner via a Sentry working-memory layer, and adding a knowledge store that can be actively managed and corrected from human feedback, offers a concrete recipe for long-horizon, non-Markovian embodied control.
Links
- ICML 2026: https://icml.cc/virtual/2026/poster/60897
← Back to ICML-2026