Review RoboMME - Heungwoo/research GitHub Wiki
Venue: ICML 2026 (Oral) Β· arXiv: 2603.04639 Β· Code: RoboMME/robomme_benchmark Affiliations: University of Michigan Β· Stanford Β· Figure AI Traction (2026-06-09): β 111 GitHub Β· 6 citations (arXiv preprint) One-page version: ICML-2026-RoboMME Β· See also VLA Memory review
RoboMME is the first benchmark that turns "which memory mechanism should a VLA use?" into a controlled, apples-to-apples comparison. It contributes (1) a cognitively-motivated task taxonomy (temporal / spatial / object / procedural memory) over 16 manipulation tasks, and (2) the MME-VLA suite β 14 memory-augmented policies all built on the same Ο0.5 backbone, factored as 3 memory representations Γ instantiations Γ 3 integration mechanisms. The headline empirical result: no single memory design wins everywhere; the best overall non-oracle policy, FrameSamp + Modulator (perceptual memory, 44.51 % avg), beats all symbolic and recurrent variants and the prior SOTA, but symbolic memory still wins on counting / event-salient tasks. Humans score 90.5 %, so headroom is large.

Generalist VLAs (Ο0.5, OpenVLA, etc.) are largely Markovian β they condition on the current observation only. Real manipulation is history-dependent: counting repeated actions, re-finding an object that was occluded, imitating a demonstrated motion strategy. Memory-augmented VLAs have appeared (MemoryVLA, HAMLET, SAM2Act, MemERβ¦), but each is evaluated on its own narrow, non-standardized setup, so the community cannot say which memory representation helps which kind of task. RoboMME fixes the evaluation substrate and then sweeps the design space on a single backbone.
| Benchmark | Non-Markov | Partial Obs | Dynamic | Video | Memory types | #Tasks | #Demos | Avg #Steps |
|---|---|---|---|---|---|---|---|---|
| RLBench18 / CALVIN / LIBERO / VLABench | β | β | β | β | T | 18β130 | β | 115β584 |
| RoboCerebra | β | β | β | β | T | 100 | 1,000 | 2,972 |
| MemoryBench (SAM2Act) | β | β | β | β | T+S | 3 | 300 | 312 |
| MIKASA-robo | β | β | β | β | T+S+O | 12 | 1,250 | 72 |
| RoboMME (ours) | β | β | β | β | T+S+O+P | 16 | 1,600 | 481 |
RoboMME is the only suite that is simultaneously non-Markovian, partially observable, dynamic, video-conditioned, and covers all four memory types with subgoal + keyframe annotations.
16 tasks in 4 suites, each targeting one primary memory faculty (T = temporal, S = spatial, O = object, P = procedural):
| Suite | Memory | Tasks | What it stresses |
|---|---|---|---|
| Counting | T | PickXTimes, BinFill, SwingXTimes, StopCube | counting repetitions over time; StopCube is time-critical |
| Permanence | S | VideoUnmask, ButtonUnmask, VideoUnmaskSwap, ButtonUnmaskSwap | object permanence under occlusion / masking; Swap adds tracking through container shuffles |
| Reference | O (+T) | PickHighlight, VideoRepick, VideoPlaceButton, VideoPlaceOrder | resolving visual / action / language references to a specific object or target seen earlier in a video |
| Imitation | P (+O) | MoveCube, InsertPeg, PatternLock, RouteStick | replicating a demonstrated motion strategy (pick-place vs push vs hook; linear vs circular trajectories) |
Tasks range from 208 to 1,134 average timesteps (VideoPlaceOrder is the longest, a long-video ordinal-reference task), so the suite spans short reactive tasks to very long-horizon video-conditioned ones.
This is the core technical contribution: a factored design space on a frozen-recipe Ο0.5 backbone, so differences are attributable to the memory design rather than the base model.

(A) Symbolic memory β history is summarized as interpretable language subgoals. At each step an auxiliary VLM conditions on the current image + prior subgoals and emits the next subgoal. Two instantiations:
-
SimpleSG β plain instructions (
pick up the green cube). -
GroundSG β grounded subgoals that append image coordinates (
pick up the green cube at [63, 152]), which helps spatial reasoning. - Subgoal generators: Gemini-2.5-Pro (prompt-only), a fine-tuned Qwen3-VL-4B (QwenVL) trained on the dataset's subgoal annotations, or Oracle (simulator ground truth, used as an upper bound). Symbolic memory requires extra subgoal annotations to train the VLM.
(B) Perceptual memory β history is a sequence of visual tokens from past images, extracted by the Ο0.5 vision encoder (no proprioception needed). Two token-selection strategies:
- TokenDrop β token dropping (Γ la TimeChat): removes temporally redundant patches by RGB difference.
- FrameSamp β uniform frame sampling (Γ la VideoLLM): evenly downsamples frames and concatenates their tokens.
(C) Recurrent memory β compress the visual-token sequence into a fixed-size latent state via recurrence. Two models:
- TTT β Test-Time Training: updates fast weights online with a self-supervised loss, then reads out features.
- RMT β Recurrent Memory Transformer: processes the stream in segments and recurrently updates learnable memory slots with a transformer.
Symbolic memory integrates trivially (concatenate subgoals onto the task instruction). The neural memory tokens (perceptual / recurrent) need architectural plumbing, so the paper studies three ways to inject them into Ο0.5's VLM + action-expert structure:
| Mechanism | How memory enters the policy |
|---|---|
| Memory-as-Context | memory tokens are concatenated with the original inputs and jointly processed by the VLM expert β directly shifts VLM features. |
| Memory-as-Modulator | uses AdaLN: action features cross-attend to memory tokens, and the result is projected into scale & shift parameters that modulate the normalized action features before each feed-forward block. Lightweight, feature-wise conditioning that preserves the base policy. |
| Memory-as-Expert | a dedicated lightweight memory expert processes memory tokens on its own pathway, with a blockwise causal attention scheme β the action expert attends to both the VLM and memory experts, but VLM and memory experts do not attend to each other (limits interference, preserves VLM behavior). |
| Representation | Instantiations Γ Integration | # policies |
|---|---|---|
| Symbolic | SimpleSG, GroundSG (concat) | 2 |
| Perceptual | {TokenDrop, FrameSamp} Γ {Context, Modul, Expert} | 6 |
| Recurrent | {TTT, RMT} Γ {Context, Modul, Expert} | 6 |
| Total | 14 |
Prior methods compared (4): (1) Ο0.5 no-memory; (2) Ο0.5 w/ past actions (concatenates past actions as explicit symbolic memory, UniVLA-style); (3) SAM2Act+ (SAM2 memory bank + discrete keyframe actions executed by an external motion planner); (4) MemER (VLM infers symbolic subgoals from a buffer of selected keyframes β a hybrid of perceptual + symbolic memory).
Evaluation protocol: fixed memory budget of 512 tokens (matching the current-observation image-token count) for fair comparison; 50 episodes/task (800 total) with held-out seeds; single multi-task model, β€1,300 steps/episode; scores averaged over the last 3 checkpoints Γ 3 seeds (9 runs).
Grounded in the paper's Appendix A (math) and the official code in RoboMME/robomme_policy_learning β the model suite lives there (an openpi/Ο0.5 fork), while the benchmark env is in robomme_benchmark. Weights: Yinpei/mme_vla_suite.
Ο0.5 is a Mixture-of-Transformers (MoT) dual-expert model. In code, HistoryPi0 subclasses openpi's Pi0 (src/mme_vla_suite/models/integration/history_pi0.py) and runs two Gemma experts under one blockwise-causal attention:
-
VLM expert (PaliGemma = SigLIP + Gemma-2B) β a SigLIP image encoder (
PaliGemma.img) turns each camera view into vision tokens; these plus the tokenized language prompt form the prefix. This is the "understand the scene" expert (paligemma_variant = "gemma_2b"). -
Action expert (Gemma-300M) β takes the noised action chunk + a flow-matching timestep and forms the suffix; it predicts the flow-matching velocity that denoises the action chunk. In Ο0.5 it is conditioned on the diffusion timestep via AdaRMSNorm (
use_adarms=[False, True],adarms_cond = time_emb).
Attention is blockwise-causal: action (suffix) tokens attend to the VLM prefix; the mask is built so vision tokens never attend to memory tokens, but action tokens can (make_attn_mask, mask_na). The training loss is the standard flow-matching objective on the action expert's predicted velocity.
From a streaming horizon of 16 past frames (streaming_obs_horizon: 16, single view), FrameSamp evenly samples frames (S_t = EvenSample({0β¦t}; N), Appendix Eq. 14) and mean-pools each to 16 tokens (token_per_image: 16, pool_type: mean), under a 512-token budget. The mem_encoder builds each memory token from the SigLIP image embedding (2048-d) + a positional embedding (768-d) β and in the best config no proprioceptive state (use_state_emb: false, confirming the paper's "visual tokens alone suffice") β then projects to memory_token_dim: 1024, yielding the memory tokens M_t (obs.static_image_emb / static_pos_emb in embed_memory). (Config: models/config/robomme/perceptual-framesamp-modul.yaml. TokenDrop instead keeps the least-redundant patches by RGB/Jaccard similarity; recurrent variants feed recur_* streams through a TTT or RMT encoder. memory_token_dim is 2048 for context-integration but 1024 for modulation/expert.)
Crucially, in modulation mode the memory tokens do not enter the token sequence (that is the context variant, which prepends M_t to the prefix). Instead memory conditions the AE through feature-wise AdaLN, implemented in history_gemma.py and gated by mem_mods=[False, True] so that only the action expert is modulated β the VLM expert is left untouched, preserving its pretrained behaviour. The injection happens inside HistoryBlock β i.e. in every transformer layer β right before the FFN sublayer, and only on the action-expert stream (the code guard is if i == len(xs)-1 and integration_type=="modulation", where the loop index i runs over experts, so len(xs)-1 selects the action expert, not a particular layer). At each such layer (Appendix Eqs. 21β23):
-
Cross-attention, action β memory (
MemoryAttention): query = current action featuressΜβ, key/value = memory tokensMβ, with a softmax masked over valid memory slots βrβ = Attn_mod(Q = sΜβ, K = Mβ, V = Mβ). -
Produce scale & shift (
MemoryRMSNormβ a singleDense(2Β·width)onrβ):(Ξ³β, Ξ²β) = MLP_mod(rβ). -
AdaLN-modulate the action features:
Εβ = Ξ³β β Norm(sΜβ) + Ξ²β(code:normed * (1 + scale) + shift, with near-zero init so the block starts as identity).
So the action expert is conditioned on both signals simultaneously: Ο0.5's existing time-AdaRMSNorm (diffusion timestep) and the new memory-AdaLN (history). Because the conditioning is a lightweight per-feature scale/shift rather than extra sequence tokens, it barely perturbs the base policy β which is exactly why the paper finds modulation the most robust integration and why TokenDrop's lost global context hurts less here than under context integration.
flowchart LR
subgraph IN[Inputs]
OBS["current obs<br/>front + wrist RGB"]
LANG["language goal"]
HIST["frame history"]
end
HIST -->|"FrameSamp: 16-frame horizon,<br/>mean-pool 16 tok/img, β€512 budget"| MENC["mem_encoder<br/>(SigLIP 2048-d + pos 768-d,<br/>no state) β 1024-d"]
MENC --> MT["memory tokens Mβ"]
OBS --> SIG["SigLIP encoder"]
SIG --> VLM["VLM expert<br/>(PaliGemma: SigLIP + Gemma-2B)"]
LANG --> VLM
VLM -->|"prefix KV (blockwise-causal)"| AE
NOISE["noised action chunk + flow time Ο"] --> AE["Action expert (Gemma-300M)"]
TIME["time_emb"] -->|"AdaRMSNorm"| AE
MT -->|"cross-attn β Ξ³,Ξ² (AdaLN)"| AE
AE --> V["flow velocity vΞΈ β denoise β action chunk aβ"]
| Field | Value | Meaning |
|---|---|---|
representation_type |
perceptual |
visual-token memory |
perceptual_memory.type |
frame_sampling |
even frame sampling (vs token_dropping) |
integration_type |
modulation |
AdaLN on the action expert |
streaming_obs_horizon |
16 |
past frames fed to the memory encoder |
token_per_image |
16 |
mean-pooled tokens kept per frame (vs 64 default) |
pool_type |
mean |
pooling over patch tokens |
budget |
512 |
total memory-token cap |
num_views |
1 |
single (fixed) camera view for memory |
use_pos_emb / use_state_emb
|
true / false
|
+ positional emb, no proprio state |
memory_feature.img.input_dim |
2048 |
SigLIP image-embedding dim from Ο0.5 |
memory_token_dim |
1024 |
memory-token width (2048 for context) |
| VLM / action / memory expert |
gemma_2b / gemma_300m / gemma_150m
|
gemma_150m only for the expert variant |
Key code pointers: models/integration/history_pi0.py (HistoryPi0, embed_memory, embed_prefix/suffix, compute_loss routing on integration_type β {context, modulation, expert}) Β· models/integration/history_gemma.py (HistoryBlock, MemoryAttention, MemoryRMSNorm, applying mem_attn + AdaLN before the FFN of every layer on the action-expert stream) Β· src/openpi/models/pi0.py (base MoT Ο0.5). The expert variant instead adds a third Gemma-150M memory expert with the blockwise rule "action expert attends to VLM + memory experts; VLM and memory do not attend each other".
| Representation | Policy | Integ./VLM | Counting | Permanence | Reference | Imitation | AVG |
|---|---|---|---|---|---|---|---|
| No Memory | Ο0.5 | β | 28.78 | 17.00 | 17.17 | 8.78 | 17.93 |
| Symbolic | SimpleSG | Oracle | 82.56 | 21.56 | 32.28 | 61.94 | 49.58 |
| Symbolic | GroundSG | Oracle | 83.86 | 93.31 | 95.17 | 63.98 | 84.08 |
| Symbolic | SimpleSG | QwenVL | 44.61 | 19.61 | 25.22 | 26.56 | 29.00 |
| Symbolic | GroundSG | QwenVL | 38.00 | 39.34 | 31.56 | 21.89 | 32.70 |
| Perceptual | TokenDrop | Modul | 52.33 | 26.83 | 34.72 | 38.28 | 38.04 |
| Perceptual | FrameSamp | Modul | 65.22 | 25.11 | 36.33 | 51.39 | π 44.51 |
| Perceptual | FrameSamp | Expert | 66.78 | 25.17 | 24.11 | 28.94 | 36.25 |
| Recurrent | TTT | Expert | 36.00 | 22.95 | 19.67 | 10.78 | 22.35 |
| Recurrent | RMT | Modul | 34.14 | 15.78 | 21.11 | 9.67 | 20.17 |
| Prior | SAM2Act+ | β | 35.33 | 26.00 | 16.83 | 7.33 | 21.37 |
| Prior | MemER | β | 48.83 | 53.17 | 38.00 | 29.50 | 42.38 |
(Oracle rows are upper bounds β they use simulator-ground-truth subgoals. π = best non-oracle.)
Reading the table:
- Memory helps a lot: the best learned policy (44.51 %) is ~2.5Γ the no-memory Ο0.5 baseline (17.93 %).
- Perceptual > Symbolic > Recurrent on average (for learned, non-oracle settings). FrameSamp + Modul (44.51 %) is the best learned policy and edges out the strongest prior method MemER (42.38 %).
- Recurrent memory underperforms β TTT/RMT sit at 18β22 %, barely above no-memory, suggesting fixed-size latent compression loses too much for these tasks.
- The oracle gap is huge: GroundSG + Oracle hits 84.08 %, so if perfect subgoals were available, symbolic memory would dominate β the bottleneck is the subgoal-predicting VLM, not the representation.
- TokenDrop < FrameSamp because aggressive patch pruning removes global spatial context (hurts e.g. StopCube, which needs object-distance awareness).
- Memory-as-Modulator is the best integration for both perceptual variants β its lightweight AdaLN conditioning preserves the base policy.

Regrouping the 16 tasks by functional demand exposes that different memory designs are complementary:
| Policy | Motion-Centric | Time-Sensitive | Short-Horizon | Long-Horizon | Scene-Change | Event-Salient |
|---|---|---|---|---|---|---|
| No Memory Ο0.5 | 11.17 | 6.67 | 16.39 | 28.45 | 13.41 | 32.96 |
| Symbolic GroundSG+QwenVL | 5.83 | 0.00 | 48.22 | 42.89 | 17.70 | 72.08 |
| Symbolic SimpleSG+QwenVL | 7.34 | 0.44 | 24.96 | 29.22 | 15.33 | 84.96 |
| Perceptual FrameSamp+Modul | 54.95 | 42.00 | 29.18 | 46.00 | 22.07 | 68.22 |
| MemER (prior) | 23.67 | 0.00 | 48.22 | 28.00 | 54.67 | 72.89 |
- Perceptual (FrameSamp+Modul) dominates motion-centric, time-sensitive, long-horizon tasks β where extended visual history and low-level motion cues matter.
- Symbolic wins short-horizon and event-salient tasks β where a clear subgoal is enough (note time-sensitive β 0 % for symbolic: language can't express "press exactly now").
- MemER wins dynamic scene-change tasks, because it keeps raw keyframe images during online execution.
Four history-dependent real-robot tasks (success out of 40):
| Method | Put Fruits | Track Cube | Repick Block | Draw Pattern | Total |
|---|---|---|---|---|---|
| Ο0.5 (no memory) | 2/10 | 1/10 | 1/10 | 0/10 | 4/40 |
| GroundSG + QwenVL (symbolic) | 9/10 | 3/10 | 5/10 | 2/10 | 19/40 |
| FrameSamp + Modul (perceptual) | 6/10 | 5/10 | 6/10 | 8/10 | 25/40 |
The simulation ordering holds on real hardware: perceptual memory generalizes best overall (25/40), symbolic memory is strongest on the language-friendly Put Fruits task, and no-memory Ο0.5 essentially fails (4/40).
Reformulating each task as an online VideoQA problem (participants pick the next high-level action; an oracle planner guarantees flawless low-level control) yields 90.5 % human success β but humans still fail consistently on long-horizon (PatternLock) and time-critical (StopCube) tasks. RoboMME is therefore hard even with perfect motor control: the memory/reasoning demand alone is non-trivial.
- Q1 β best design? Perceptual memory overall; FrameSamp + Modul = 44.51 %; Memory-as-Modulator is the best integration; TokenDrop < FrameSamp.
- Q2 β is symbolic reasoning alone enough? No β even oracle subgoals degrade on manipulation-intensive (StopCube, InsertPeg) and cluttered scenes (BinFill, PickHighlight); precise visuomotor control becomes the bottleneck.
- Q3 β human ceiling? 90.5 %, still imperfect β rigorous testbed.
- Q4 β task-dependence? No single representation dominates; strengths are complementary (symbolic β counting/short-horizon/event-salient; perceptual β imitation/motion/time-sensitive/long-horizon; MemER β scene-change).
- Q5 β efficiency? Memory adds compute; the paper sweeps a 64β512-token memory budget and reports the efficiencyβperformance (TFLOPs vs success) trade-off.
RoboMME is an insight/benchmark paper, not a new architecture β and that is the point. By holding the Ο0.5 backbone, the 512-token budget, and the task taxonomy constant, it converts an ad-hoc design folklore ("add a memory bank") into a measurable science. Three messages land:
- Match memory to task structure. There is no universal best; the practical recommendation is perceptual memory via AdaLN modulation as the strongest single default, with symbolic memory for counting/reference and a future path toward combining representations.
- The subgoal predictor, not the symbolic representation, is the ceiling. GroundSG+Oracle (84 %) vs GroundSG+QwenVL (33 %) shows most of the loss is in VLM subgoal prediction β a concrete target for improvement.
- Recurrent compression is not yet competitive for manipulation history at this budget.
In the 2026 VLA landscape β where long-horizon and occlusion-heavy manipulation is a frontier and memory modules are proliferating (MemoryVLA, HAMLET, VLA Memory review) β RoboMME provides the standardized yardstick the subfield was missing, which is why it landed as an ICML 2026 Oral.
- arXiv: 2603.04639
- ICML 2026: https://icml.cc/virtual/2026/poster/65933
- Code / benchmark: RoboMME/robomme_benchmark
- Figures 1β3 reproduced from the RoboMME paper (Univ. of Michigan / Stanford / Figure AI, 2026) for scholarly review; Β© the authors.
- One-page summary: ICML-2026-RoboMME
- VLA Memory review Β· MemoryVLA Β· HAMLET
- ICML 2026 index