Review RoboMME - Heungwoo/research GitHub Wiki

RoboMME β€” In-Depth Review: Benchmarking & Understanding Memory for Robotic Generalist Policies

Venue: ICML 2026 (Oral) Β· arXiv: 2603.04639 Β· Code: RoboMME/robomme_benchmark Affiliations: University of Michigan Β· Stanford Β· Figure AI Traction (2026-06-09): β˜…111 GitHub Β· 6 citations (arXiv preprint) One-page version: ICML-2026-RoboMME Β· See also VLA Memory review

TL;DR

RoboMME is the first benchmark that turns "which memory mechanism should a VLA use?" into a controlled, apples-to-apples comparison. It contributes (1) a cognitively-motivated task taxonomy (temporal / spatial / object / procedural memory) over 16 manipulation tasks, and (2) the MME-VLA suite β€” 14 memory-augmented policies all built on the same Ο€0.5 backbone, factored as 3 memory representations Γ— instantiations Γ— 3 integration mechanisms. The headline empirical result: no single memory design wins everywhere; the best overall non-oracle policy, FrameSamp + Modulator (perceptual memory, 44.51 % avg), beats all symbolic and recurrent variants and the prior SOTA, but symbolic memory still wins on counting / event-salient tasks. Humans score 90.5 %, so headroom is large.

RoboMME overview (Figure 1 from the RoboMME paper, 2026)


1. Why a memory benchmark

Generalist VLAs (Ο€0.5, OpenVLA, etc.) are largely Markovian β€” they condition on the current observation only. Real manipulation is history-dependent: counting repeated actions, re-finding an object that was occluded, imitating a demonstrated motion strategy. Memory-augmented VLAs have appeared (MemoryVLA, HAMLET, SAM2Act, MemER…), but each is evaluated on its own narrow, non-standardized setup, so the community cannot say which memory representation helps which kind of task. RoboMME fixes the evaluation substrate and then sweeps the design space on a single backbone.

Benchmark comparison (Table 2)

Benchmark Non-Markov Partial Obs Dynamic Video Memory types #Tasks #Demos Avg #Steps
RLBench18 / CALVIN / LIBERO / VLABench βœ— βœ— βœ— βœ— T 18–130 β€” 115–584
RoboCerebra βœ— βœ— βœ“ βœ— T 100 1,000 2,972
MemoryBench (SAM2Act) βœ“ βœ— βœ— βœ— T+S 3 300 312
MIKASA-robo βœ“ βœ“ βœ“ βœ— T+S+O 12 1,250 72
RoboMME (ours) βœ“ βœ“ βœ“ βœ“ T+S+O+P 16 1,600 481

RoboMME is the only suite that is simultaneously non-Markovian, partially observable, dynamic, video-conditioned, and covers all four memory types with subgoal + keyframe annotations.

2. The cognitively-motivated task taxonomy (Β§3.1)

16 tasks in 4 suites, each targeting one primary memory faculty (T = temporal, S = spatial, O = object, P = procedural):

Suite Memory Tasks What it stresses
Counting T PickXTimes, BinFill, SwingXTimes, StopCube counting repetitions over time; StopCube is time-critical
Permanence S VideoUnmask, ButtonUnmask, VideoUnmaskSwap, ButtonUnmaskSwap object permanence under occlusion / masking; Swap adds tracking through container shuffles
Reference O (+T) PickHighlight, VideoRepick, VideoPlaceButton, VideoPlaceOrder resolving visual / action / language references to a specific object or target seen earlier in a video
Imitation P (+O) MoveCube, InsertPeg, PatternLock, RouteStick replicating a demonstrated motion strategy (pick-place vs push vs hook; linear vs circular trajectories)

Tasks range from 208 to 1,134 average timesteps (VideoPlaceOrder is the longest, a long-video ordinal-reference task), so the suite spans short reactive tasks to very long-horizon video-conditioned ones.

3. Memory implementation β€” the MME-VLA suite (Β§4)

This is the core technical contribution: a factored design space on a frozen-recipe Ο€0.5 backbone, so differences are attributable to the memory design rather than the base model.

Framework of the MME-VLA suite (Figure 2 from the RoboMME paper, 2026)

3.1 Three memory representations (each with 2 instantiations)

(A) Symbolic memory β€” history is summarized as interpretable language subgoals. At each step an auxiliary VLM conditions on the current image + prior subgoals and emits the next subgoal. Two instantiations:

  • SimpleSG β€” plain instructions (pick up the green cube).
  • GroundSG β€” grounded subgoals that append image coordinates (pick up the green cube at [63, 152]), which helps spatial reasoning.
  • Subgoal generators: Gemini-2.5-Pro (prompt-only), a fine-tuned Qwen3-VL-4B (QwenVL) trained on the dataset's subgoal annotations, or Oracle (simulator ground truth, used as an upper bound). Symbolic memory requires extra subgoal annotations to train the VLM.

(B) Perceptual memory β€” history is a sequence of visual tokens from past images, extracted by the Ο€0.5 vision encoder (no proprioception needed). Two token-selection strategies:

  • TokenDrop β€” token dropping (Γ  la TimeChat): removes temporally redundant patches by RGB difference.
  • FrameSamp β€” uniform frame sampling (Γ  la VideoLLM): evenly downsamples frames and concatenates their tokens.

(C) Recurrent memory β€” compress the visual-token sequence into a fixed-size latent state via recurrence. Two models:

  • TTT β€” Test-Time Training: updates fast weights online with a self-supervised loss, then reads out features.
  • RMT β€” Recurrent Memory Transformer: processes the stream in segments and recurrently updates learnable memory slots with a transformer.

3.2 Three memory integration mechanisms (Β§4.2)

Symbolic memory integrates trivially (concatenate subgoals onto the task instruction). The neural memory tokens (perceptual / recurrent) need architectural plumbing, so the paper studies three ways to inject them into Ο€0.5's VLM + action-expert structure:

Mechanism How memory enters the policy
Memory-as-Context memory tokens are concatenated with the original inputs and jointly processed by the VLM expert β€” directly shifts VLM features.
Memory-as-Modulator uses AdaLN: action features cross-attend to memory tokens, and the result is projected into scale & shift parameters that modulate the normalized action features before each feed-forward block. Lightweight, feature-wise conditioning that preserves the base policy.
Memory-as-Expert a dedicated lightweight memory expert processes memory tokens on its own pathway, with a blockwise causal attention scheme β€” the action expert attends to both the VLM and memory experts, but VLM and memory experts do not attend to each other (limits interference, preserves VLM behavior).

3.3 Counting the 14 variants

Representation Instantiations Γ— Integration # policies
Symbolic SimpleSG, GroundSG (concat) 2
Perceptual {TokenDrop, FrameSamp} Γ— {Context, Modul, Expert} 6
Recurrent {TTT, RMT} Γ— {Context, Modul, Expert} 6
Total 14

Prior methods compared (4): (1) Ο€0.5 no-memory; (2) Ο€0.5 w/ past actions (concatenates past actions as explicit symbolic memory, UniVLA-style); (3) SAM2Act+ (SAM2 memory bank + discrete keyframe actions executed by an external motion planner); (4) MemER (VLM infers symbolic subgoals from a buffer of selected keyframes β€” a hybrid of perceptual + symbolic memory).

Evaluation protocol: fixed memory budget of 512 tokens (matching the current-observation image-token count) for fair comparison; 50 episodes/task (800 total) with held-out seeds; single multi-task model, ≀1,300 steps/episode; scores averaged over the last 3 checkpoints Γ— 3 seeds (9 runs).

3.4 Implementation detail β€” the best config (FrameSamp + Modulator) on the VLM+AE stack

Grounded in the paper's Appendix A (math) and the official code in RoboMME/robomme_policy_learning β€” the model suite lives there (an openpi/Ο€0.5 fork), while the benchmark env is in robomme_benchmark. Weights: Yinpei/mme_vla_suite.

The base VLM + Action-Expert (AE) backbone

Ο€0.5 is a Mixture-of-Transformers (MoT) dual-expert model. In code, HistoryPi0 subclasses openpi's Pi0 (src/mme_vla_suite/models/integration/history_pi0.py) and runs two Gemma experts under one blockwise-causal attention:

  • VLM expert (PaliGemma = SigLIP + Gemma-2B) β€” a SigLIP image encoder (PaliGemma.img) turns each camera view into vision tokens; these plus the tokenized language prompt form the prefix. This is the "understand the scene" expert (paligemma_variant = "gemma_2b").
  • Action expert (Gemma-300M) β€” takes the noised action chunk + a flow-matching timestep and forms the suffix; it predicts the flow-matching velocity that denoises the action chunk. In Ο€0.5 it is conditioned on the diffusion timestep via AdaRMSNorm (use_adarms=[False, True], adarms_cond = time_emb).

Attention is blockwise-causal: action (suffix) tokens attend to the VLM prefix; the mask is built so vision tokens never attend to memory tokens, but action tokens can (make_attn_mask, mask_na). The training loss is the standard flow-matching objective on the action expert's predicted velocity.

The perceptual-memory branch (FrameSamp)

From a streaming horizon of 16 past frames (streaming_obs_horizon: 16, single view), FrameSamp evenly samples frames (S_t = EvenSample({0…t}; N), Appendix Eq. 14) and mean-pools each to 16 tokens (token_per_image: 16, pool_type: mean), under a 512-token budget. The mem_encoder builds each memory token from the SigLIP image embedding (2048-d) + a positional embedding (768-d) β€” and in the best config no proprioceptive state (use_state_emb: false, confirming the paper's "visual tokens alone suffice") β€” then projects to memory_token_dim: 1024, yielding the memory tokens M_t (obs.static_image_emb / static_pos_emb in embed_memory). (Config: models/config/robomme/perceptual-framesamp-modul.yaml. TokenDrop instead keeps the least-redundant patches by RGB/Jaccard similarity; recurrent variants feed recur_* streams through a TTT or RMT encoder. memory_token_dim is 2048 for context-integration but 1024 for modulation/expert.)

How "Memory-as-Modulator" injects M_t into the AE

Crucially, in modulation mode the memory tokens do not enter the token sequence (that is the context variant, which prepends M_t to the prefix). Instead memory conditions the AE through feature-wise AdaLN, implemented in history_gemma.py and gated by mem_mods=[False, True] so that only the action expert is modulated β€” the VLM expert is left untouched, preserving its pretrained behaviour. The injection happens inside HistoryBlock β€” i.e. in every transformer layer β€” right before the FFN sublayer, and only on the action-expert stream (the code guard is if i == len(xs)-1 and integration_type=="modulation", where the loop index i runs over experts, so len(xs)-1 selects the action expert, not a particular layer). At each such layer (Appendix Eqs. 21–23):

  1. Cross-attention, action β†’ memory (MemoryAttention): query = current action features sΜƒβ‚œ, key/value = memory tokens Mβ‚œ, with a softmax masked over valid memory slots β†’ rβ‚œ = Attn_mod(Q = sΜƒβ‚œ, K = Mβ‚œ, V = Mβ‚œ).
  2. Produce scale & shift (MemoryRMSNorm β†’ a single Dense(2Β·width) on rβ‚œ): (Ξ³β‚œ, Ξ²β‚œ) = MLP_mod(rβ‚œ).
  3. AdaLN-modulate the action features: Εβ‚œ = Ξ³β‚œ βŠ™ Norm(sΜƒβ‚œ) + Ξ²β‚œ (code: normed * (1 + scale) + shift, with near-zero init so the block starts as identity).

So the action expert is conditioned on both signals simultaneously: Ο€0.5's existing time-AdaRMSNorm (diffusion timestep) and the new memory-AdaLN (history). Because the conditioning is a lightweight per-feature scale/shift rather than extra sequence tokens, it barely perturbs the base policy β€” which is exactly why the paper finds modulation the most robust integration and why TokenDrop's lost global context hurts less here than under context integration.

End-to-end data flow

flowchart LR
  subgraph IN[Inputs]
    OBS["current obs<br/>front + wrist RGB"]
    LANG["language goal"]
    HIST["frame history"]
  end
  HIST -->|"FrameSamp: 16-frame horizon,<br/>mean-pool 16 tok/img, ≀512 budget"| MENC["mem_encoder<br/>(SigLIP 2048-d + pos 768-d,<br/>no state) β†’ 1024-d"]
  MENC --> MT["memory tokens Mβ‚œ"]
  OBS --> SIG["SigLIP encoder"]
  SIG --> VLM["VLM expert<br/>(PaliGemma: SigLIP + Gemma-2B)"]
  LANG --> VLM
  VLM -->|"prefix KV (blockwise-causal)"| AE
  NOISE["noised action chunk + flow time Ο„"] --> AE["Action expert (Gemma-300M)"]
  TIME["time_emb"] -->|"AdaRMSNorm"| AE
  MT -->|"cross-attn β†’ Ξ³,Ξ² (AdaLN)"| AE
  AE --> V["flow velocity vΞΈ β†’ denoise β†’ action chunk aβ‚œ"]
Loading

Exact config of the best model (perceptual-framesamp-modul.yaml)

Field Value Meaning
representation_type perceptual visual-token memory
perceptual_memory.type frame_sampling even frame sampling (vs token_dropping)
integration_type modulation AdaLN on the action expert
streaming_obs_horizon 16 past frames fed to the memory encoder
token_per_image 16 mean-pooled tokens kept per frame (vs 64 default)
pool_type mean pooling over patch tokens
budget 512 total memory-token cap
num_views 1 single (fixed) camera view for memory
use_pos_emb / use_state_emb true / false + positional emb, no proprio state
memory_feature.img.input_dim 2048 SigLIP image-embedding dim from Ο€0.5
memory_token_dim 1024 memory-token width (2048 for context)
VLM / action / memory expert gemma_2b / gemma_300m / gemma_150m gemma_150m only for the expert variant

Key code pointers: models/integration/history_pi0.py (HistoryPi0, embed_memory, embed_prefix/suffix, compute_loss routing on integration_type ∈ {context, modulation, expert}) Β· models/integration/history_gemma.py (HistoryBlock, MemoryAttention, MemoryRMSNorm, applying mem_attn + AdaLN before the FFN of every layer on the action-expert stream) Β· src/openpi/models/pi0.py (base MoT Ο€0.5). The expert variant instead adds a third Gemma-150M memory expert with the blockwise rule "action expert attends to VLM + memory experts; VLM and memory do not attend each other".

4. Performance comparison (Β§5.2)

4.1 Main results β€” success rate (%) by task suite (Table 3)

Representation Policy Integ./VLM Counting Permanence Reference Imitation AVG
No Memory Ο€0.5 – 28.78 17.00 17.17 8.78 17.93
Symbolic SimpleSG Oracle 82.56 21.56 32.28 61.94 49.58
Symbolic GroundSG Oracle 83.86 93.31 95.17 63.98 84.08
Symbolic SimpleSG QwenVL 44.61 19.61 25.22 26.56 29.00
Symbolic GroundSG QwenVL 38.00 39.34 31.56 21.89 32.70
Perceptual TokenDrop Modul 52.33 26.83 34.72 38.28 38.04
Perceptual FrameSamp Modul 65.22 25.11 36.33 51.39 πŸ† 44.51
Perceptual FrameSamp Expert 66.78 25.17 24.11 28.94 36.25
Recurrent TTT Expert 36.00 22.95 19.67 10.78 22.35
Recurrent RMT Modul 34.14 15.78 21.11 9.67 20.17
Prior SAM2Act+ – 35.33 26.00 16.83 7.33 21.37
Prior MemER – 48.83 53.17 38.00 29.50 42.38

(Oracle rows are upper bounds β€” they use simulator-ground-truth subgoals. πŸ† = best non-oracle.)

Reading the table:

  • Memory helps a lot: the best learned policy (44.51 %) is ~2.5Γ— the no-memory Ο€0.5 baseline (17.93 %).
  • Perceptual > Symbolic > Recurrent on average (for learned, non-oracle settings). FrameSamp + Modul (44.51 %) is the best learned policy and edges out the strongest prior method MemER (42.38 %).
  • Recurrent memory underperforms β€” TTT/RMT sit at 18–22 %, barely above no-memory, suggesting fixed-size latent compression loses too much for these tasks.
  • The oracle gap is huge: GroundSG + Oracle hits 84.08 %, so if perfect subgoals were available, symbolic memory would dominate β€” the bottleneck is the subgoal-predicting VLM, not the representation.
  • TokenDrop < FrameSamp because aggressive patch pruning removes global spatial context (hurts e.g. StopCube, which needs object-distance awareness).
  • Memory-as-Modulator is the best integration for both perceptual variants β€” its lightweight AdaLN conditioning preserves the base policy.

4.2 By functional requirement (Table 11) β€” complementary strengths

Performance across task characteristics (Figure 3 from the RoboMME paper, 2026)

Regrouping the 16 tasks by functional demand exposes that different memory designs are complementary:

Policy Motion-Centric Time-Sensitive Short-Horizon Long-Horizon Scene-Change Event-Salient
No Memory Ο€0.5 11.17 6.67 16.39 28.45 13.41 32.96
Symbolic GroundSG+QwenVL 5.83 0.00 48.22 42.89 17.70 72.08
Symbolic SimpleSG+QwenVL 7.34 0.44 24.96 29.22 15.33 84.96
Perceptual FrameSamp+Modul 54.95 42.00 29.18 46.00 22.07 68.22
MemER (prior) 23.67 0.00 48.22 28.00 54.67 72.89
  • Perceptual (FrameSamp+Modul) dominates motion-centric, time-sensitive, long-horizon tasks β€” where extended visual history and low-level motion cues matter.
  • Symbolic wins short-horizon and event-salient tasks β€” where a clear subgoal is enough (note time-sensitive β‰ˆ 0 % for symbolic: language can't express "press exactly now").
  • MemER wins dynamic scene-change tasks, because it keeps raw keyframe images during online execution.

4.3 Real-world validation (Table 4)

Four history-dependent real-robot tasks (success out of 40):

Method Put Fruits Track Cube Repick Block Draw Pattern Total
Ο€0.5 (no memory) 2/10 1/10 1/10 0/10 4/40
GroundSG + QwenVL (symbolic) 9/10 3/10 5/10 2/10 19/40
FrameSamp + Modul (perceptual) 6/10 5/10 6/10 8/10 25/40

The simulation ordering holds on real hardware: perceptual memory generalizes best overall (25/40), symbolic memory is strongest on the language-friendly Put Fruits task, and no-memory Ο€0.5 essentially fails (4/40).

4.4 Human study

Reformulating each task as an online VideoQA problem (participants pick the next high-level action; an oracle planner guarantees flawless low-level control) yields 90.5 % human success β€” but humans still fail consistently on long-horizon (PatternLock) and time-critical (StopCube) tasks. RoboMME is therefore hard even with perfect motor control: the memory/reasoning demand alone is non-trivial.

5. Key findings (the five research questions)

  • Q1 β€” best design? Perceptual memory overall; FrameSamp + Modul = 44.51 %; Memory-as-Modulator is the best integration; TokenDrop < FrameSamp.
  • Q2 β€” is symbolic reasoning alone enough? No β€” even oracle subgoals degrade on manipulation-intensive (StopCube, InsertPeg) and cluttered scenes (BinFill, PickHighlight); precise visuomotor control becomes the bottleneck.
  • Q3 β€” human ceiling? 90.5 %, still imperfect β†’ rigorous testbed.
  • Q4 β€” task-dependence? No single representation dominates; strengths are complementary (symbolic β†’ counting/short-horizon/event-salient; perceptual β†’ imitation/motion/time-sensitive/long-horizon; MemER β†’ scene-change).
  • Q5 β€” efficiency? Memory adds compute; the paper sweeps a 64–512-token memory budget and reports the efficiency–performance (TFLOPs vs success) trade-off.

6. Significance

RoboMME is an insight/benchmark paper, not a new architecture β€” and that is the point. By holding the Ο€0.5 backbone, the 512-token budget, and the task taxonomy constant, it converts an ad-hoc design folklore ("add a memory bank") into a measurable science. Three messages land:

  1. Match memory to task structure. There is no universal best; the practical recommendation is perceptual memory via AdaLN modulation as the strongest single default, with symbolic memory for counting/reference and a future path toward combining representations.
  2. The subgoal predictor, not the symbolic representation, is the ceiling. GroundSG+Oracle (84 %) vs GroundSG+QwenVL (33 %) shows most of the loss is in VLM subgoal prediction β€” a concrete target for improvement.
  3. Recurrent compression is not yet competitive for manipulation history at this budget.

In the 2026 VLA landscape β€” where long-horizon and occlusion-heavy manipulation is a frontier and memory modules are proliferating (MemoryVLA, HAMLET, VLA Memory review) β€” RoboMME provides the standardized yardstick the subfield was missing, which is why it landed as an ICML 2026 Oral.

Links

Related pages

← Back to ICML-2026 Β· Home

⚠️ **GitHub.com Fallback** ⚠️