ICLR 2026 Memory Benchmark Robots - Heungwoo/research GitHub Wiki

MIKASA — a memory benchmark for RL agents and tabletop robots

Venue: ICLR 2026 (Poster) · Authors: Egor Cherepanov, Nikita Kachaev, Alexey Kovalev, Aleksandr Panov · arXiv:2502.10550 · Category: Data / memory / representation for manipulation · Trend tag: Memory / benchmarks

Official title: Memory, Benchmark & Robots: A Benchmark for Solving Complex Tasks with Reinforcement Learning.

Approach diagram

flowchart LR
  Cls[Memory-task classification<br/>object / spatial / sequential / capacity] --> Base[MIKASA-Base<br/>unified memory-RL tasks]
  Cls --> Robo[MIKASA-Robo<br/>32 tabletop manipulation tasks]
  Base --> Eval[Systematic evaluation<br/>of memory-enhanced agents]
  Robo --> Eval
  Eval --> Out[Diagnose what kind of memory<br/>a policy actually has]
Loading

Problem

Memory is essential for tasks with temporal and spatial dependencies under partial observability, yet RL lacks a universal benchmark to measure an agent's memory across diverse scenarios. This gap is especially acute in tabletop robotic manipulation, where no standardized memory benchmark existed.

Method

MIKASA (Memory-Intensive Skills Assessment Suite for Agents) makes three contributions:

  • A classification framework for memory-intensive RL tasks, separating distinct kinds of required memory.
  • MIKASA-Base — a unified benchmark aggregating memory-RL tasks for systematic evaluation of memory-enhanced agents across scenarios.
  • MIKASA-Robo — a new suite of 32 carefully designed memory-intensive tabletop manipulation tasks (installable via pip install mikasa-robo-suite), built on ManiSkill.

Results

The benchmark is the contribution: it standardizes how memory capability is measured rather than proposing a single policy. Specific per-task baseline numbers are reported in the full paper and omitted here pending confirmation. MIKASA-Robo's 32 tasks span varied memory demands so that an agent's success profile reveals which memory type it possesses.

Significance

Gives the field a principled, reproducible way to test whether a "memory" policy genuinely uses memory, separating object/spatial/sequential/capacity demands instead of conflating them. Complements model-side work like MemoryVLA and HAMLET by providing the evaluation substrate; MIKASA-Robo's tabletop tasks are already used as a memory testbed elsewhere in this wiki.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️