IROS 2026 RoboSSM - Heungwoo/research GitHub Wiki

IROS 2026 — RoboSSM: Scalable In-Context Imitation Learning via State-Space Models

Venue: IROS 2026 (Pittsburgh) · paper #3133 · KAIST · UT Austin (Yoo, Hu, Zhu, Liu, Liu, Martín-Martín, Peter Stone). Paper: arXiv 2509.19658 · OpenReview · code. The state-space datapoint for in-context imitation — replace the Transformer backbone of in-context imitation learning with an SSM (Longhorn) for linear-time inference and strong long-prompt extrapolation. Companions: In-Context Imitation · RoboTTT · VLA Memory · IROS 2026 survey.

RoboSSM vs the Transformer-based ICL baseline (ICRT) as the number of test-time demonstrations grows to 32: RoboSSM (orange) stays flat/rising across LIBERO-Object and LIBERO-90 scenes, while ICRT (purple) collapses toward 0 once prompts exceed the training length — SSMs extrapolate to long prompts where Transformers do not (results figure from Yoo et al., arXiv 2509.19658, © the authors)

1. Problem

In-context imitation learning (ICIL) lets a robot learn a task from a prompt of a few demonstrations — no deployment-time parameter updates, so it supports few-shot adaptation to novel tasks. But recent ICIL methods are Transformer-based, which has quadratic computational limits and underperforms when the prompt is longer than those seen at training (poor length extrapolation).

2. Method

RoboSSM is a scalable ICIL recipe built on state-space models:

  • Replaces the Transformer with Longhorn — a state-of-the-art SSM giving linear-time inference and strong extrapolation — making it well-suited to long-context prompts (many/long demos).
  • The recurrent SSM state is the memory: context is compressed into a fixed-size state rather than an ever-growing attention cache — the same "memory as parametric state, not cache" idea as RoboTTT's fast weights, but via SSM recurrence.

3. Results

  • On LIBERO, RoboSSM processes prompts up to 16× longer than seen in training and outperforms the Transformer-based ICIL baseline (ICRT) on unseen and long-horizon tasks — the gap widens exactly as the demonstration count grows (see figure: ICRT collapses past its training length while RoboSSM holds).
  • "For the first time, SSMs are an efficient and scalable backbone for ICIL." Code released.

4. Why it matters (in-context / memory lens)

RoboSSM is IROS 2026's clearest evidence that the long-context latency wall of in-context imitation (In-Context Imitation §6) has an architectural escape: SSMs give linear cost + length extrapolation, so more/longer demos help instead of blowing up compute. It sits in the In-Context Imitation §2 taxonomy as a token-sequence-ICL (E) method with an SSM (not Transformer) backbone, and pairs with the other two IROS "efficient long-context memory" answers — TempoFit (KV-cache memory) and RoboTTT (fast-weight TTT). Together they show the field converging on parametric/recurrent context over ever-larger attention windows.

Limitations (reviewer): LIBERO-only evaluation; SSM extrapolation on real multi-hour prompts and contact-rich tasks untested; Longhorn's fixed-size state may bottleneck very information-dense demos.

5. Links

← Back to IROS 2026 survey · Home