IROS 2026 TempoFit - Heungwoo/research GitHub Wiki

IROS 2026 — TempoFit: Plug-and-Play Layer-Wise Temporal KV Memory for Long-Horizon VLA

Venue: IROS 2026 (Pittsburgh) · paper #1170 · Xi'an Jiaotong-Liverpool University (Sun, Yang, Zhang, Ma, Wu, … Chen). Paper: arXiv 2603.07647 (Mar 2026). The training-free memory retrofit of IROS 2026 — give a frozen VLA history by reusing its own prefix-attention K/V as a content-addressable runtime state — no new tokens, no trainable modules. Companions: VLA Memory · In-Context Imitation · RoboTTT · IROS 2026 survey.

TempoFit — (A) a frozen VLA runs per timestep; at one selected memory layer a compact memory state is carried forward across timesteps t−2 → t−1 → t; (B) inside that layer: current K/V are stored to a FIFO history, parameter-free K-to-K retrieval produces history logits that a Frame-Gap Temporal Bias (recency-weighted) adjusts, softmax-pools the history K/V, and injects them via norm-preserving residual loading before self-attention (method figure from Sun et al., arXiv 2603.07647, © the authors)

1. Problem

Pretrained VLAs are strong at single-step manipulation but their inference is largely memoryless — brittle in non-Markovian long-horizon settings (occlusion, state aliasing, subtle post-action changes). Prior fixes inject history either by stacking frames (scales visual tokens + latency, adds near-duplicate pixels) or by learning extra temporal interfaces (require (re-)training, may break the original single-frame inference graph).

2. Method

TempoFit is a training-free temporal retrofit that upgrades frozen VLAs through state-level memory. Key insight: "prefix-attention K/V already forms a model-native, content-addressable runtime state; reusing them across timesteps introduces history without new tokens or trainable modules."

  • Layer-wise FIFO prefix K/V stored at selected intermediate layers.
  • Parameter-free K-to-K retrieval with Frame-Gap Temporal Bias (FGTB) — a fixed recency bias (inspired by NLP positional biases) that keeps decisions present-dominant.
  • Pre-attention residual loading with norm-preserving rescaling injects the retrieved context without distribution shift under frozen weights.

3. Results

  • LIBERO-Long: up to +4.0% average success on strong pretrained backbones, at near-real-time latency.
  • Transfers consistently to CALVIN and real-robot long-horizon tasks.
  • No retraining, no architecture change — a pure inference-time upgrade.

4. Why it matters (memory lens)

TempoFit is the "KV-cache is the memory" answer in IROS 2026's memory bifurcation (survey §5.2, VLA Memory): where structured methods build explicit stores (GaussMemory's 3D-Gaussian scene, PROMPT's memory trees), TempoFit exploits the model's own attention state — the cheapest possible retrofit, and training-free (unlike RoboTTT's fast weights, which need TTT training, or frame-stacking, which needs retraining). It's the pragmatic middle path: parametric/content memory with zero training cost. The FGTB recency-bias echoes the "selective history beats full context" lesson from VLA Memory's RSS-2026 trend.

Limitations (reviewer): +4.0% is a modest lift; FIFO/recency bias may drop genuinely old-but-relevant evidence (no learned retrieval); tuned on LIBERO-Long/CALVIN; frozen-weight residual injection assumes the K/V space is stable across the horizon.

5. Links

← Back to IROS 2026 survey · Home