IROS 2026 TempoFit - Heungwoo/research GitHub Wiki
IROS 2026 — TempoFit: Plug-and-Play Layer-Wise Temporal KV Memory for Long-Horizon VLA
Venue: IROS 2026 (Pittsburgh) · paper #1170 · Xi'an Jiaotong-Liverpool University (Sun, Yang, Zhang, Ma, Wu, … Chen). Paper: arXiv 2603.07647 (Mar 2026). The training-free memory retrofit of IROS 2026 — give a frozen VLA history by reusing its own prefix-attention K/V as a content-addressable runtime state — no new tokens, no trainable modules. Companions: VLA Memory · In-Context Imitation · RoboTTT · IROS 2026 survey.

1. Problem
Pretrained VLAs are strong at single-step manipulation but their inference is largely memoryless — brittle in non-Markovian long-horizon settings (occlusion, state aliasing, subtle post-action changes). Prior fixes inject history either by stacking frames (scales visual tokens + latency, adds near-duplicate pixels) or by learning extra temporal interfaces (require (re-)training, may break the original single-frame inference graph).
2. Method
TempoFit is a training-free temporal retrofit that upgrades frozen VLAs through state-level memory. Key insight: "prefix-attention K/V already forms a model-native, content-addressable runtime state; reusing them across timesteps introduces history without new tokens or trainable modules."
- Layer-wise FIFO prefix K/V stored at selected intermediate layers.
- Parameter-free K-to-K retrieval with Frame-Gap Temporal Bias (FGTB) — a fixed recency bias (inspired by NLP positional biases) that keeps decisions present-dominant.
- Pre-attention residual loading with norm-preserving rescaling injects the retrieved context without distribution shift under frozen weights.
3. Results
- LIBERO-Long: up to +4.0% average success on strong pretrained backbones, at near-real-time latency.
- Transfers consistently to CALVIN and real-robot long-horizon tasks.
- No retraining, no architecture change — a pure inference-time upgrade.
4. Why it matters (memory lens)
TempoFit is the "KV-cache is the memory" answer in IROS 2026's memory bifurcation (survey §5.2, VLA Memory): where structured methods build explicit stores (GaussMemory's 3D-Gaussian scene, PROMPT's memory trees), TempoFit exploits the model's own attention state — the cheapest possible retrofit, and training-free (unlike RoboTTT's fast weights, which need TTT training, or frame-stacking, which needs retraining). It's the pragmatic middle path: parametric/content memory with zero training cost. The FGTB recency-bias echoes the "selective history beats full context" lesson from VLA Memory's RSS-2026 trend.
Limitations (reviewer): +4.0% is a modest lift; FIFO/recency bias may drop genuinely old-but-relevant evidence (no learned retrieval); tuned on LIBERO-Long/CALVIN; frozen-weight residual injection assumes the K/V space is stable across the horizon.
5. Links
- Official program: IROS 2026 (paper #1170) · survey: IROS 2026
- Related: VLA Memory · RoboSSM · RoboTTT · In-Context Imitation
← Back to IROS 2026 survey · Home