RSS 2026 Long Context Robot Imitation Learning by - Heungwoo/research GitHub Wiki

Long-Context Robot Imitation Learning by Focusing on Key History Frames

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Imitation learning 3 · paper #201 Authors: Max Sobol Mark, Maria Attarian, Chuyuan Fu, Debidatta Dwibedi, Jacky Liang, Dhruv Shah, Aviral Kumar arXiv: 2602.15010 · program page

Summary compiled from the arXiv paper (v2, 18 Feb 2026; titled "BPP: Long-Context Robot Imitation Learning by Focusing on Key History Frames"); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Big Picture Policies keyframe conditioning (Figure 1 of arXiv 2602.15010, © the authors)

Fig. 1: A slightly out-of-distribution rollout (left) drifts off the dataset support, but VLM-detected keyframes stay close to the training distribution. A naive history-conditioned policy sees OOD inputs, whereas BPP conditions on keyframes; the right bar chart reports average real-world success — BPP 54% vs a naive+auxiliary-task policy 32%, naive history 13% and current-obs-only 14% — the "70% improvement."

Problem

Many robot tasks require attending to the history of past observations (e.g. remembering which places have been searched), yet the best-performing policies condition only on the current observation. Naively conditioning on past observations fails due to spurious correlations: policies latch onto incidental features of training histories that do not generalize. This stems from limited coverage of the space of possible histories during training, which grows exponentially with horizon, and existing regularization gives inconsistent benefits.

Method

Big Picture Policies (BPP) conditions on a minimal set of meaningful keyframes detected by an off-the-shelf vision-language model. By projecting diverse rollouts onto a compact representation of task-relevant events, BPP substantially reduces distribution shift between training and deployment without sacrificing expressivity.

Results

BPP is evaluated on four challenging real-world manipulation tasks and three simulation tasks, all requiring history conditioning. It achieves 70% higher success rates than the best comparison on the real-world evaluations.

Significance

Directly targets the history-coverage problem that undermines long-context imitation learning, using a VLM as a keyframe selector. Connects to Review-Human-Video-Transfer and Review-Realtime-Execution.

← Back to RSS 2026 survey · RSS-2026-Papers · Home