ICML 2026 VLA Forgetting - Heungwoo/research GitHub Wiki
Pretrained VLAs are Surprisingly Resistant to Forgetting — continual learning in large-scale Vision-Language-Action models
Venue: ICML 2026 (Oral) Category: Analysis-Insight Affiliations: Huihan Liu, Changyeon Kim, Bo Liu, Minghuan Liu, Yuke Zhu (2026) Traction (2026-06): 4 citations (arXiv)

Problem
Continual learning — acquiring new skills over time without catastrophically forgetting old ones — has been studied mostly in small behavior-cloning (BC) policies trained from scratch, where forgetting is severe. Whether the same dynamics hold for modern large-scale pretrained Vision-Language-Action (VLA) models was underexplored. This empirical study asks whether the conventional stability–plasticity trade-off still governs VLAs, and what role large-scale pretraining plays.
Method
The authors run continual-learning experiments on LIBERO (four task suites: Spatial, 10, Object, Goal), training a separate model per suite with a fixed task ordering and carrying weights across tasks. They evaluate two pretrained VLA backbones — Pi0 and GR00T N1.5 (differing in architecture, parameter count, and pretraining data) — against a non-pretrained BC-Transformer (plus BC-ViT, BC-Diffusion-Policy). The main continual-learning strategy is simple Experience Replay (ER): at task k, training mixes the current task with a buffer of M randomly sampled past-task transitions (default M=1000). Metrics are average success rate (SR) and Negative Backward Transfer (NBT) — positive NBT = forgetting, ≤0 = retention or positive transfer.
To isolate pretraining, they compare three Pi0 variants — VL+Action (full robot-data pretraining on a PaliGemma backbone), VL-only (PaliGemma backbone, no robot pretraining), and from scratch — sweeping replay-buffer size. A knowledge-retention probe (Figure 6) decomposes the model into a vision-language (VL) backbone and an action head, then swaps components across training stages and measures recovery speed when re-finetuning a "forgotten" task.

Results
Pretrained VLAs with ER achieve near-zero or even positive backward transfer across LIBERO suites — learning new tasks can improve prior-task performance, challenging the classic stability–plasticity trade-off. The effect is consistent across both Pi0 and GR00T N1.5, suggesting it is a general property of pretrained VLAs rather than an architecture artifact. The advantage is starkest in low-replay regimes: at a 2% buffer (100 samples/task), pretrained VLAs hold NBT around 0.1–0.2 while non-pretrained baselines deteriorate to 0.4–0.5 (2–4× more forgetting), and the small models need ~20% replay to match. A Pareto-frontier analysis (forgetting vs. buffer size) shows robot-data pretraining shifts the curve toward zero forgetting. Ablations indicate the training objective barely matters (Pi0 flow-matching vs. ℓ2 on LIBERO-Spatial: NBT −0.0003 vs. 0.016), whereas model size does — from-scratch Pi0 NBT improves from 0.110 (17M LLM + ResNet vision) to −0.052 (250M + SigLIP-So400M/14). Finally, the component-swap study shows seemingly forgotten skills are retained, not erased: re-finetuning recovers prior-task performance rapidly (Pi0 recovers within ~20% of the training steps that BC-Transformer needs).
Significance
The paper reframes continual learning for VLAs: large-scale pretraining fundamentally changes the dynamics, making simple Experience Replay a surprisingly strong baseline that can reach zero forgetting at small buffer sizes. The insight that degraded performance reflects suppressed but retained knowledge — recoverable with minimal finetuning — suggests practitioners need far less replay data than the from-scratch literature implies, and that scale, not specialized anti-forgetting regularizers (e.g., EWC), drives robustness.
Links
- arXiv: 2603.03818
- ICML 2026: https://icml.cc/virtual/2026/oral/71117
← Back to ICML-2026