NeurIPS 2025 Knowledge Insulation - Heungwoo/research GitHub Wiki

Knowledge Insulating Vision-Language-Action Models — Train Fast, Run Fast, Generalize Better

Venue: NeurIPS 2025 (Spotlight) · Authors: Danny Driess, Jost Tobias Springenberg, Brian Ichter, Lili Yu, Adrian Li-Bell, Karl Pertsch, Allen Z. Ren, Homer Walke, Quan Vuong, Lucy Xiaoyang Shi, Sergey Levine — Physical Intelligence · arXiv: 2505.23705 (submitted 29 May 2025) Category: VLA Architecture / Training Significance: Published formalism bridging π0.5 → π0.6 → π0.7

Approach diagram

flowchart LR
  V[Vision] --> VLM[VLM backbone]
  L[Language] --> VLM
  VLM -- FAST-tokenized actions --> CE[Cross-entropy loss<br/>trains VLM representation]
  VLM -- conditioning features --> AE[Continuous action expert<br/>flow-matching]
  AE --> A[Actions]
  AE -. 🚫 NO gradient propagation back to VLM .-> VLM
Loading

Problem

Training a continuous flow-matching action expert attached to a VLM has a quiet failure mode: the action-expert loss back-propagates into the VLM and corrupts language-grounded representations (the thing you paid web-scale pretraining to get). Naive joint training degrades generalization; training the VLM only on language and bolting on a frozen action head leaves performance on the table. Prior π-series work had the recipe informally; it needed to be published and ablated.

Method

Knowledge Insulation (KI) does two things at once:

  1. VLM training signal via FAST-discretized actions. Continuous actions are tokenized (FAST, DCT-based) and used as discrete prediction targets for the VLM — training its representations with the stable cross-entropy objective it was built for.
  2. Continuous action expert attends the VLM's conditioning features and is trained with flow matching to produce continuous actions.
  3. The gradient from the action expert is blocked from propagating back into the VLM. The VLM is trained only by its own cross-entropy loss on FAST tokens + standard co-training tasks.

Net effect: same VLM learns from discrete action supervision AND is used as conditioning for continuous control — but its parameters are insulated from the continuous-head gradient.

Results

  • KI is built on the π0 / π0.5-FAST architecture (VLM initialized from PaliGemma, as in π0, plus a flow-matching action expert). The paper turns π0.5's two-stage "FAST-pretrain then add action expert" recipe into a single-stage one; the evaluated model is branded π0.5 + KI on the project page.
  • Training speed: the KI recipe converges about as fast as the purely autoregressive π0-FAST, whereas π0 (continuous action expert only) needs roughly 7.5× as many training steps to reach comparable performance. The dual discrete+continuous objective adds ~20% per-step compute, offset by faster convergence.
  • Simulation benchmarks: state-of-the-art LIBERO-90 96.0% and LIBERO-Spatial 98.0%; DROID 0.55 ± 0.09 vs π0 0.49 ± 0.09 and π0-FAST 0.45 ± 0.09.
  • Generalization / language following: stop-gradient insulation plus co-training on web VLM data markedly improves instruction following and out-of-distribution object generalization; a frozen backbone reaches ~0%, confirming VLM pretraining alone lacks sufficient robotics representations.
  • The recipe is later used in π0.6 (Nov 2025) and π0.7 (Apr 2026).

Significance

The cleanest published bridge from CoRL 2025 to ICLR 2026:

  • π0.5 (Oral, Apr 2025) used a two-stage version (FAST-pretrain, then add action expert in post-training);
  • KI (arXiv May 2025; NeurIPS 2025 Spotlight) formalizes it into a single-stage recipe and ablates it on the π0/π0.5-FAST architecture;
  • π0.6 (Nov 2025) and π0.7 (Apr 2026) build on it as part of their training recipe.

Also a conceptual generalization: "insulate pretrained representations from task-specific gradients" is now the default pattern across FASTER, OmniSAT, and HyperVLA (different realizations of the same principle).

Links

Related pages

← Back to NeurIPS-2025

⚠️ **GitHub.com Fallback** ⚠️