ICML 2026 SpecPrune VLA - Heungwoo/research GitHub Wiki
SpecPrune-VLA — Action-aware self-speculative token pruning for faster VLA inference
Venue: ICML 2026 (Poster) Category: Efficiency Traction (2026-06): 24 citations (arXiv)

Problem
Pruning visual tokens is a promising way to accelerate compute-bound Vision-Language-Action (VLA) inference, but existing pruning methods only look at local information in the current action generation and ignore the global information carried by previous actions. The paper reports that this myopia causes "a reduction of more than 20% in the success rate and limited speedup in some scenarios." The baseline of comparison is the high-performing OpenVLA-OFT model.
Method
SpecPrune-VLA is a training-free pruning method built on the insight that information across consecutive actions is highly similar, so token selection should combine local (current-action) and global (previous-action) cues. It has three components:
- Static token pruning at the action level. Token redundancy is assessed using global information from previous actions together with local information from the current generation, statically reducing the number of visual tokens for the whole action.
- Dynamic token pruning at the layer level. Exploiting the relevance between tokens and model layers, tokens are dynamically pruned according to their layer-specific importance.
- Lightweight action-aware controller. Generated actions are categorized into coarse-grained vs. fine-grained based on speed; fine-grained actions are sensitive to pruning error, so a lightweight controller identifies the current action granularity and adjusts the pruning strategy accordingly.

Results
Compared with OpenVLA-OFT on the LIBERO simulation benchmark, SpecPrune-VLA achieves an average 1.46× speedup on an NVIDIA A800 GPU and 1.57× speedup on an NVIDIA GeForce RTX 3090 GPU "with negligible loss on task success rate." The paper's headline result: "1.57x speedup in LIBERO simulation and 1.70x on real-world tasks".
Significance
As 2026 VLA models push toward real-time control, training-free acceleration that preserves success rate is highly valuable for deployment. SpecPrune-VLA's reuse of cross-action global context — rather than treating each action in isolation — addresses a concrete failure mode of prior token-pruning approaches and shows that meaningful speedups are attainable on commodity hardware (RTX 3090) as well as datacenter GPUs.
Links
- arXiv: 2509.05614
- ICML 2026: https://icml.cc/virtual/2026/poster/64525
← Back to ICML-2026