ICML 2026 SpecPrune VLA - Heungwoo/research GitHub Wiki

SpecPrune-VLA — Action-aware self-speculative token pruning for faster VLA inference

Venue: ICML 2026 (Poster) Category: Efficiency Traction (2026-06): 24 citations (arXiv)

Latency and arithmetic-intensity breakdown motivating pruning in VLA models (Figure 1 from Author et al., 2026)

Problem

Pruning visual tokens is a promising way to accelerate compute-bound Vision-Language-Action (VLA) inference, but existing pruning methods only look at local information in the current action generation and ignore the global information carried by previous actions. The paper reports that this myopia causes "a reduction of more than 20% in the success rate and limited speedup in some scenarios." The baseline of comparison is the high-performing OpenVLA-OFT model.

Method

SpecPrune-VLA is a training-free pruning method built on the insight that information across consecutive actions is highly similar, so token selection should combine local (current-action) and global (previous-action) cues. It has three components:

  • Static token pruning at the action level. Token redundancy is assessed using global information from previous actions together with local information from the current generation, statically reducing the number of visual tokens for the whole action.
  • Dynamic token pruning at the layer level. Exploiting the relevance between tokens and model layers, tokens are dynamically pruned according to their layer-specific importance.
  • Lightweight action-aware controller. Generated actions are categorized into coarse-grained vs. fine-grained based on speed; fine-grained actions are sensitive to pruning error, so a lightweight controller identifies the current action granularity and adjusts the pruning strategy accordingly.

Overview of the SpecPrune-VLA two-level pruning pipeline with the action-aware controller (Figure 2 from Author et al., 2026)

Results

Compared with OpenVLA-OFT on the LIBERO simulation benchmark, SpecPrune-VLA achieves an average 1.46× speedup on an NVIDIA A800 GPU and 1.57× speedup on an NVIDIA GeForce RTX 3090 GPU "with negligible loss on task success rate." The paper's headline result: "1.57x speedup in LIBERO simulation and 1.70x on real-world tasks".

Significance

As 2026 VLA models push toward real-time control, training-free acceleration that preserves success rate is highly valuable for deployment. SpecPrune-VLA's reuse of cross-action global context — rather than treating each action in isolation — addresses a concrete failure mode of prior token-pruning approaches and shows that meaningful speedups are attainable on commodity hardware (RTX 3090) as well as datacenter GPUs.

Links

← Back to ICML-2026