ICRA 2026 Token Pruning - Heungwoo/research GitHub Wiki
Venue: ICRA 2026 ยท Authors: Titong Jiang, Xuefeng Jiang, Yuan Ma, Xin Wen, Bailin Li, Kun Zhan, Peng Jia, Yahui Liu, Sheng Sun, Xianpeng Lang ยท arXiv: 2509.12594 Category: Efficient VLA (token pruning) Trend tag: Differentiable token pruning
The method is called LightVLA. Built on OpenVLA-OFT, it learns which visual tokens to keep โ pruning is driven by task performance rather than heuristics, so accuracy improves while compute drops.
flowchart LR
IMG[Input images] --> ViT[ViT encoder<br/>visual tokens H_v]
LANG[Language tokens H_l] --> Q
ViT --> Q["Query generation<br/>Q = softmax(H_v H_l^T/โD) H_l"]
Q --> S[Token scoring<br/>S = Q H_v^T / โD]
S --> GS[Gumbel-softmax selection<br/>hard argmax + soft path<br/>differentiable in backward]
GS --> KEEP[Keep informative tokens<br/>~78 avg vs full set]
KEEP --> LLM[VLA backbone<br/>OpenVLA-OFT]
LLM --> ACT[Robot action]
VLA inference is bottlenecked by attention over large sets of visual tokens. Prior visual-token pruning was designed for VLMs and, when ported to VLA, tends to underperform: it relies on heuristic keep-ratios and "magic numbers" that ignore which tokens actually matter for action execution.
- Dynamic queries are generated by cross-attention between visual and language tokens, then used to score every visual token's importance.
- Gumbel-softmax selection makes the discrete keep/drop decision differentiable end-to-end, combining a hard argmax (forward) with a soft path (backward), so the model learns an adaptive, input-dependent token budget during fine-tuning.
-
No magic numbers, no extra trainable parameters in the base variant โ keeping it compatible with standard inference frameworks. A
LightVLA*variant adds optional learnable query / LayerNorm weights.
On the LIBERO benchmark, built on OpenVLA-OFT:
- Success rate: 97.4% avg (Spatial 98.4 / Object 98.4 / Goal 98.2 / Long 94.6) vs OpenVLA-OFT 94.5% โ +2.9%.
- Efficiency: โ59.1% FLOPs, โ38.2% latency, retaining ~78 visual tokens on average.
- vs other pruning methods (success rate): FlashVLA 73.7%, SP-VLA 74.9%, VLA-Cache 74.7% โ all far below LightVLA's 97.4%, showing VLM-oriented pruners degrade on VLA tasks.
LightVLA inverts the usual efficiency/accuracy trade-off: by making pruning performance-driven and differentiable, it simultaneously cuts compute and raises success rate. It is the parameter-free, hyperparameter-free counterpart to scheduler- and quantization-based VLA efficiency lines, slotting cleanly into the small/efficient VLA category.
- arXiv: https://arxiv.org/abs/2509.12594
- Project page: https://liauto-research.github.io/LightVLA/
- Action-aware Dynamic Pruning โ token-level efficiency, heuristic-driven counterpart
- SP-VLA โ joint scheduling + spatio-semantic pruning
- VLA Architectures review โ small/efficient category
- ICRA 2026 Survey
โ Back to ICRA-2026