ICLR 2026 Action Aware Pruning - Heungwoo/research GitHub Wiki

Action-aware Dynamic Pruning (ADP) โ€” Trajectory-Gated Token Pruning for VLA

Venue: ICLR 2026 Category: VLA Architecture โ€” Efficiency Trend tag: Efficiency

Approach diagram

flowchart LR
  Img[Visual tokens] --> TXT[Text-driven token selection]
  Hist[Recent action history] --> GATE[Action-aware<br/>trajectory gate]
  TXT --> GATE
  GATE --> KEEP[Adaptive keep ratio<br/>coarse: prune more ยท fine: prune less]
  KEEP --> POL[VLA policy]
Loading

Problem

Static token pruning ignores that visual redundancy varies across manipulation phases: it is high during coarse approach motions and low during fine-grained contact phases. A fixed keep-ratio either over-prunes during precision steps or under-prunes during transit, capping the achievable speed/accuracy frontier.

Method

ADP combines two signals:

  • Text-driven token selection filters visual tokens by relevance to the language instruction.
  • Action-aware trajectory gating uses recent motion history to dynamically tune the keep ratio โ€” pruning aggressively during coarse motion, retaining detail during precise operations. The gating signal is a windowed trajectory distance computed from end-effector pose changes.

The result is a phase-adaptive pruning schedule that is training-free / plug-and-play, layered onto an existing VLA without retraining the backbone.

Results

  • 1.35ร— speedup on OpenVLA-OFT (at a 30% visual-token keep ratio, FLOPs โ†“ to 5.85).
  • LIBERO per-suite success at 30% keep ratio on OpenVLA-OFT: Spatial 97.6 / Object 98.4 / Goal 97.4 / Long 84.2 (โ‰ˆ94.4% average), competitive with the unpruned baseline.
  • +25.8% success-rate improvement with OpenVLA vs. its baseline, in addition to the latency/FLOPs savings.
  • Reduces FLOPs and action-inference latency while maintaining competitive success rate vs. baselines on LIBERO suites and real-world tasks.

Significance

Reframes token pruning as a task-phase-conditioned problem rather than a static compression problem, exploiting an empirical structure (coarse vs. fine phases have different visual demands). Plug-in compatible with existing policies, complementing weight-side acceleration like AutoQVLA and joint scheduling work like SP-VLA.

Links

Related pages

โ† Back to ICLR-2026

โš ๏ธ **GitHub.com Fallback** โš ๏ธ