ICML 2026 EcoVLA - Heungwoo/research GitHub Wiki
EcoVLA — Environment-aware adaptive channel pruning for VLA inference
Venue: ICML 2026 (Poster) Category: Efficiency Affiliations: Yuting Huang, Leilei Ding, Zhipeng Tang, Zenghuan Zhu, Jiajun Deng, Xinrui Lin, Shuo Liu, Haojie Ren, Jianmin Ji, Yanyong Zhang Traction (2026-06): 1 citations (arXiv)

Problem
Vision-Language-Action (VLA) models have large parameter counts that produce substantial inference latency, the primary bottleneck for real-time manipulation. Parameter sparsification is a natural remedy, but as the environment evolves during VLA execution, the optimal sparsity pattern changes accordingly. Static pruning lacks the adaptability required for these dynamics, while fixed-interval dynamic layer pruning suffers from coarse granularity and high retraining overheads. Prior acceleration work has focused mostly on token pruning, leaving structured model pruning comparatively unexplored.
Method
EcoVLA is a training-free, plug-and-play adaptive pruning framework designed to combine orthogonally with existing VLA accelerators. It has two components:
- Environment-aware Adaptive Pruning (EAP): a lightweight adaptive channel pruning method. A lightweight environment-aware sparsity-variations predictor perceives real-time dynamics, and Temporal Consistency Pruning exploits the temporal consistency of VLA execution in physical environments. On triggering a sparsity update at frame t, EcoVLA runs a dense inference, computes instantaneous features from the current input, aggregates them with historical features, and applies the new sparsity pattern starting from frame t+1 for sparse inference.
- Interleaved Inference Orchestration (I²O): leverages the FLOPs "bubbles" inherent in VLA inference to schedule the pruning computation in parallel, hiding pruning overhead so the cost has negligible impact on latency.
- Hardware-efficient implementation: sparse efficient kernels plus dense-metric acceleration.

Results
"up to 1.60x speedup with only 0.4% drop in success rate (2.18x combined with token pruning)". Evaluated across three VLA models (OpenVLA-OFT, π0.5, CogACT) and two benchmarks (LIBERO, SIMPLER), EcoVLA achieves 1.6x speedup with only a 0.4% reduction in success rate. On OpenVLA-OFT/LIBERO it reaches 1.26x and 1.41x speedups at 25% and 40% pruning ratios, with success-rate losses of only 0.35% and 2.8% respectively, and notably outperforms the structured-pruning baseline Wanda on the pruning-sensitive LIBERO-Goal split. Integrating with token pruning, EcoVLA lifts speedup from 1.21x (FastV at 50% pruning) to 2.18x while recovering FastV's accuracy drop, narrowing the gap to the vanilla baseline to just 0.5%. The authors also validate EcoVLA on a real Kinova Gen3 robot platform.
Significance
EcoVLA targets an underexplored axis of VLA acceleration—structured channel pruning that adapts to environment dynamics—rather than the well-studied token-pruning axis. Because it is training-free and composes orthogonally with token pruners, it offers near-free latency reductions that stack with existing methods, an attractive property for deploying large VLAs on real-time robot hardware.
Links
- arXiv: 2602.00780
- ICML 2026: https://icml.cc/virtual/2026/poster/66030
← Back to ICML-2026