ICML 2026 Characterizing Vision Language Action Models across XPUs - Heungwoo/research GitHub Wiki
Characterizing VLA Models across XPUs — model-hardware co-characterization and acceleration for on-robot deployment
Venue: ICML 2026 (Poster) Category: Efficiency Affiliations: Kaijun Zhou, Qiwei Chen, Da Peng, Zhiyang Li, Xijun Li, Jinyu Gu — School of Computer Science, Shanghai Jiao Tong University Traction (2026-06): 0 citations (arXiv)

Problem
VLA models are promising for generalist robot control, but on-robot deployment is bottlenecked by real-time inference under tight cost and energy budgets. Most prior evaluations use desktop-grade GPUs, which obscures the trade-offs and opportunities of heterogeneous edge accelerators (GPUs/XPUs/NPUs). There is no systematic framework that pairs VLA models to the right-sized hardware or that exploits VLA-specific computational structure for acceleration.
Method
The work delivers three contributions. (1) Cross-accelerator leaderboard. Model-hardware pairs are evaluated under a CET metric (Cost, Energy, Time), showing that right-sized edge devices can beat flagship GPUs on cost-/energy-efficiency while still meeting control-rate constraints. (2) Dual-phase characterization. Fine-grained Nsight profiling of π0 on RTX 4090, AGX Orin, and Jetson Thor reveals a consistent two-phase inference pattern — a compute-bound VLM backbone (SM utilization typically >90%) followed by a memory-bound Action Expert (SM utilization 20–40%). The VLM achieves ~3x higher hardware utilization than the Action Expert, yet the Action Expert dominates latency (~2x that of the VLM), so high-performance compute units idle during the inefficient phase (Finding #3). (3) Two accelerators guided by these insights:
- DP-Cache (Diffusion Policy Cache): exploits temporal redundancy in the diffusion trajectory. A stable segment (relative L1 distance between consecutive diffusion steps stays low, ~steps 20–80, fixed via offline profiling) lets the cached result be broadcast to skip redundant steps. Hyperparameter
Ssets the cache stride. - V-AEFusion (VLM-Action Expert pipeline parallelism): pipelines the two phases. At timestep
t, while the VLM processesObs_t, the Action Expert concurrently runs early denoising "stale steps" on the KV cache fromObs_{t-1}, then switches to the freshObs_tcache for the final refinement steps. High temporal coherence under closed-loop control (the arm pose changes little within one action chunk) justifies using stale features without overshooting.

Results
Headline: "up to 2.9x speedup on GPUs and 3.3x on edge NPUs with only marginal success degradation". DP-Cache on RTX 4090 (Table 4) gives 1.89x speedup at S=4 (378→200 ms, 2.6→5.0 Hz) with LIBERO success improving on Spatial (75.4→78.9%) and Long (74.1→76.9%) while Object dips slightly (52.4→51.4%); at S=8 it reaches 2.09x (181 ms, 5.6 Hz) with a larger accuracy cost. Combined with V-AEFusion's pipeline parallelism, the orthogonal techniques fuse to the headline end-to-end speedups across the GPU/NPU tiers. An example leaderboard is hosted at vla-leaderboard-01.vercel.app.
Significance
This is one of the first systematic, hardware-aware studies of where and how to run VLA policies on edge accelerators rather than data-center GPUs. The dual-phase (compute-bound VLM + memory-bound Action Expert) finding is a reusable lens, and DP-Cache + V-AEFusion show that VLA-specific temporal redundancy can be cashed in for 2–3x speedups without retraining — directly relevant to low-cost, energy-constrained robot deployment.
Links
- arXiv: 2604.24447
- ICML 2026: https://icml.cc/virtual/2026/poster/65223
← Back to ICML-2026