ICLR 2026 AutoQVLA - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 (poster) Category: VLA Architecture — Efficiency Trend tag: Efficiency Paper: "QVLA: Not All Channels Are Equal in Vision-Language-Action Model's Quantization" (arXiv 2602.03782). AutoQVLA is the name the paper gives to the automated bit-allocation algorithm; the framework itself is QVLA. Authors: Yuhao Xu, Yantai Yang, Zhenyang Fan, Yufan Liu, Yuming Li, Bing Li, Zhipeng Zhang (Shanghai Jiao Tong University, Chinese Academy of Sciences, Alipay)
flowchart LR
W[VLA weights] --> S[Per-channel<br/>action-space sensitivity<br/>Jacobian / Taylor]
S --> G[Greedy demotion<br/>bits in 0,2,4,8,16]
G --> H[High-sensitivity channels]
G --> L[Low-sensitivity channels]
H --> HK[Keep in high precision]
L --> LQ[Aggressively quantize<br/>0-bit = prune]
HK --> M[Mixed-precision VLA<br/>29.2% VRAM, 98.9% perf, 1.49x]
LQ --> M
LLM-derived quantization (e.g., SmoothQuant, AWQ, OmniQuant) optimizes for data fidelity — minimizing weight/activation reconstruction error — but ignores that in a VLA, minor action deviations compound into catastrophic task failure. Uniform-bit quantization is therefore mismatched to embodied control. The paper's analysis shows significant intra-layer channel heterogeneity: individual channels contribute very unequally to the final action output, so treating them identically is wasteful.
Channel-wise mixed-precision bit allocation guided by action-space sensitivity:
- For each weight channel, the impact on the final action output is estimated via a first-order Taylor / Jacobian approximation, yielding a per-channel importance score in action space (not weight-reconstruction space).
- A greedy demotion algorithm starts every channel at full precision and iteratively lowers the least-sensitive channels through the bit-width set {0, 2, 4, 8, 16} until the memory budget is met.
- 0-bit = pruning, so quantization and structured pruning are unified in one framework. Activations are kept at a uniform bit-width (e.g., 8-bit) for hardware efficiency.
On LIBERO (4 task suites), base model OpenVLA-OFT, W4A4 setting:
- VRAM drops to 29.2% of the original (4.5 GB vs 15.4 GB → 70.8% reduction) while retaining 98.9% of original performance, with a 1.49× speedup. W8A8 gives a −0.7% drop at 7.2 GB / 1.36×.
- Strongly beats LLM-derived baselines at equal budget: at W4A4, QVLA −0.5% vs SmoothQuant −13.3%; at W4A16, QVLA −0.4% vs AWQ −4.5%; vs OmniQuant −3.2%.
- Ablations: channel-wise > layer-wise (76.5% vs 74.8% at INT4); enabling 0-bit pruning matches/exceeds no-pruning while lowering VRAM (76.8% @ 7.0 GB vs uniform-8bit 74.6%).
- Validated beyond LIBERO: UniVLA-7B (W4A16) 95.1% vs AWQ 92.6%; real-robot π₀ retains its 63.3% success rate (W8A16, 1.28× speedup).
Reframes VLA compression around action-centric rather than data-fidelity error, the first systematic quantization study tailored to VLAs. Makes larger VLAs deployable on commodity GPUs / on-device. The per-channel action-sensitivity map is itself a reusable artifact for future efficient-VLA work.
- arXiv: https://arxiv.org/abs/2602.03782
- OpenReview: https://openreview.net/forum?id=TpL2nXanru
- Code: https://github.com/AutoLab-SAI-SJTU/QVLA
← Back to ICLR-2026