ICLR 2026 DIVA Fast dVLA - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Category: VLA Architecture — Efficiency Trend tag: Trend 1
flowchart LR
subgraph STU[Selective Token Unmasking baseline]
N0[Mask] --> N1[Unmask 1 highest-conf token] --> N2[...] --> Nout[Out, fragmented]
end
subgraph DIVA
D0[Mask] --> D1[Selective Group Unmasking<br/>unmask whole group by aggregate conf]
D1 --> Dout[Out, coherent chunks]
end
subgraph FastdVLA[Fast-dVLA]
F0[Block-causal attention + KV reuse] --> Fopt[Diffusion forcing<br/>per-block noise + asymmetric distillation]
Fopt --> Fout[Pipelined parallel decoding, 30 Hz]
end
Discrete-diffusion VLAs (Discrete Diffusion VLA, Unified Diffusion VLA) match or beat flow matching on accuracy. The two papers here attack different gaps: DIVA targets action coherence in the discrete-diffusion decode (the naive confidence-first token-by-token unmask fragments action sequences); Fast-dVLA targets wall-clock latency, since iterative dVLA decoding runs far below the ~30 Hz real-time requirement of physical robots.
Note: these are two distinct papers, NOT a complementary DIVA+Fast-dVLA acceleration pair. DIVA is a base discrete-diffusion VLA architecture; Fast-dVLA is an inference-acceleration method built on existing dVLA bases (UD-VLA / DD-VLA).
DIVA (ICLR 2026 submission) — a discrete-diffusion VLA built on OpenVLA with three designs: (1) a learnable discrete action tokenization bridging continuous actions to the multimodal token space; (2) a latent-driven policy learning strategy that aligns the VLA backbone and the policy head via joint (multi-objective) optimization; (3) a Selective Group Unmasking (SGU) decoding strategy that unmasks the highest-aggregate-confidence group of action tokens per step (vs Selective Token Unmasking which refines one token at a time), preserving spatiotemporal coherence. DIVA optimizes for accuracy/coherence, not latency.
Fast-dVLA (arXiv 2603.25661) — an acceleration method that constrains attention to causal action blocks so a fully decoded block's KV cache is frozen and reused; applies diffusion forcing (progressively increasing per-block noise so earlier blocks finish first while later blocks refine in parallel); uses asymmetric distillation to transfer a bidirectional teacher into the block-wise student; and a pipelined parallel decoding algorithm for real-time control.
- DIVA: LIBERO success rates 98.0 / 98.8 / 97.6 / 95.2% (Spatial/Object/Goal/Long), avg 97.4% (+2.0% on LIBERO-Long over second best); real-world avg 45% → 60% vs π0 baseline.
- Fast-dVLA: 2.8×–4.1× speedup while preserving action performance; 2.8× over UD-VLA on long-horizon CALVIN at 186.7 tokens/s; LIBERO avg 96.6%, LIBERO-Long 92.0% → 92.8%; SimplerEnv 366.4 tokens/s; stable 30 Hz real-world control. Baselines include UD-VLA, DD-VLA, GR00T-N1, π0 (not π0.6).
DIVA shows that group-level unmasking, not token-by-token confidence-first decoding, is the right primitive for coherent action chunks under discrete diffusion. Fast-dVLA removes the standing latency objection to discrete-diffusion VLAs, reaching 30 Hz parity with optimized continuous-action policies while staying competitive with frontier AR and flow-matching VLAs — strengthening the single-transformer unification argument on cost grounds.
- Fast-dVLA arXiv: https://arxiv.org/abs/2603.25661 — project page: https://chris1220313648.github.io/Fast-dVLA/
- DIVA (ICLR 2026): https://openreview.net/forum?id=mNya9d1DA2
- Discrete Diffusion VLA
- Unified Diffusion VLA (UD-VLA — Fast-dVLA base/baseline)
← Back to ICLR-2026