ICLR 2026 DIVA Fast dVLA - Heungwoo/research GitHub Wiki

DIVA & Fast-dVLA — Accelerating Discrete-Diffusion VLAs

Venue: ICLR 2026 Category: VLA Architecture — Efficiency Trend tag: Trend 1

Approach diagram

flowchart LR
  subgraph STU[Selective Token Unmasking baseline]
    N0[Mask] --> N1[Unmask 1 highest-conf token] --> N2[...] --> Nout[Out, fragmented]
  end
  subgraph DIVA
    D0[Mask] --> D1[Selective Group Unmasking<br/>unmask whole group by aggregate conf]
    D1 --> Dout[Out, coherent chunks]
  end
  subgraph FastdVLA[Fast-dVLA]
    F0[Block-causal attention + KV reuse] --> Fopt[Diffusion forcing<br/>per-block noise + asymmetric distillation]
    Fopt --> Fout[Pipelined parallel decoding, 30 Hz]
  end
Loading

Problem

Discrete-diffusion VLAs (Discrete Diffusion VLA, Unified Diffusion VLA) match or beat flow matching on accuracy. The two papers here attack different gaps: DIVA targets action coherence in the discrete-diffusion decode (the naive confidence-first token-by-token unmask fragments action sequences); Fast-dVLA targets wall-clock latency, since iterative dVLA decoding runs far below the ~30 Hz real-time requirement of physical robots.

Note: these are two distinct papers, NOT a complementary DIVA+Fast-dVLA acceleration pair. DIVA is a base discrete-diffusion VLA architecture; Fast-dVLA is an inference-acceleration method built on existing dVLA bases (UD-VLA / DD-VLA).

Method

DIVA (ICLR 2026 submission) — a discrete-diffusion VLA built on OpenVLA with three designs: (1) a learnable discrete action tokenization bridging continuous actions to the multimodal token space; (2) a latent-driven policy learning strategy that aligns the VLA backbone and the policy head via joint (multi-objective) optimization; (3) a Selective Group Unmasking (SGU) decoding strategy that unmasks the highest-aggregate-confidence group of action tokens per step (vs Selective Token Unmasking which refines one token at a time), preserving spatiotemporal coherence. DIVA optimizes for accuracy/coherence, not latency.

Fast-dVLA (arXiv 2603.25661) — an acceleration method that constrains attention to causal action blocks so a fully decoded block's KV cache is frozen and reused; applies diffusion forcing (progressively increasing per-block noise so earlier blocks finish first while later blocks refine in parallel); uses asymmetric distillation to transfer a bidirectional teacher into the block-wise student; and a pipelined parallel decoding algorithm for real-time control.

Results

  • DIVA: LIBERO success rates 98.0 / 98.8 / 97.6 / 95.2% (Spatial/Object/Goal/Long), avg 97.4% (+2.0% on LIBERO-Long over second best); real-world avg 45% → 60% vs π0 baseline.
  • Fast-dVLA: 2.8×–4.1× speedup while preserving action performance; 2.8× over UD-VLA on long-horizon CALVIN at 186.7 tokens/s; LIBERO avg 96.6%, LIBERO-Long 92.0% → 92.8%; SimplerEnv 366.4 tokens/s; stable 30 Hz real-world control. Baselines include UD-VLA, DD-VLA, GR00T-N1, π0 (not π0.6).

Significance

DIVA shows that group-level unmasking, not token-by-token confidence-first decoding, is the right primitive for coherent action chunks under discrete diffusion. Fast-dVLA removes the standing latency objection to discrete-diffusion VLAs, reaching 30 Hz parity with optimized continuous-action policies while staying competitive with frontier AR and flow-matching VLAs — strengthening the single-transformer unification argument on cost grounds.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️