ICLR 2026 Discrete Diffusion VLA - Heungwoo/research GitHub Wiki

Discrete Diffusion VLA β€” Bringing Discrete Diffusion to Action Decoding

Venue: ICLR 2026 Β· arXiv: 2508.20072 Category: VLA Architecture Trend tag: Trend 1

Approach diagram

flowchart LR
  V[Vision tokens] --> T[Single Transformer<br/>SAME as VLM]
  L[Language tokens] --> T
  M[Masked action tokens] --> T
  T --> P{Prediction<br/>cross-entropy}
  P -- unmask high-confidence first --> R1[Round 1 unmasking]
  R1 --> RM[Secondary re-masking<br/>of low-confidence]
  RM --> Done[Final action chunk]
Loading

See Figure 1 of the original paper (link below) for the authors' architecture diagram with adaptive ordering.

Problem

Existing non-AR VLA decoders β€” most notably flow matching in the Ο€0 family β€” attach a separate continuous action expert outside the VLM backbone. This causes:

  • Parameter / interface mismatch between the VLM (token-trained) and the action head (vector-field-trained)
  • A fixed denoising trajectory: no adaptive decoding order
  • No error correction: once a denoising step commits, the model cannot revisit

Method

Single-transformer policy. Built on the OpenVLA architecture (Prismatic-7B VLM with SigLIP+DINOv2 visual encoders and a Llama-2 backbone); the original autoregressive causal-attention backbone is converted into a bidirectional transformer so discrete diffusion over actions runs inside the same VLM. Continuous controls (translation, rotation, gripper) are discretized with a 256-bin quantile-based scheme (1st–99th percentile) β€” 7 tokens per timestep β€” and grouped into chunks. These tokens are generated by masked discrete diffusion inside the same transformer as the VLM, trained with the VLM's own cross-entropy objective. At inference it runs a small fixed number of refinement rounds (12 by default) with a cosine mask schedule. Two tricks:

  • Adaptive decoding order β€” confident tokens unmasked first, hard tokens later
  • Secondary re-masking β€” revisit uncertain earlier predictions across refinement rounds

Results

  • LIBERO: 96.3% average success (+0.9% vs OpenVLA-OFT discrete)
  • SimplerEnv-Fractal: 71.2% visual matching / 56.9% variant aggregation β†’ 64.1% overall
  • SimplerEnv-Bridge: 54.2% overall (+14.7% over Ο€0, +6.4% over Ο€0-FAST) Beats AR, MLP-head, and continuous-diffusion baselines under matched training conditions, while using fewer NFEs than AR.

Significance

Crystallizes the discrete-diffusion VLA design philosophy: unify action generation into the VLM transformer, reuse the VLM training objective, and inherit the entire token-based decoding ecosystem (MaskGIT-style sampling, dLLM tricks). The most compelling 2026 architectural alternative to flow-matching action experts.

Links

πŸ“– In-depth review

For a long-form review with the paper's architecture figure, full LIBERO (96.3%) + SimplerEnv (64.1% Fractal / 54.2% Bridge) numbers, all ablations (adaptive ordering, re-masking, temperature schedule, OOD robustness), and limitations: In-Depth Review of Discrete Diffusion VLA.

πŸ“– In-depth architecture review

For how discrete diffusion compares against AR / flow matching / continuous diffusion / world-model backbones / dual-system / reasoning-augmented / tokenizer-centric / small-efficient / 3D-tactile / hybrid approaches β€” 10+ categories across ~60 VLA papers: VLA Architectures Review.

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️