ICLR 2026 Unified Diffusion VLA - Heungwoo/research GitHub Wiki

Unified Diffusion VLA (UD-VLA) โ€” Joint Discrete Denoising Diffusion of Frames + Actions

Venue: ICLR 2026 (poster) ยท OpenReview: UvQOcw2oCD ยท arXiv: 2511.01718 Authors: Jiayi Chen, Wenxuan Song, Pengxiang Ding, Ziyang Zhou, Han Zhao, Feilong Tang, Donglin Wang, Haoang Li (HKUST-GZ, Westlake University, Zhejiang University, Monash University) Category: VLA Architecture Trend tag: Trend 1

Approach diagram

flowchart LR
  Cur[Current image tokens] --> T[Single Transformer<br/>hybrid attention<br/>intra-block bidir / inter-block causal]
  L[Text tokens] --> T
  Mf[Masked future image tokens] --> T
  Ma[Masked action tokens] --> T
  T --> Df[Denoised future frames]
  T --> Da[Denoised action chunk]
  Da -. acting block attends to .-> Df
Loading

Sequence layout: [ text ; current image ; future image ; action ], where text and current-image tokens are input and future-image and action tokens are the output to be denoised. Modalities are unified by discrete tokenization โ€” a VQ visual tokenizer for images and the FAST action tokenizer for actions โ€” over a pretrained Emu3 VLM backbone (originally autoregressive/causal).

Problem

Even with discrete diffusion action decoding (cf. Discrete Diffusion VLA), visual prediction and action prediction are typically separate subsystems. To reason about "what will the world look like after my action" in a single forward pass, you need to denoise future image tokens and action tokens together.

Method

Apply the Joint Discrete Denoising Diffusion Process (JD3P) to future image tokens and action tokens within the same synchronous denoising step. The key is a hybrid attention mask (Fig. 2): output tokens are split into a generation (future-image) block and an acting (action) block; within each block attention is bidirectional, while across blocks it is causal โ€” the generation block attends only to the input, the acting block attends to both the input and the future-image block, and no information flows backward into the input. This recasts end-to-end control as two coupled processes: (i) a foresight process predicting the next visual state and (ii) an inverse-kinematics process inferring actions conditioned on that prediction. At every denoising iteration all action tokens causally attend to all future image tokens, so actions are progressively refined under sufficient visual guidance.

Training is two-stage on the Emu3 VLM backbone: stage (i) post-trains on large-scale video ([text ; current image ; future image]) to inject world-model future-frame prediction; stage (ii) jointly trains image generation + action prediction on robot data. The objective is a single-step mask-predict cross-entropy (the explicit multi-step diffusion corruption chain is discarded), using a shift operation so the diffusion decoder retains the next-token-prediction capacity of the pretrained backbone. Inference uses parallel decoding with adaptive (cosine) masking plus prefix KV-cache, prefilled special tokens (<BOI>/<EOI>/<BOA>), confidence-guided decoding, and decoding-space mapping that restricts each modality to its codebook range.

Results

  • CALVIN ABCDโ†’D: avg length 4.64 (best in Table 2; vs MDT 4.52, UP-VLA 4.42, UniVLA* 4.26).
  • LIBERO: 92.7% average (Spatial 94.1 / Object 95.7 / Goal 91.2 / Long 89.6), SOTA-class alongside DreamVLA (92.6).
  • SimplerEnv-WidowX: 59.4% overall, outperforming all baselines (e.g. ฯ€0-FAST 48.3, F1 59.4 surpassed on overall, SpatialVLA 42.7).
  • Inference: ~4ร— faster than autoregressive decoding (JD3P 219.3 tok/s, 4.64 avg len vs AR 50.2 tok/s, 4.18; Jacobi 2ร—, independent-diffusion ID 2.9ร— but lower 4.35).
  • Predicted future frames double as interpretability output (you can see what the policy "imagines"); paper notes generation is task-level coherent but lacks fine pixel fidelity due to no large-scale generative pretraining and compressed image tokens.

Decisive ablations. Attention scheme: Hybrid 4.64 > Bidirectional 4.32 (leaks action info across modalities) > Causal 4.04. Visual-generation target: Future-image 4.64 > Current-image 4.39 > none 4.21 โ€” confirming foresight, not mere reconstruction, drives the gain.

Significance

Operationalizes "world-model-as-policy" inside a single network โ€” no separate learned dynamics model needed โ€” and is the first unified VLA to perform joint discrete diffusion decoding of both visual and action tokens at inference (prior unified VLAs like WorldVLA/UniVLA model both in post-training but decode only actions autoregressively at inference; see the paper's Table 1). The clearest demonstration that discrete diffusion's natural fit with image tokens enables tight visual-prediction โ†” control coupling, while also recovering the autoregressive speed gap via parallel masked decoding.

Links

Related pages

โ† Back to ICLR-2026

โš ๏ธ **GitHub.com Fallback** โš ๏ธ