ICLR 2026 Reflective DDVLA - Heungwoo/research GitHub Wiki

ReflectDrive — Discrete Diffusion VLA for Driving

Venue: ICLR 2026 Category: VLA Architecture — Discrete diffusion / Autonomous driving Trend tag: Diffusion · Reflective inference · Safety constraints Affiliations: LiAuto, Tsinghua University (Li, Zheng, Wang, Wang, Zhao, Liu, Zhan, Zhan, Lang)

Approach diagram

flowchart LR
  Obs[Multi-view cameras<br/>front, front-left, front-right<br/>+ ego state + nav instruction] --> Tok[Trajectory tokenizer<br/>uniform 1D codebook per axis]
  Tok --> DLM[Discrete Diffusion LM<br/>LLaDA-V backbone, fine-tuned]
  DLM --> Goal[Goal-Conditioned Generation<br/>Top-K' goals, NMS dNMS = 0.9 m, K = 3]
  Goal --> Cand[K candidate trajectories]
  Cand --> Sg[Global Scorer Sglobal]
  Sg --> Best[Best initial trajectory τ*]
  Best --> Loop{Safety-Guided<br/>Regeneration Loop<br/>≤ 10 iterations}
  Loop -->|Safety Scorer Ssafe| Viol[Identify violating waypoints V]
  Viol -->|Local Scorer Slocal| Anchor[Manhattan local search δ ≤ 10]
  Anchor --> Inpaint[Diffusion inpainting<br/>around safe anchors]
  Inpaint --> Loop
  Loop -->|safe| Out[Final trajectory]
Loading

Problem

End-to-end driving stacks struggle to encode physical / safety rules during training. Existing fixes are either (1) rule-based post-processing that demands human engineering, (2) RL with simulation-only training and unsafe online rollouts, or (3) gradient-based diffusion guidance that is computationally expensive and brittle around safety constraints. The paper targets a gradient-free way to inject safety into a learned VLA-style planner.

Detailed Method

1. Trajectory discretization

Each 2D waypoint (x, y) is quantized independently with a uniform 1D codebook A = {a₁, a₂, …} over spatial range [-M, M] = [-100, 100] m (Table 4) at resolution Δg; the paper defines Δg symbolically but does not report a numeric value for it or the resulting codebook cardinality. A trajectory of N waypoints is flattened into a length-2N token sequence

y = Q(τ) = (y1,x, y1,y, ..., yN,x, yN,y) ∈ A^{2N}

This makes BEV-space search efficient and supports the inpainting story of discrete diffusion.

2. Discrete diffusion language model

Forward process masks tokens with a cosine-style noise schedule. Reverse process is trained with the standard masked-token NLL

L(θ) = E[ -Σ_{i: m_i^(s) = 1} log p_θ(y_i | ỹ^(s), c, s) ]

conditioned on context c = three camera views + ego state + language navigation command. Backbone is LLaDA-V (You et al., 2025), a pretrained Diffusion Language Model. Training uses classifier-free guidance.

Training hyperparameters (Table 4):

Parameter Value
Dataset NAVSIM training split (Table 4 does not name the split or sample count)
Batch size 16
Gradient accumulation 1
Learning rate 1 × 10⁻⁵
Epochs 3
Max context length 8192
LR scheduler Cosine, warmup ratio 0.03
Weight decay 0.0
Precision bfloat16

Inference config (Table 3): 5 diffusion steps, answer length 32, block length 32, low-confidence remasking, K = 3 goal candidates, NMS distance threshold 0.9 m, max refinement iterations 10.

3. Reflective Inference

Three scoring functions (implementation in Appendix C):

  • Sglobal(τ): overall trajectory quality including safety + coherence; returns 0 if any critical rule is violated.
  • Ssafe(τ): safety oracle that scores each waypoint by worst-violation-in-local-window.
  • Slocal(ax, ay): per-token-pair scorer used during local search.

Stage 1 — Goal-Conditioned Generation. The model produces a distribution over the terminal waypoint pθ(yN | c, s). Top-K' goal candidates are extracted, then Non-Maximum Suppression with threshold dNMS = 0.9 m returns K = 3 spatially diverse goals G = {G1, G2, G3}. For each goal Gk, a full trajectory is generated by inpainting the intermediate waypoints conditioned on Gk. The K trajectories are scored with Sglobal; the highest-scoring τ* proceeds to Stage 2.

Stage 2 — Safety-Guided Regeneration. Iterative loop:

  1. Ssafe identifies violating waypoint indices V = {t | Ssafe(τ*)t < τ_safe}.
  2. For each t ∈ T* ⊆ V, search the Manhattan neighborhood Nδ (δ ≤ 10 tokens) for the corrected pair maximizing Slocal — these become safety anchors.
  3. Diffusion inpainting regenerates the surrounding tokens conditioned on the safety anchors.
  4. Repeat until V is empty or budget = 10 iterations is hit; fall back to highest-Ssafe candidate otherwise.

Empirically most violations are resolved in 1–3 iterations (the paper does not report a numeric average).

Comprehensive Results

NAVSIM closed-loop (Table 1)

PDMS aggregates NC (no-collision), DAC (drivable-area compliance), TTC (time-to-collision), Comfort, EP (ego progress).

Method Paradigm Input NC↑ DAC↑ TTC↑ Comf.↑ EP↑ PDMS↑
UniAD E2E Cam 97.8 91.9 92.9 100.0 78.8 83.4
PARA-Drive E2E Cam 97.9 92.4 93.0 99.8 79.3 84.0
Transfuser E2E C+L 97.7 92.8 92.8 100.0 79.2 84.0
Hydra-MDP Augmented C+L 98.3 96.0 94.6 100.0 78.7 86.5
DiffusionDrive Diffusion C+L 98.2 96.2 94.7 100.0 82.2 88.1
GoalFlow Diffusion C+L 98.4 98.3 94.6 100.0 85.0 90.3
AutoVLA (Post-RFT) Autoregressive VLA Cam 98.4 95.6 98.0 99.9 81.9 89.1
ReflectDrive (w/o R.I.) Discrete Diffusion Cam 96.9 95.4 92.2 100.0 79.0 84.8
ReflectDrive (Ours) Discrete Diffusion Cam 97.7 99.3 93.5 100.0 86.9 91.1
ReflectDrive† (GT agents) Discrete Diffusion Cam 99.7 99.5 99.1 99.9 88.9 94.7
Human – – 100.0 100.0 100.0 99.9 87.5 94.8

Camera-only ReflectDrive surpasses Camera+LiDAR Augmented planners on PDMS, with the largest margin on DAC = 99.3 (drivable-area compliance), where reflection injects hard safety. With ground-truth-agent oracle (†), the system matches human performance within noise.

Reflection improves over the no-RI baseline by +3.9 DAC, +1.3 TTC, +0.8 NC, +7.9 EP → +6.3 PDMS.

Ablation: components of Reflective Inference (Table 2)

Goal-Cond. Safety-Guided NC↑ DAC↑ TTC↑ Comf.↑ EP↑ PDMS↑
✗ ✗ 96.9 95.4 92.2 100.0 79.0 84.8
✓ ✗ 96.6 96.5 91.5 100.0 83.8 87.4
✗ ✓ 98.1 98.9 94.8 99.9 84.1 90.3
✓ ✓ 97.7 99.3 93.5 99.9 86.9 91.1

Goal-Conditioned generation primarily lifts EP (progress); Safety-Guided primarily lifts DAC and TTC. They are complementary.

Other ablations

  • Generation steps (Fig. 4a). Non-monotonic: PDMS peaks at 5 steps, then declines.
  • Goal candidates K and NMS dNMS (Fig. 4b). Multi-modal goal generation gives a wider option set; scaling K beyond 3 yields diminishing returns.
  • Exploration steps and max iterations (Fig. 4c). PDMS scales positively with both → "inference-time scaling" exists for discrete-diffusion planners.

Limitations (as stated by authors, Appendix D)

  1. Single-frame inputs. Three-view images of the current frame only; velocity / motion of surrounding agents is unobserved. Future: incorporate history and predict obstacle trajectories jointly.
  2. Reflection design.
    • Goal-Conditioned Generation reuses the PDM scorer without task-specific tuning of high-level navigation/efficiency objectives.
    • Safety-Guided Regeneration: more iterations don't always help. Failure modes:
      • Boundary oscillation in narrow drivable space (discrete-token rounding error compounds).
      • Navigation correctness not encoded in the reward.
      • Goal-point selection is suboptimal in some scenarios; constrained search range can't recover.
  3. Sample efficiency. Limited engineering optimization so far; the authors note substantial room to improve inference efficiency (no concrete latency figures are reported in the paper).

Significance & Positioning

Provenance note (important): This is the only "discrete-diffusion VLA" actually at ICLR 2026, and it lives in the autonomous driving family — distinct from the manipulation-focused DDVLA paper at NeurIPS 2025 which the community sometimes conflates with it. The contribution is not "discrete diffusion for VLA" per se; it is the gradient-free reflection / inpainting loop that uses the discrete-token structure to perform local search and regeneration around safety anchors.

Compared to neighbors:

  • AutoVLA (Post-RFT): same VLA-for-driving paradigm, autoregressive instead of discrete-diffusion, slightly weaker on PDMS (89.1 vs 91.1) and notably weaker on EP (81.9 vs 86.9). Reflection lets ReflectDrive correct without RL fine-tuning.
  • Hydra-MDP / DiffusionDrive / GoalFlow: rely on trajectory anchors or rule-based guidance and Camera+LiDAR. ReflectDrive matches/beats them with camera only.
  • Diffusion Planner / classifier-guided diffusion: gradient-based guidance during denoising — expensive and parameter-sensitive. ReflectDrive's discrete inpainting is gradient-free.
  • Manipulation discrete-diffusion VLAs (DDVLA NeurIPS 2025): shares the discrete-action-codebook idea but optimizes for sampling efficiency, not safety constraint enforcement.

The broader bet is that discrete action spaces + masked diffusion make it natural to compose learned generators with rule-based search — the same recipe that worked for code (FIM, infilling) now applied to safety-critical planning.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️