CoRL 2025 DexVLA - Heungwoo/research GitHub Wiki

DexVLA โ€” Plug-in Diffusion Action Expert for VLAs

Venue: CoRL 2025 ยท arXiv: 2502.05855 Category: VLA Architecture Trend tag: Separate action expert (the pattern ICLR 2026 challenges)

Approach diagram

flowchart LR
  V[Vision] --> VLM[Qwen2-VL 2B]
  L[Language] --> VLM
  VLM -. features .-> DE[Plug-in ScaleDP<br/>~1B Diffusion<br/>Action Expert]
  noise --> DE
  DE --> A[Action chunk<br/>cross-embodiment]
Loading

Problem

VLAs have generally either (a) produced actions autoregressively from the VLM directly (slow at scale) or (b) used small MLP/diffusion heads (no generative capacity). A large, expressive action generator decoupled from the VLM could marry both.

Method

DexVLA attaches a ~1B-parameter diffusion action expert (ScaleDP โ€” a transformer-based scaling of Diffusion Policy: 32 layers, hidden 1280, 16 heads) as a plug-in module to a pretrained Qwen2-VL 2B backbone. The expert is trained across robot embodiments (single-arm, bimanual, dexterous hands), so a single expert drives different hardware. A smaller 410M ScaleDP variant is offered to save memory.

Training uses a three-stage embodiment curriculum:

  1. Cross-embodiment pre-training of the diffusion expert alone on diverse robot data (using DistilBERT + ViT for conditioning, no VLM).
  2. Embodiment-specific alignment โ€” jointly train VLM, projection layers, and diffusion expert on single-embodiment data, with the VLM's visual encoder frozen.
  3. Task-specific adaptation on downstream demonstrations, with sub-step language reasoning generated as output.

Results

Strong cross-embodiment performance on long-horizon, dexterous tasks (laundry folding from crumpled clothing, dryer unloading, bin picking, table bussing, drink pouring with a dexterous hand) using only direct language prompting. Reported scores include 0.92 on shirt folding and 0.90 on a novel embodiment with 100 demonstrations, substantially exceeding Octo, OpenVLA, and Diffusion Policy. Establishes that billion-scale action generators are feasible and beneficial.

Significance

Canonical example of the "VLM + separate action expert" pattern โ€” the same architecture ฯ€0/ฯ€0.5/ฯ€0.6/ฯ€0.7 use. Exactly this pattern is what ICLR 2026's discrete-diffusion VLAs explicitly challenge by unifying action generation back into the same transformer as the VLM:

Links

Related pages

โ† Back to CoRL-2025

โš ๏ธ **GitHub.com Fallback** โš ๏ธ