CVPR 2026 AnchorVLA - Heungwoo/research GitHub Wiki

AnchorVLA β€” Anchored Diffusion for Efficient End-to-End Mobile Manipulation

Venue: CVPR 2026 (likely β€” pending CVF virtual-page confirmation) Category: Diffusion Policy / Mobile Manipulation Trend tag: Trend 1 Affiliations: UQMM Lab, University of Queensland + Robotics & Autonomous Systems Group, CSIRO (Brisbane) Authors: Jia Syuen Lim, Zhizhen Zhang, Peter Bohm, Brendan Tidd, Zi Huang, Yadan Luo

Approach diagram

flowchart LR
  OBS["observation"] --> BB["VLA-Adapter backbone<br/>(Qwen-2.5 0.5B, LoRA r=64)"]
  VOCAB["anchor vocabulary<br/>M=20 (K-Means, offline)"] --> DEN
  BB --> DEN["anchored diffusion<br/>perturb all anchors,<br/>truncated schedule S_tr=10"]
  DEN --> SCORE["anchor scoring head<br/>pick highest-score chunk"]
  SCORE --> RES["residual self-correction<br/>module (per-step MLP)"]
  RES --> OUT["final action"]
Loading

Problem

Mobile manipulation needs a policy that keeps multiple viable action options (multimodality) while staying reactive during execution. Diffusion policies model multimodal action distributions well, but full iterative denoising is too expensive at control time on compute-constrained mobile platforms, and action chunking trades reactivity for efficiency.

Method

Backbone: VLA-Adapter with a Qwen-2.5 (0.5B) language model, LoRA fine-tuned (rank 64) on the frozen VLM (~726M total params). Only the action head differs from the original VLA-Adapter.

Anchored diffusion:

  • Build an anchor vocabulary offline: segment training trajectories into fixed-length action chunks (horizon H) and K-Means cluster them into M=20 representative anchor trajectories. (This is not per-observation nearest-neighbor retrieval β€” the anchor set is fixed.)
  • At inference, perturb all M anchors (noise added to the anchor, not pure Gaussian noise: A = βˆšαΎ±Β·Δ€ + √(1βˆ’αΎ±)Β·Ξ΅) and denoise them in parallel with a truncated schedule (S_tr = 10 vs. 50 for a full baseline) β€” fewer steps suffice because each starting point is already close to data.
  • A learned anchor scoring head predicts a confidence score per denoised chunk and selects the highest-scoring one.
  • A lightweight residual self-correction module (~57K-param MLP) makes per-step, high-frequency micro-adjustments during chunk execution to counter chunking-induced drift.

Result: full action chunk in a fraction of standard diffusion-policy steps.

Results

  • ManiSkill-HAB (6 tasks): 64.0% avg success at H=2, 61.5% at H=5, vs. AC-DiT baseline 55.6%.
  • Real-world (Unitree Go2 + SO101, 2 tasks): 40% success vs. SmolVLA 15%.
  • Efficiency: ~89.8 Hz at H=5; ~80% compute reduction vs. per-step (H=1) operation.
  • Ablations: anchored prior is decisive at low step counts β€” 42.9% vs. 0% success at 10 denoising steps with vs. without the anchor; residual module adds 42.9% vs. 33.6%.

Significance

The "anchored diffusion" trick is generic β€” instead of denoising every action chunk from pure Gaussian noise, AnchorVLA pre-clusters demonstrations into a small fixed anchor vocabulary and denoises locally around those anchors with a truncated schedule, then lets a scoring head pick the best. This preserves the multimodality of diffusion while collapsing inference cost, and the per-step residual head recovers reactivity lost to action chunking. The recipe is likely to migrate from mobile manipulation to other latency-constrained VLA settings. Note this is anchor-vocabulary selection (scored among M=20 candidates), not per-observation nearest-neighbor retrieval.

Links

Related pages

← Back to CVPR-2026

⚠️ **GitHub.com Fallback** ⚠️