CVPR 2026 ACoT VLA - Heungwoo/research GitHub Wiki

ACoT-VLA โ€” Action Chain-of-Thought for Vision-Language-Action Models

Venue: CVPR 2026 Category: Reasoning / CoT VLA (action-space reasoning) Trend tag: Trend 4 Affiliations: AgiBot + Beihang University (BUAA) Authors: Linqing Zhong, Yi Liu, Yifei Wei, Ziyu Xiong, Maoqing Yao, Si Liu, Guanghui Ren (Si Liu, Guanghui Ren = corresponding)

Approach diagram

flowchart LR
  OBS["obs + instruction"] --> BB["VLA backbone"]
  BB --> EAR["Explicit Action Reasoner<br/>coarse trajectory intent"]
  BB --> IAR["Implicit Action Reasoner<br/>latent action patterns"]
  EAR --> FUSE["fuse"]
  IAR --> FUSE
  FUSE --> AH["action head"]
  AH --> ACT["fine action chunk"]
Loading

Problem

Chain-of-thought has been bolted onto VLAs by emitting text (CoT-VLA, ECoT) or visual frames (CoT-VLA, dVLA). Both add tokens to the inference path and require the model to think outside the action space before re-entering it. The question: can CoT happen entirely inside the action space?

Method

Two reasoners side-by-side:

  • Explicit Action Reasoner (EAR) โ€” a light-weight transformer that synthesizes coarse reference action trajectories as explicit reasoning steps, using self-attention over noisy action sequences plus cross-attention with the VLM's key-value cache.
  • Implicit Action Reasoner (IAR) โ€” extracts latent action priors via cross-attention with learnable query matrices over downsampled VLM representations from each VLM layer, aggregated by average pooling and an MLP projection.

Both reasoners' outputs undergo dual cross-attention with the noisy action query, are concatenated and passed through self-attention fusion, then condition the final action head, which emits a fine-grained action chunk. Reasoning never leaves the action manifold.

Results

  • LIBERO: 98.5% average success (Spatial 99.4 / Object 99.6 / Goal 98.8 / Long 96.0), first across all four suites; baseline 96.9%.
  • LIBERO-Plus: 86.6% zero-shot transfer, 88.0% under SFT.
  • VLABench: 63.5% Intention Score, 47.4% Progress Score across five tracks (+12.6% in unseen-texture conditions).
  • Real-world (AgiBot G1): 66.7% average over three tasks (Wipe Stain, Pour Water, Open-set Pick), beating baselines by ~5โ€“30 pts.

Compared against 30+ methods including Diffusion Policy, Octo, CoT-VLA, WorldVLA, DreamVLA, OpenVLA, ฯ€0, ฯ€0.5, SpatialVLA, UniVLA, GR00T-N1. Ablation: EAR alone +1.4%, IAR alone +1.2%, combined +1.6% over baseline on LIBERO.

Significance

ACoT-VLA's bet: the action space is rich enough to encode reasoning if you give it a coarse-to-fine decomposition. Sister paper to Fast-ThinkAct (which compresses text CoT into a latent), and predecessor-style to it (which compresses the same reasoning into the same space). Likely to influence how production VLAs structure their action heads.

Links

Related pages

โ† Back to CVPR-2026

โš ๏ธ **GitHub.com Fallback** โš ๏ธ