CVPR 2026 ACoT VLA - Heungwoo/research GitHub Wiki
Venue: CVPR 2026 Category: Reasoning / CoT VLA (action-space reasoning) Trend tag: Trend 4 Affiliations: AgiBot + Beihang University (BUAA) Authors: Linqing Zhong, Yi Liu, Yifei Wei, Ziyu Xiong, Maoqing Yao, Si Liu, Guanghui Ren (Si Liu, Guanghui Ren = corresponding)
flowchart LR
OBS["obs + instruction"] --> BB["VLA backbone"]
BB --> EAR["Explicit Action Reasoner<br/>coarse trajectory intent"]
BB --> IAR["Implicit Action Reasoner<br/>latent action patterns"]
EAR --> FUSE["fuse"]
IAR --> FUSE
FUSE --> AH["action head"]
AH --> ACT["fine action chunk"]
Chain-of-thought has been bolted onto VLAs by emitting text (CoT-VLA, ECoT) or visual frames (CoT-VLA, dVLA). Both add tokens to the inference path and require the model to think outside the action space before re-entering it. The question: can CoT happen entirely inside the action space?
Two reasoners side-by-side:
- Explicit Action Reasoner (EAR) โ a light-weight transformer that synthesizes coarse reference action trajectories as explicit reasoning steps, using self-attention over noisy action sequences plus cross-attention with the VLM's key-value cache.
- Implicit Action Reasoner (IAR) โ extracts latent action priors via cross-attention with learnable query matrices over downsampled VLM representations from each VLM layer, aggregated by average pooling and an MLP projection.
Both reasoners' outputs undergo dual cross-attention with the noisy action query, are concatenated and passed through self-attention fusion, then condition the final action head, which emits a fine-grained action chunk. Reasoning never leaves the action manifold.
- LIBERO: 98.5% average success (Spatial 99.4 / Object 99.6 / Goal 98.8 / Long 96.0), first across all four suites; baseline 96.9%.
- LIBERO-Plus: 86.6% zero-shot transfer, 88.0% under SFT.
- VLABench: 63.5% Intention Score, 47.4% Progress Score across five tracks (+12.6% in unseen-texture conditions).
- Real-world (AgiBot G1): 66.7% average over three tasks (Wipe Stain, Pour Water, Open-set Pick), beating baselines by ~5โ30 pts.
Compared against 30+ methods including Diffusion Policy, Octo, CoT-VLA, WorldVLA, DreamVLA, OpenVLA, ฯ0, ฯ0.5, SpatialVLA, UniVLA, GR00T-N1. Ablation: EAR alone +1.4%, IAR alone +1.2%, combined +1.6% over baseline on LIBERO.
ACoT-VLA's bet: the action space is rich enough to encode reasoning if you give it a coarse-to-fine decomposition. Sister paper to Fast-ThinkAct (which compresses text CoT into a latent), and predecessor-style to it (which compresses the same reasoning into the same space). Likely to influence how production VLAs structure their action heads.
- arXiv: 2601.11404
- Code:
AgibotTech/ACoT-VLA
- VLA Architecture review ยงG (reasoning-augmented)
- Fast-ThinkAct ยท CoT-VLA
- CVPR 2026 survey
โ Back to CVPR-2026