NeurIPS 2025 Chain of Action - Heungwoo/research GitHub Wiki
Venue: NeurIPS 2025 ยท Author: ByteDance Seed ยท arXiv: 2506.09990 Category: VLA Architecture / CoT
flowchart RL
Goal[Goal keyframe] --> KT[Token at t+H]
KT --> KTm1[Token at t+H-1]
KTm1 --> Dots[...]
Dots --> Kt1[Token at t+1]
Kt1 --> Kt[Token at t]
Kt --> Exec[Execute]
note[Generation is BACKWARD:<br/>goal โ current, not current โ goal]
Standard AR VLAs generate actions forward in time (t โ t+H). But manipulation tasks are often defined by their end state (goal keyframe). Forward AR has the model reason "what do I do first?" โ a harder question than "what's the last step before the goal?"
Generate the trajectory backward โ from a predicted goal keyframe to the current state โ in a single autoregressive model. This is action-level CoT: each step is "what action would produce the next keyframe on the path from goal to current?"
At inference, the model:
- Predicts the goal keyframe โ the first token is a stable keyframe action encoding the task goal (predicted by the model, not externally provided).
- Autoregressively generates trajectory tokens backward from the goal, enforcing a global-to-local structure where each local action is constrained by the final goal.
- Executes the earliest (current-time) action first.
Actions are emitted as continuous action tokens (regression via linear projection, not discrete tokenization or a separate diffusion/flow head) within a single unified autoregressive transformer. The backbone is lightweight: a ResNet18 (ImageNet-pretrained) vision encoder over multi-view 128ร128 RGB, feeding a ~512-d, 8-head encoder-decoder transformer (4 encoder / 7 decoder layers, including a multi-token-prediction layer). The keyframe is defined as a step where the gripper state changes or joint velocities approach zero.
- SOTA on 60 RLBench tasks + 8 real tasks (described as the most comprehensive RLBench evaluation to date).
- Outperforms ACT by 16% and Diffusion Policy by 23% averaged across the 60 RLBench tasks; surpasses ACT by 15% in real-world manipulation.
- On the 60-task suite the absolute averages are roughly CoA 55% vs ACT 39% vs Diffusion Policy 33%; real-world 8-task kitchen suite ~61% (CoA) vs ~46% (ACT).
- Ablations show backward (reverse) AR is the core driver (โ0.76 vs 0.67 forward), and each of the four supporting designs โ continuous action tokens, dynamic stopping (variable-length trajectories), reverse temporal ensemble, and multi-token prediction โ contributes positively.
- Decisive ablation: regularizing the latent action space via a latent consistency loss is critical (โ0.76 vs 0.21 without it) โ far more important than direct action-space supervision.
A conceptually simple change (flip the AR direction) with large empirical impact. Feeds into the ICLR 2026 action-CoT cluster alongside ThinkAct and Robot-R1, informing Embodied-R1's pointing primitives and Actions as Language's framing.
- arXiv: https://arxiv.org/abs/2506.09990
- NeurIPS virtual page: https://neurips.cc/virtual/2025/loc/san-diego/poster/116620
- GitHub: https://github.com/ByteDance-Seed/Chain-of-Action
- ECoT-Lite (CoRL 2025 ancestor)
- ThinkAct (RL-for-reasoning sibling)
- Embodied-R1 (ICLR 2026 descendant)
- Review: VLA Architectures โ ยง5.G
โ Back to NeurIPS-2025