NeurIPS 2025 Chain of Action - Heungwoo/research GitHub Wiki

Chain-of-Action (CoA) โ€” Trajectory Autoregressive Modeling

Venue: NeurIPS 2025 ยท Author: ByteDance Seed ยท arXiv: 2506.09990 Category: VLA Architecture / CoT

Approach diagram

flowchart RL
  Goal[Goal keyframe] --> KT[Token at t+H]
  KT --> KTm1[Token at t+H-1]
  KTm1 --> Dots[...]
  Dots --> Kt1[Token at t+1]
  Kt1 --> Kt[Token at t]
  Kt --> Exec[Execute]
  note[Generation is BACKWARD:<br/>goal โ†’ current, not current โ†’ goal]
Loading

Problem

Standard AR VLAs generate actions forward in time (t โ†’ t+H). But manipulation tasks are often defined by their end state (goal keyframe). Forward AR has the model reason "what do I do first?" โ€” a harder question than "what's the last step before the goal?"

Method

Generate the trajectory backward โ€” from a predicted goal keyframe to the current state โ€” in a single autoregressive model. This is action-level CoT: each step is "what action would produce the next keyframe on the path from goal to current?"

At inference, the model:

  1. Predicts the goal keyframe โ€” the first token is a stable keyframe action encoding the task goal (predicted by the model, not externally provided).
  2. Autoregressively generates trajectory tokens backward from the goal, enforcing a global-to-local structure where each local action is constrained by the final goal.
  3. Executes the earliest (current-time) action first.

Actions are emitted as continuous action tokens (regression via linear projection, not discrete tokenization or a separate diffusion/flow head) within a single unified autoregressive transformer. The backbone is lightweight: a ResNet18 (ImageNet-pretrained) vision encoder over multi-view 128ร—128 RGB, feeding a ~512-d, 8-head encoder-decoder transformer (4 encoder / 7 decoder layers, including a multi-token-prediction layer). The keyframe is defined as a step where the gripper state changes or joint velocities approach zero.

Results

  • SOTA on 60 RLBench tasks + 8 real tasks (described as the most comprehensive RLBench evaluation to date).
  • Outperforms ACT by 16% and Diffusion Policy by 23% averaged across the 60 RLBench tasks; surpasses ACT by 15% in real-world manipulation.
  • On the 60-task suite the absolute averages are roughly CoA 55% vs ACT 39% vs Diffusion Policy 33%; real-world 8-task kitchen suite ~61% (CoA) vs ~46% (ACT).
  • Ablations show backward (reverse) AR is the core driver (โ‰ˆ0.76 vs 0.67 forward), and each of the four supporting designs โ€” continuous action tokens, dynamic stopping (variable-length trajectories), reverse temporal ensemble, and multi-token prediction โ€” contributes positively.
  • Decisive ablation: regularizing the latent action space via a latent consistency loss is critical (โ‰ˆ0.76 vs 0.21 without it) โ€” far more important than direct action-space supervision.

Significance

A conceptually simple change (flip the AR direction) with large empirical impact. Feeds into the ICLR 2026 action-CoT cluster alongside ThinkAct and Robot-R1, informing Embodied-R1's pointing primitives and Actions as Language's framing.

Links

Related pages

โ† Back to NeurIPS-2025

โš ๏ธ **GitHub.com Fallback** โš ๏ธ