ICML 2026 VLA ATTC - Heungwoo/research GitHub Wiki
VLA-ATTC — Adaptive test-time compute that deliberates only when a VLA is uncertain
Venue: ICML 2026 (Poster) Category: RL for VLA / Efficiency Traction (2026-06): 5 citations (arXiv)

Problem
Vision-Language-Action (VLA) models make decisions through a fast, instinctive ("System 1") inference pass that lacks deliberation. This works for easy scenarios but produces suboptimal or even catastrophic actions in complex or ambiguous situations that demand deeper thought. Prior attempts to add deliberation fall into two camps, both flawed: sequential approaches (e.g. Chain-of-Thought) require costly fine-tuning, laborious CoT annotation, and can degrade action quality by forcing action-centric models to emit text; parallel approaches generate many candidates and score them, but apply deliberation indiscriminately to every state and rely on unstable absolute value estimation, making them prohibitively expensive for real robots. The authors frame three challenges: (i) quantifying situational difficulty for adaptive deliberation, (ii) minimizing multi-action sampling overhead, and (iii) designing a lightweight yet high-fidelity action critic.
Method
VLA-ATTC equips a frozen base VLA with adaptive test-time compute (TTC) without modifying the base model. A Cognitive Clutch monitors the base policy's uncertainty: at each step it samples two action chunks from the shared VLM context using different noise seeds and measures their Dynamic Time Warping (DTW) distance as an uncertainty score. If the score is below a threshold τ (set as the K-th percentile over an offline dataset), the model executes reflexively; if above, it enters the TTC Deliberation Phase, which generates N candidate actions in parallel by amortizing the expensive VLM pre-fill, then runs a tournament. A novel Relative Action Critic (RAC) — a lightweight Transformer matching the VLM's depth — performs iterative pairwise comparisons rather than unstable absolute scoring. The RAC ingests four inputs through dedicated MLPs (Action i, Action j, their difference, and proprioceptive state) and uses a multi-branch attention block fusing self-attention, raw cross-attention to VLM features, and query cross-attention to N_q learnable distilled-context queries. Training preference pairs are curated automatically by reducing ODE integration steps in the flow-matching action head to create graded "expert vs. sub-optimal" pairs — no manual annotation.

Results
On LIBERO-LONG, VLA-ATTC raises PI0.5 average success from 90.6% to 95.4% (Full) / 94% and PI0 from 82.8% to 92.2% / 90.6% — reducing the SOTA model's failure rate by over 50%. On the hardest "Both pots on stove" task PI0 jumps 40%→58%, and "Mug in microwave and close" jumps 62%→88%. It consistently beats the prior parallel-deliberation method Robomonkey (56.5% avg). On a real Agilex Piper arm, PI0+ATTC improves average success by +17.3% (46%→63.3%). Ablations show every RAC component matters (removing learnable queries is most damaging), N=16 candidates is the efficiency sweet spot, and minimal N=2 DTW sampling agrees with human difficulty rankings 89.2% of the time at a fraction of the cost. Efficiency: VLA-ATTC sustains 20.8 Hz control vs. 23.3 Hz baseline and just 1.5 Hz for Robomonkey.
Significance
VLA-ATTC shows that "thinking slow" can be applied surgically — only on the sparse hard states a cognitive clutch flags — turning System-2 deliberation from an always-on luxury into a near-real-time, plug-in capability for frozen VLAs. The relative pairwise critic sidesteps the instability of absolute value learning, and the automated ODE-step preference pipeline removes the annotation bottleneck. Code and weights are to be open-sourced.
Links
- arXiv: 2605.01194
- ICML 2026: https://icml.cc/virtual/2026/poster/61157
← Back to ICML-2026