ICML 2026 Think Less Act Early - Heungwoo/research GitHub Wiki
Venue: ICML 2026 (Poster) Category: Reasoning Affiliations: Dianqiao Lei, Lianlei Shan
Vision-Language-Action (VLA) models increasingly rely on explicit chain-of-thought-style reasoning before acting, which improves decision quality but adds substantial inference latency — a poor fit for embodied control where the robot must act in real time. Generating long explicit reasoning traces for every step wastes computation, and full-reasoning baselines can also be unstable. The paper asks how a VLA can reason effectively yet cheaply, terminating its deliberation as soon as it is confident enough to act.
The authors propose AVA-VLA, a framework that replaces explicit reasoning with latent variable sequences rather than verbalized chains of thought. Two ideas drive the approach:
- Reinforcement-learned latent reasoning. Reasoning is treated as a sequence of latent variables optimized via reinforcement learning using task-level rewards, so the latent reasoning is shaped directly by downstream task success rather than by imitating textual rationales.
- Confidence-based early exit. An early-exit mechanism adaptively terminates reasoning based on state confidence, letting the model "think less" and "act early" when it is already certain — cutting computational latency on embodied decision-making tasks.
flowchart LR
O[Vision + Language Observation] --> L[Latent reasoning step]
L --> C{State confidence<br/>high enough?}
C -- no --> L
C -- yes --> A[Emit action / act early]
R[Task-level reward] -. RL optimization .-> L
Schematic derived from the paper's abstract; AVA-VLA refines latent reasoning variables via RL and exits early once state confidence is sufficient.
The paper reports that AVA-VLA achieves superior stability and success rates compared to full-reasoning baselines while reducing computational latency on embodied decision-making tasks. (Quantitative tables are not available in the public ICML abstract; specific numbers will be added once the full paper is released.)
AVA-VLA targets a central tension in reasoning-augmented VLAs: deliberation improves decisions but is too slow for real-time control. By moving reasoning into a learned latent space optimized by RL — rather than expensive explicit text generation — and adding a confidence-gated early-exit, the work aims to retain the accuracy benefits of reasoning while restoring the low latency needed for embodied deployment. This places it alongside other 2026 efforts to make VLA reasoning adaptive and compute-aware.
- ICML 2026: https://icml.cc/virtual/2026/poster/61571
← Back to ICML-2026