CVPR 2026 Fast ThinkAct - Heungwoo/research GitHub Wiki
Venue: CVPR 2026 Category: Reasoning / CoT VLA Trend tag: Trend 4 (action-CoT vs. sketch-CoT) Affiliations: NVIDIA · National Taiwan University · UIUC (NVIDIA Research Taiwan group; authors incl. Chi-Pin Huang, Fu-En Yang, Min-Hung Chen, Zhiding Yu, Jan Kautz, Yu-Chiang Frank Wang)
flowchart LR
OBS["obs + instruction"] --> TEACHER["teacher VLA<br/>explicit text CoT"]
TEACHER --> SLOW["slow chain-of-thought<br/>many tokens"]
SLOW --> DISTILL["distillation"]
OBS --> STUDENT["student VLA<br/>latent-CoT tokens"]
DISTILL --> STUDENT
STUDENT --> LATENT["verbalizable latent CoT<br/>compressed, decodable"]
LATENT --> ACT["action head"]
Reasoning VLAs (CoT-VLA, ThinkAct, ECoT-Lite) emit explicit text reasoning before producing actions. That text adds 10–100× tokens to the inference path, which kills real-time control. The trade-off looks forced: keep CoT (slow) or drop it (less robust).
Distill the explicit text CoT into a verbalizable latent CoT — a small set of compact continuous latents that compress both the linguistic and visual planning trace while remaining decodable: a dedicated verbalizer module recovers the text reasoning post-hoc for interpretability. Training uses a preference-guided distillation objective that aligns the student's manipulation trajectories with the teacher's, so the student inherits both planning modalities rather than just imitating text tokens. The student VLA matches the teacher's action distribution while running with a fraction of the token budget.
Up to 89.3 % inference-latency reduction (≈9.3× speedup) over state-of-the-art reasoning VLAs while maintaining performance. Evaluated across manipulation (SimplerEnv, LIBERO, RoboTwin2.0 bimanual), embodied reasoning (EgoPlan-Bench2, RoboVQA, OpenEQA), 10-shot few-shot adaptation on RoboTwin2.0, and failure recovery (RoboFAC-Sim/Real). The latent-CoT tokens remain decodable via the verbalizer into compact reasoning text.
The latency-vs-interpretability trade-off in reasoning VLAs has been the main blocker to deploying CoT VLAs in production. Fast-ThinkAct shows the trade-off is not Pareto-binding — most of the CoT benefit can be recovered with a small, decodable latent. Likely to be adopted into NVIDIA's own GR00T series and into production releases generally. Cousin recipe: π0.7's metadata + CFG (also compresses an interpretable signal into a few tokens).
- arXiv: 2601.09708
- Project: NVIDIA Research Taiwan publications page
- ThinkAct · ECoT-Lite · ACoT-VLA (sister action-CoT paper at CVPR 2026)
- VLA Architecture review §G (reasoning-augmented)
- CVPR 2026 survey
← Back to CVPR-2026