CVPR 2026 Fast ThinkAct - Heungwoo/research GitHub Wiki

Fast-ThinkAct — Efficient VLA Reasoning via Verbalizable Latent Planning

Venue: CVPR 2026 Category: Reasoning / CoT VLA Trend tag: Trend 4 (action-CoT vs. sketch-CoT) Affiliations: NVIDIA · National Taiwan University · UIUC (NVIDIA Research Taiwan group; authors incl. Chi-Pin Huang, Fu-En Yang, Min-Hung Chen, Zhiding Yu, Jan Kautz, Yu-Chiang Frank Wang)

Approach diagram

flowchart LR
  OBS["obs + instruction"] --> TEACHER["teacher VLA<br/>explicit text CoT"]
  TEACHER --> SLOW["slow chain-of-thought<br/>many tokens"]
  SLOW --> DISTILL["distillation"]
  OBS --> STUDENT["student VLA<br/>latent-CoT tokens"]
  DISTILL --> STUDENT
  STUDENT --> LATENT["verbalizable latent CoT<br/>compressed, decodable"]
  LATENT --> ACT["action head"]
Loading

Problem

Reasoning VLAs (CoT-VLA, ThinkAct, ECoT-Lite) emit explicit text reasoning before producing actions. That text adds 10–100× tokens to the inference path, which kills real-time control. The trade-off looks forced: keep CoT (slow) or drop it (less robust).

Method

Distill the explicit text CoT into a verbalizable latent CoT — a small set of compact continuous latents that compress both the linguistic and visual planning trace while remaining decodable: a dedicated verbalizer module recovers the text reasoning post-hoc for interpretability. Training uses a preference-guided distillation objective that aligns the student's manipulation trajectories with the teacher's, so the student inherits both planning modalities rather than just imitating text tokens. The student VLA matches the teacher's action distribution while running with a fraction of the token budget.

Results

Up to 89.3 % inference-latency reduction (≈9.3× speedup) over state-of-the-art reasoning VLAs while maintaining performance. Evaluated across manipulation (SimplerEnv, LIBERO, RoboTwin2.0 bimanual), embodied reasoning (EgoPlan-Bench2, RoboVQA, OpenEQA), 10-shot few-shot adaptation on RoboTwin2.0, and failure recovery (RoboFAC-Sim/Real). The latent-CoT tokens remain decodable via the verbalizer into compact reasoning text.

Significance

The latency-vs-interpretability trade-off in reasoning VLAs has been the main blocker to deploying CoT VLAs in production. Fast-ThinkAct shows the trade-off is not Pareto-binding — most of the CoT benefit can be recovered with a small, decodable latent. Likely to be adopted into NVIDIA's own GR00T series and into production releases generally. Cousin recipe: π0.7's metadata + CFG (also compresses an interpretable signal into a few tokens).

Links

  • arXiv: 2601.09708
  • Project: NVIDIA Research Taiwan publications page

Related pages

← Back to CVPR-2026

⚠️ **GitHub.com Fallback** ⚠️