NeurIPS 2025 ThinkAct - Heungwoo/research GitHub Wiki

ThinkAct — VLA Reasoning via Reinforced Visual Latent Planning

Venue: NeurIPS 2025 · Authors: Chi-Pin Huang, Yueh-Hua Wu, Min-Hung Chen, Yu-Chiang Frank Wang, Fu-En Yang (NVIDIA + National Taiwan University) · arXiv: 2507.16815 Category: RL / CoT Reasoning / Dual-System

Approach diagram

flowchart LR
  Obs[Observation + instruction] --> MLLM[MLLM reasoner]
  MLLM -- generates plan --> P[Reasoning plan]
  P -- RL reward --> R{Reward:<br/>goal completion +<br/>trajectory consistency}
  MLLM -- plan --> L[Visual latent plan]
  L --> AE[Action model]
  AE --> A[Action]
  R -- action-aligned RL --> MLLM
Loading

Problem

Chain-of-thought VLAs emit text reasoning but:

  1. Text tokens are slow to decode (autoregressive serial);
  2. Reasoning quality is weakly supervised — there's no guarantee it helps actions. A better signal would be reinforcement that shapes reasoning toward plans that actually work in the environment.

Method

Two-stage pipeline with RL-shaped visual plans:

  1. MLLM plans in language as embodied chain-of-thought. Backbone = Qwen2.5-VL 7B (also tested at 3B).
  2. Plan is rewarded by GRPO RL (6K iterations, batch 64, lr 1e-6, rollout 5) using action-aligned visual rewards: a goal reward (predicted vs. detected start/end positions) + a trajectory reward (dynamic-time-warping distance to the visual trajectory distribution), plus a format-correctness term.
  3. Reinforced plan is compressed to a visual plan latent that conditions a downstream DiT-based action model (~432M params; DINOv2 image encoder + CLIP text encoder, 1024-d embeddings). During action adaptation the reasoning MLLM is frozen — only the action model's state encoder, latent projector, and policy head are trained (dual-system: System-2 plan, System-1 act).

At inference the reasoner and action model run asynchronously ("slow thinking, fast control"): each visual plan latent conditions N action steps (N=75 on LIBERO, N=15 on SimplerEnv).

Trains the MLLM to reason in ways that lead to successful actions, not just fluent text.

Results

  • LIBERO overall success 84.4% (Spatial 88.3 / Object 91.4 / Goal 87.1 / Long 70.9), beating DiT-Policy and CoT-VLA.
  • SimplerEnv highest overall scores: Google-VM 71.5%, Google-VA 65.1%, Bridge-VM 43.8% (+15.5 / +16.9 / +11.4 over the baseline action model).
  • Embodied reasoning: EgoPlan-Bench2 48.2%, RoboVQA 59.8 BLEU, OpenEQA 56.2%.
  • Few-shot adaptation on novel tasks; strong long-horizon planning; self-correction when initial plans fail (the RL reward gradient teaches plan revision).

Significance

Published NeurIPS version of the RL-for-reasoning-in-VLA pattern. Along with Chain-of-Action and Robot-R1, ThinkAct seeds the ICLR 2026 R1-style VLA cluster:

Links

Related pages

← Back to NeurIPS-2025

⚠️ **GitHub.com Fallback** ⚠️