NeurIPS 2025 ThinkAct - Heungwoo/research GitHub Wiki
Venue: NeurIPS 2025 · Authors: Chi-Pin Huang, Yueh-Hua Wu, Min-Hung Chen, Yu-Chiang Frank Wang, Fu-En Yang (NVIDIA + National Taiwan University) · arXiv: 2507.16815 Category: RL / CoT Reasoning / Dual-System
flowchart LR
Obs[Observation + instruction] --> MLLM[MLLM reasoner]
MLLM -- generates plan --> P[Reasoning plan]
P -- RL reward --> R{Reward:<br/>goal completion +<br/>trajectory consistency}
MLLM -- plan --> L[Visual latent plan]
L --> AE[Action model]
AE --> A[Action]
R -- action-aligned RL --> MLLM
Chain-of-thought VLAs emit text reasoning but:
- Text tokens are slow to decode (autoregressive serial);
- Reasoning quality is weakly supervised — there's no guarantee it helps actions. A better signal would be reinforcement that shapes reasoning toward plans that actually work in the environment.
Two-stage pipeline with RL-shaped visual plans:
- MLLM plans in language as embodied chain-of-thought. Backbone = Qwen2.5-VL 7B (also tested at 3B).
- Plan is rewarded by GRPO RL (6K iterations, batch 64, lr 1e-6, rollout 5) using action-aligned visual rewards: a goal reward (predicted vs. detected start/end positions) + a trajectory reward (dynamic-time-warping distance to the visual trajectory distribution), plus a format-correctness term.
- Reinforced plan is compressed to a visual plan latent that conditions a downstream DiT-based action model (~432M params; DINOv2 image encoder + CLIP text encoder, 1024-d embeddings). During action adaptation the reasoning MLLM is frozen — only the action model's state encoder, latent projector, and policy head are trained (dual-system: System-2 plan, System-1 act).
At inference the reasoner and action model run asynchronously ("slow thinking, fast control"): each visual plan latent conditions N action steps (N=75 on LIBERO, N=15 on SimplerEnv).
Trains the MLLM to reason in ways that lead to successful actions, not just fluent text.
- LIBERO overall success 84.4% (Spatial 88.3 / Object 91.4 / Goal 87.1 / Long 70.9), beating DiT-Policy and CoT-VLA.
- SimplerEnv highest overall scores: Google-VM 71.5%, Google-VA 65.1%, Bridge-VM 43.8% (+15.5 / +16.9 / +11.4 over the baseline action model).
- Embodied reasoning: EgoPlan-Bench2 48.2%, RoboVQA 59.8 BLEU, OpenEQA 56.2%.
- Few-shot adaptation on novel tasks; strong long-horizon planning; self-correction when initial plans fail (the RL reward gradient teaches plan revision).
Published NeurIPS version of the RL-for-reasoning-in-VLA pattern. Along with Chain-of-Action and Robot-R1, ThinkAct seeds the ICLR 2026 R1-style VLA cluster:
- Embodied-R1 — RL on pointing primitives
- SimpleVLA-RL — scaling infrastructure
- VLA-RFT — RL in world-model
- arXiv: https://arxiv.org/abs/2507.16815
- NeurIPS virtual page: https://neurips.cc/virtual/2025/loc/san-diego/poster/119747
- Project: https://jasper0314-huang.github.io/thinkact-vla/
- Fast-in-Slow · ChatVLA-2 (dual-system siblings)
- Embodied-R1 · SimpleVLA-RL (ICLR 2026 descendants)
- RL for VLA (cross-venue RL topic landing)
- Review: VLA Architectures — §5.F + §5.G
← Back to NeurIPS-2025