NeurIPS 2025 Fast in Slow - Heungwoo/research GitHub Wiki
Fast-in-Slow (FiS-VLA) โ Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning
Venue: NeurIPS 2025 ยท arXiv: 2506.01953 Authors / affiliations: Hao Chen, Jiaming Liu, Chenyang Gu, Zhuoyang Liu, Renrui Zhang, Xiaoqi Li, Xiao He, Yandong Guo, Chi-Wing Fu, Shanghang Zhang, Pheng-Ann Heng โ Peking University, CUHK, BAAI, AI2Robotics Category: Dual-System VLA Backbone: Prismatic-initialized 7B VLM (LLaMA2-7B LLM + SigLIP & DINOv2 vision encoders), plus a lightweight 3D point-cloud tokenizer (FPS + kNN)
flowchart LR
subgraph S2[System 2: slow VLM, full LLaMA2-7B]
VLM[VLM backbone<br/>reasoning, high capacity]
S1param[โก System 1<br/>final transformer blocks reused]
end
Obs1[High-freq obs:<br/>2D images + state + 3D point cloud] --> S1param
Obs2[Low-freq obs:<br/>2D images + language] --> VLM
S1param --> A[Actions @ 117.7 Hz<br/>action chunk = 8]
VLM -. latent conditioning, 1:4 ratio .-> S1param
Dual-system VLAs (GR00T N1, RoboDual, Hi-Robot) cascade System 1 and System 2 โ S2 produces a plan/latent, S1 produces actions. But cascading introduces interface latency and requires a hand-designed interface. Can we get the benefits of dual-system (fast control + slow reasoning) in one unified model?
Embed System-1 inside System-2 via partial parameter sharing. Rather than attaching a separate policy network as System 1 (as in prior dual-system VLAs), FiS repurposes the final transformer blocks of the LLM as the fast execution module (ablations find ~2 shared blocks optimal) while the full network remains System 2 for reasoning. The two systems take heterogeneous inputs at asynchronous frequencies (a 1:4 System-2 : System-1 ratio):
- System 1 (fast path): high-frequency robot state, 2D images, and 3D point clouds โ produces action chunks via a flow/diffusion denoising objective, conditioned on System 2's latent features.
- System 2 (slow path): low-frequency 2D images + language โ multimodal latent representations that guide System 1.
Training uses a dual-aware co-training strategy: a diffusion denoising objective equips System 1 with action generation, while an autoregressive objective preserves System 2's contextual reasoning.
This is an embedded dual-system โ System 1 and System 2 are the same model with different activation patterns, not two separate models.
- Sim (RLBench, 10 tasks): ~69% mean success vs. CogACT 61% / ฯ0 55% โ +8% over prior SOTA. Baselines include ManipLLM, OpenVLA, CogACT, ฯ0.
- Real (dual-arm Agilex + AlphaBot, 8 tasks): 68% (Agilex) and 74% (AlphaBot) vs. ฯ0's ~59%/61% โ +11% avg.
- 117.7 Hz control rate with action chunk = 8 (21.9 Hz at chunk = 1, vs. CogACT 9.8 Hz / ฯ0 13.8 Hz) on a single NVIDIA RTX 4090 โ real-time servo loop despite a large VLM in the pipeline.
- Key ablations: shared-block count (~2 optimal), System-2:System-1 frequency ratio (1:4 best), heterogeneous System-1 modalities (state + 2D + 3D each help), and the dual-aware co-training loss (the System-2 reasoning loss is necessary to preserve VLM knowledge).
Resolves the "dual-system overhead" problem without giving up dual-system benefits. Together with ChatVLA-2 (MoE separation) and ThinkAct (RL visual plans), it establishes three distinct architectural patterns for the dual-system design space that ICLR 2026's WholeBodyVLA / HiMoE-VLA / AdaMoE then pick from.
- arXiv: https://arxiv.org/abs/2506.01953
- NeurIPS virtual page: https://neurips.cc/virtual/2025/loc/san-diego/poster/119931
For a long-form review with the paper's architecture figure, RLBench + real-robot accuracy numbers, ablations (shared-block count, 1:4 frequency ratio, loss components), and limitations: In-Depth Review of Fast-in-Slow.
- ChatVLA-2 ยท ThinkAct (dual-system siblings)
- WholeBodyVLA ยท HiMoE-VLA (ICLR 2026 descendants)
- Review: VLA Architectures โ ยง5.F
โ Back to NeurIPS-2025