Review Fast in Slow - Heungwoo/research GitHub Wiki
Paper: Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning ยท Chen, Liu, Gu, Liu, Zhang, Li, He, Guo, Fu, Zhang, Heng ยท NeurIPS 2025 ยท arXiv 2506.01953 Project: https://fast-in-slow.github.io Related summary: Fast-in-Slow
This page is the long-form companion to the Fast-in-Slow summary. Everything is sourced from arXiv 2506.01953 and the project page.
Fast-in-Slow is one of three distinct architectural patterns for the dual-system design space published at NeurIPS 2025. Each solves the reasoning-vs-action conflict differently:
| Paper | Pattern | Mechanism |
|---|---|---|
| GR00T N1 / N1.5 (NVIDIA, Mar 2025) | Cascaded | Separate VLM (S2) โ diffusion transformer (S1) with a hand-designed latent interface |
| Fast-in-Slow (this review) | Embedded | Shared parameters โ the last 2 transformer blocks of the LLM are repurposed as S1 |
| ChatVLA-2 (NeurIPS 2025) | MoE-routed | Dynamic MoE that routes reasoning vs. action tokens to different experts |
| ThinkAct (NeurIPS 2025, NVIDIA) | RL-visual-plan | MLLM emits RL-rewarded plans compressed to a visual latent; separate action head |
FiS-VLA's claim is that embedding > cascading โ System 1 inherits VLM pretraining through shared parameters, not through a bottleneck interface.
Most dual-system VLAs (GR00T N1, RoboDual, Hi-Robot) cascade System 1 and System 2 as two separate networks joined by a hand-designed interface. This introduces latency, forces S1 to be trained from scratch, and prevents S1 from leveraging VLM pretraining. FiS-VLA inverts this: System 1 is not a separate model โ it's the final 2 transformer blocks of the VLM, repurposed as the fast action head. The same LLM plays both roles; S2 runs the full stack at low frequency, S1 runs only the last 2 blocks at high frequency. Result: 117.7 Hz control (NVIDIA 4090) with chunk=8; +8% simulation / +11% real-world average success rate over prior SOTA. Trained with a dual-aware co-training loss that keeps S2's autoregressive reasoning intact while S1 learns diffusion-based action generation.
Dual-system VLAs work, but the cascaded design (GR00T N1, RoboDual, Hi-Robot) has four structural problems:
- System 1 can't inherit VLM pretraining. S1 is typically a small diffusion or flow head trained from scratch on robot data โ it doesn't see the VLM's web-scale language/vision grounding.
- Interface is hand-designed and lossy. S2 must compress its thinking into a fixed-size latent/token that S1 consumes. Design of this interface is an open research question with no consensus.
- Parameter count balloons. S1 + S2 are two separate networks โ the total is larger than either alone.
- Synchronization is fragile. S1 runs at servo rate (50โ100 Hz), S2 at 5โ10 Hz; tracking when S1 should refresh from S2 requires extra engineering.
FiS-VLA's thesis: don't have two networks โ have one network with two activation patterns.

Figure 2 of the FiS-VLA paper (Chen et al., NeurIPS 2025). Left: at low frequency (1/n steps), the full LLM processes language + 2D images โ this is System 2's contextual reasoning pass. Middle: at high frequency (every step), only the final 2 transformer blocks (blocks 31โ32 of a 32-layer LLM) are re-run with multimodal high-frequency input โ 2D images, 3D point clouds (via a lightweight 3D tokenizer), robot state, and noised actions. Right: two heads โ diffusion (โ_fast, continuous actions) and autoregressive (โ_slow, discrete reasoning tokens) โ are co-trained. The shared encoder is used for both 2D images and point-cloud features. Included for scholarly review.
flowchart TB
subgraph S2[System 2 โ Slow path, 1/n steps]
L[Language instruction] --> FullLLM[Full LLM:<br/>Blocks 1 โ 32]
V2D[2D images] --> FullLLM
FullLLM --> LatentCond[Latent condition<br/>guides next nรH actions]
end
subgraph S1[System 1 โ Fast path, every step]
V2Dhf[2D images hi-freq] --> Last[LLM Blocks 31โ32<br/>ONLY<br/>shared with S2]
PC[3D point clouds<br/>via lightweight tokenizer] --> Last
RS[Robot state] --> Last
NA[Noised actions] --> Last
LatentCond --> Last
Last --> Diff[Diffusion head<br/>โ_fast = MSE on noise]
Last --> AR[Autoregressive head<br/>โ_slow = CE on tokens]
end
Diff --> Out[Continuous actions<br/>117.7 Hz @ chunk=8]
classDef shared fill:#e8f0ff,stroke:#4060a0,color:#000
class Last shared
classDef pretrained fill:#ffeadc,stroke:#c06030,color:#000
class FullLLM pretrained
Blue box = shared parameters. S2 and S1 are the same 2 transformer blocks at positions 31โ32; S2 additionally runs blocks 1โ30.
- Backbone: a pretrained Prismatic VLM (SigLIP + DINOv2 vision encoders, LLaMA2-7B LLM) with 32 transformer blocks in the LLM.
- System 2 = full 32-block forward pass on language + 2D images at low frequency.
- System 1 = last 2 blocks only (blocks 31โ32) re-run at high frequency with heterogeneous high-freq inputs.
- S1 and S2 share parameters in those 2 blocks โ no separate "action network."
- Ablation shows that 2 blocks is optimal โ performance saturates; fewer blocks hurt capacity, more blocks hurt speed.
- S2 (slow, every n steps): language instruction + 2D images only. Produces a latent condition consumed by S1 for the next nยทH time steps.
- S1 (fast, every step): 2D images (high-freq), 3D point cloud (tokenized via a lightweight 3D tokenizer, then embedded through the same shared encoder as the 2D images), robot state, and noised actions. Plus the latent condition from the most recent S2 pass.
This is the key design choice: S1 sees more modalities than S2, because S1 does the physics and S2 does the semantics.
S2 runs once for every 4 S1 steps. Ablation confirms 1:4 is optimal โ 1:1 is wasteful (S2 computed too often), 1:8 loses semantic tracking.
โ_fast โ diffusion over continuous actions:
- Noise the ground-truth action chunk at a random timestep ฯ.
- S1's diffusion head predicts the noise.
- MSE loss: standard DDPM-style.
โ_slow โ autoregressive reasoning preservation:
- S2's autoregressive output predicts discrete reasoning tokens (the VLM's original next-token-prediction on the language side).
- Standard cross-entropy over the token vocabulary.
- This loss is what prevents catastrophic forgetting of the VLM's reasoning capability while the fast head is trained for action generation.
Asynchronous sampling during training: at a 1:4 ratio so training matches inference.
- 117.7 Hz control frequency on NVIDIA 4090 with action chunk = 8.
- Simulation: 21.9 Hz with chunk=1 (compared to CogACT's 9.8 Hz).
- 7-DoF single-arm (Franka) and 14-DoF / 16-DoF dual-arm (AgileX, AlphaBot) setups tested.
- Pretraining corpus: ~860K+ trajectories across 37 open-source datasets (Open X-Embodiment, DROID, RoboMIND, โฆ), then fine-tuned with 100 demos/task per real-world platform.
| Method | Avg SR |
|---|---|
| ฯโ (flow-matching baseline) | 55% |
| CogACT | 61% |
| FiS-VLA | 69% |
(Approximate from paper; FiS-VLA beats CogACT by 8 points and ฯโ by 14 points on RLBench.)
| Platform | Tasks | ฯโ avg | FiS-VLA avg |
|---|---|---|---|
| AgileX (14-DoF) | Pick+place ยท Lift ball+place ยท Place bottles at rack ยท Wipe blackboard | 59% | 68% |
| AlphaBot (16-DoF) | Pick bowl+place ยท Handover+place ยท Pour water+move ยท Fold towel | 61% | 74% |
Largest single-task gain: Fold towel 40% โ 60% on AlphaBot.
Headline: +8% simulation / +11% real-world average success rate over SOTA.
117.7 Hz on 4090 with chunk=8 โ over 10ร faster than cascaded dual-system baselines (CogACT ~10 Hz).
| Shared blocks | Performance |
|---|---|
| 0 (no sharing โ cascaded baseline) | Worst |
| 1 | Good |
| 2 | Best โ performance saturates |
| 4 | Same as 2, slower |
Motivation for the 2-block choice: cheapest S1 forward pass that doesn't bottleneck capacity.
| S2 : S1 ratio | Effect |
|---|---|
| 1 : 1 | Wasteful โ S2 computed too often |
| 1 : 4 | Optimal |
| 1 : 8 | Loses semantic tracking |
- Without โ_slow: success rate drops from 69% โ 62% on RLBench. Demonstrates that co-training with autoregressive reasoning loss is not just a regularizer โ it's load-bearing for preserving VLM capabilities.
- Without heterogeneous inputs: S1 performance degrades significantly. Robot state and 3D point clouds each contribute substantially.
- Robot state provides internal status โ can't be reconstructed from vision alone.
- 3D point clouds enhance geometric understanding โ particularly on tasks with occluded or cluttered objects.
- 2D images alone are insufficient for the hardest contact-rich tasks.
Under object / background / lighting variations, FiS-VLA drops roughly 19โ31% โ indicating meaningful room for improvement on OOD robustness. Crucially, FiS-VLA degrades less than ฯโ on every axis (ฯโ drops ~27โ46%), so the robustness gap is the takeaway, not the absolute drop.
- Statically-configured sharing. The number of shared blocks (2) and the frequency ratio (1:4) are fixed hyperparameters, not adaptive to task difficulty. Authors identify dynamic adaptation as future work.
- OOD performance drops 19โ29% under visual / scene perturbations.
- Still a cascade at the inter-tier level โ S2's latent condition has to be passed forward to S1, which reintroduces a synchronization boundary (just inside one network instead of between two).
- No head-to-head with ฯ0.6 or ฯ0.7. FiS-VLA's real-world baselines are CogACT and ฯโ (not ฯ0.5 / ฯ0.6). The ฯ0.6 + Knowledge Insulation recipe (same NeurIPS 2025!) is arguably a stronger baseline for the "same VLM handles both reasoning and action" thesis.
- No LIBERO numbers. Standard benchmark is conspicuously absent โ only RLBench. Makes comparison with ICLR 2026 VLAs (which report LIBERO heavily) awkward.
- Fixed sharing depth raises transferability questions. "Last 2 blocks are S1" works for the Prismatic / LLaMA2-7B 32-layer LLM. What about Gemma3-4B (ฯ-series) or different architectures? Untested.
- Reasoning-preservation claim is tested only on action-task performance. The paper shows โ_slow helps action accuracy, but doesn't show S2's reasoning outputs remain as good as the un-finetuned VLM on language-only benchmarks (e.g., MMMU, MMBench). VLM4VLA-style analysis would strengthen the case.
-
Open weights โ released post-publication (correction to earlier "closed" claim). Training/inference code is public (github.com/CHEN-H01/Fast-in-Slow, MIT) and a pretrained large-scale checkpoint was released on HuggingFace (
haosad/fisvla) on 2025-07-08, so reproducibility is no longer the limitation it was at submission (project page: fast-in-slow.github.io).
FiS-VLA defines the "embedded dual-system" architectural pattern. Along with ChatVLA-2 (MoE-routed) and ThinkAct (RL-latent-plan), it establishes three concrete alternatives to the GR00T-style cascade that ICLR 2026 dual-system papers (WholeBodyVLA, HiMoE-VLA, AdaMoE) draw from.
Key structural contribution: parameter sharing between reasoning and action. This contrasts with two competing positions in the 2025 dual-system literature:
- Separate networks + hand-designed interface (GR00T N1, RoboDual, Hi-Robot)
- Sparse MoE within one network (ChatVLA-2, AdaMoE)
FiS-VLA's dense-sharing approach may generalize better to small-backbone regimes where MoE capacity is wasteful, while MoE may win at scale.
Efficiency contribution: 117.7 Hz control is production-relevant. Matches the ฯ0.6 latency target (63 ms/chunk โ 16 Hz chunk rate = comparable) while producing longer chunks per forward pass.
Placement (see Review-VLA-Architecture ยง5.F): FiS-VLA is Category F (Hierarchical / Dual-system / MoE), sub-pattern "embedded."
- Fast-in-Slow summary โ short version
- ChatVLA-2 โ MoE-routed sibling (NeurIPS 2025)
- ThinkAct โ RL-visual-plan sibling (NeurIPS 2025)
- Review: VLA Architectures ยง5.F โ dual-system / MoE category
- ฯ0.6 / Review-pi06 โ the flow-matching baseline arguably doing "same VLM, different role" via Knowledge Insulation
- Knowledge Insulation โ orthogonal mechanism for reasoning/action reconciliation (gradient-insulation rather than parameter-sharing)
- NeurIPS 2025 survey
- arXiv: https://arxiv.org/abs/2506.01953 ยท HTML: https://arxiv.org/html/2506.01953
- NeurIPS 2025 poster page: https://neurips.cc/virtual/2025/loc/san-diego/poster/119931
- Project: https://fast-in-slow.github.io
โ Back to NeurIPS-2025-Fast-in-Slow ยท NeurIPS-2025 ยท Home