Review Fast in Slow - Heungwoo/research GitHub Wiki

In-Depth Review โ€” Fast-in-Slow: A Dual-System Foundation Model

Paper: Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning ยท Chen, Liu, Gu, Liu, Zhang, Li, He, Guo, Fu, Zhang, Heng ยท NeurIPS 2025 ยท arXiv 2506.01953 Project: https://fast-in-slow.github.io Related summary: Fast-in-Slow

This page is the long-form companion to the Fast-in-Slow summary. Everything is sourced from arXiv 2506.01953 and the project page.

๐Ÿ“Ž Dual-system VLA context

Fast-in-Slow is one of three distinct architectural patterns for the dual-system design space published at NeurIPS 2025. Each solves the reasoning-vs-action conflict differently:

Paper Pattern Mechanism
GR00T N1 / N1.5 (NVIDIA, Mar 2025) Cascaded Separate VLM (S2) โ†’ diffusion transformer (S1) with a hand-designed latent interface
Fast-in-Slow (this review) Embedded Shared parameters โ€” the last 2 transformer blocks of the LLM are repurposed as S1
ChatVLA-2 (NeurIPS 2025) MoE-routed Dynamic MoE that routes reasoning vs. action tokens to different experts
ThinkAct (NeurIPS 2025, NVIDIA) RL-visual-plan MLLM emits RL-rewarded plans compressed to a visual latent; separate action head

FiS-VLA's claim is that embedding > cascading โ€” System 1 inherits VLM pretraining through shared parameters, not through a bottleneck interface.


1. TL;DR

Most dual-system VLAs (GR00T N1, RoboDual, Hi-Robot) cascade System 1 and System 2 as two separate networks joined by a hand-designed interface. This introduces latency, forces S1 to be trained from scratch, and prevents S1 from leveraging VLM pretraining. FiS-VLA inverts this: System 1 is not a separate model โ€” it's the final 2 transformer blocks of the VLM, repurposed as the fast action head. The same LLM plays both roles; S2 runs the full stack at low frequency, S1 runs only the last 2 blocks at high frequency. Result: 117.7 Hz control (NVIDIA 4090) with chunk=8; +8% simulation / +11% real-world average success rate over prior SOTA. Trained with a dual-aware co-training loss that keeps S2's autoregressive reasoning intact while S1 learns diffusion-based action generation.

2. Motivation โ€” the cost of two separate networks

Dual-system VLAs work, but the cascaded design (GR00T N1, RoboDual, Hi-Robot) has four structural problems:

  1. System 1 can't inherit VLM pretraining. S1 is typically a small diffusion or flow head trained from scratch on robot data โ€” it doesn't see the VLM's web-scale language/vision grounding.
  2. Interface is hand-designed and lossy. S2 must compress its thinking into a fixed-size latent/token that S1 consumes. Design of this interface is an open research question with no consensus.
  3. Parameter count balloons. S1 + S2 are two separate networks โ€” the total is larger than either alone.
  4. Synchronization is fragile. S1 runs at servo rate (50โ€“100 Hz), S2 at 5โ€“10 Hz; tracking when S1 should refresh from S2 requires extra engineering.

FiS-VLA's thesis: don't have two networks โ€” have one network with two activation patterns.

3. Representative diagram

Figure 2 from the paper โ€” FiS-VLA framework

FiS-VLA framework (Figure 2 from Chen et al., 2025)

Figure 2 of the FiS-VLA paper (Chen et al., NeurIPS 2025). Left: at low frequency (1/n steps), the full LLM processes language + 2D images โ€” this is System 2's contextual reasoning pass. Middle: at high frequency (every step), only the final 2 transformer blocks (blocks 31โ€“32 of a 32-layer LLM) are re-run with multimodal high-frequency input โ€” 2D images, 3D point clouds (via a lightweight 3D tokenizer), robot state, and noised actions. Right: two heads โ€” diffusion (โ„’_fast, continuous actions) and autoregressive (โ„’_slow, discrete reasoning tokens) โ€” are co-trained. The shared encoder is used for both 2D images and point-cloud features. Included for scholarly review.

Our reconstruction as mermaid

flowchart TB
  subgraph S2[System 2 โ€” Slow path, 1/n steps]
    L[Language instruction] --> FullLLM[Full LLM:<br/>Blocks 1 โ†’ 32]
    V2D[2D images] --> FullLLM
    FullLLM --> LatentCond[Latent condition<br/>guides next nร—H actions]
  end

  subgraph S1[System 1 โ€” Fast path, every step]
    V2Dhf[2D images hi-freq] --> Last[LLM Blocks 31โ€“32<br/>ONLY<br/>shared with S2]
    PC[3D point clouds<br/>via lightweight tokenizer] --> Last
    RS[Robot state] --> Last
    NA[Noised actions] --> Last
    LatentCond --> Last
    Last --> Diff[Diffusion head<br/>โ„’_fast = MSE on noise]
    Last --> AR[Autoregressive head<br/>โ„’_slow = CE on tokens]
  end

  Diff --> Out[Continuous actions<br/>117.7 Hz @ chunk=8]
  classDef shared fill:#e8f0ff,stroke:#4060a0,color:#000
  class Last shared
  classDef pretrained fill:#ffeadc,stroke:#c06030,color:#000
  class FullLLM pretrained
Loading

Blue box = shared parameters. S2 and S1 are the same 2 transformer blocks at positions 31โ€“32; S2 additionally runs blocks 1โ€“30.


4. Method โ€” full recipe

4.1 "Embedded S1-in-S2" via partial parameter sharing

  • Backbone: a pretrained Prismatic VLM (SigLIP + DINOv2 vision encoders, LLaMA2-7B LLM) with 32 transformer blocks in the LLM.
  • System 2 = full 32-block forward pass on language + 2D images at low frequency.
  • System 1 = last 2 blocks only (blocks 31โ€“32) re-run at high frequency with heterogeneous high-freq inputs.
  • S1 and S2 share parameters in those 2 blocks โ€” no separate "action network."
  • Ablation shows that 2 blocks is optimal โ€” performance saturates; fewer blocks hurt capacity, more blocks hurt speed.

4.2 Heterogeneous input at different frequencies

  • S2 (slow, every n steps): language instruction + 2D images only. Produces a latent condition consumed by S1 for the next nยทH time steps.
  • S1 (fast, every step): 2D images (high-freq), 3D point cloud (tokenized via a lightweight 3D tokenizer, then embedded through the same shared encoder as the 2D images), robot state, and noised actions. Plus the latent condition from the most recent S2 pass.

This is the key design choice: S1 sees more modalities than S2, because S1 does the physics and S2 does the semantics.

4.3 1:4 operating frequency ratio

S2 runs once for every 4 S1 steps. Ablation confirms 1:4 is optimal โ€” 1:1 is wasteful (S2 computed too often), 1:8 loses semantic tracking.

4.4 Dual-aware co-training โ€” two losses in one model

$$\mathcal{L}_{\text{FiS-VLA}} = \mathcal{L}_{\text{fast}} + \mathcal{L}_{\text{slow}}$$

โ„’_fast โ€” diffusion over continuous actions:

  • Noise the ground-truth action chunk at a random timestep ฯ„.
  • S1's diffusion head predicts the noise.
  • MSE loss: standard DDPM-style.

โ„’_slow โ€” autoregressive reasoning preservation:

  • S2's autoregressive output predicts discrete reasoning tokens (the VLM's original next-token-prediction on the language side).
  • Standard cross-entropy over the token vocabulary.
  • This loss is what prevents catastrophic forgetting of the VLM's reasoning capability while the fast head is trained for action generation.

Asynchronous sampling during training: at a 1:4 ratio so training matches inference.

4.5 Hardware & inference

  • 117.7 Hz control frequency on NVIDIA 4090 with action chunk = 8.
  • Simulation: 21.9 Hz with chunk=1 (compared to CogACT's 9.8 Hz).
  • 7-DoF single-arm (Franka) and 14-DoF / 16-DoF dual-arm (AgileX, AlphaBot) setups tested.
  • Pretraining corpus: ~860K+ trajectories across 37 open-source datasets (Open X-Embodiment, DROID, RoboMIND, โ€ฆ), then fine-tuned with 100 demos/task per real-world platform.

5. Results

5.1 RLBench simulation โ€” 10 tasks, average success rate

Method Avg SR
ฯ€โ‚€ (flow-matching baseline) 55%
CogACT 61%
FiS-VLA 69%

(Approximate from paper; FiS-VLA beats CogACT by 8 points and ฯ€โ‚€ by 14 points on RLBench.)

5.2 Real-world dual-arm โ€” AgileX and AlphaBot (100 demos/task, 20 rollouts eval)

Platform Tasks ฯ€โ‚€ avg FiS-VLA avg
AgileX (14-DoF) Pick+place ยท Lift ball+place ยท Place bottles at rack ยท Wipe blackboard 59% 68%
AlphaBot (16-DoF) Pick bowl+place ยท Handover+place ยท Pour water+move ยท Fold towel 61% 74%

Largest single-task gain: Fold towel 40% โ†’ 60% on AlphaBot.

Headline: +8% simulation / +11% real-world average success rate over SOTA.

5.3 Control frequency

117.7 Hz on 4090 with chunk=8 โ€” over 10ร— faster than cascaded dual-system baselines (CogACT ~10 Hz).


6. Ablations

6.1 Number of shared blocks

Shared blocks Performance
0 (no sharing โ€” cascaded baseline) Worst
1 Good
2 Best โ€” performance saturates
4 Same as 2, slower

Motivation for the 2-block choice: cheapest S1 forward pass that doesn't bottleneck capacity.

6.2 Asynchronous frequency ratio

S2 : S1 ratio Effect
1 : 1 Wasteful โ€” S2 computed too often
1 : 4 Optimal
1 : 8 Loses semantic tracking

6.3 Loss components

  • Without โ„’_slow: success rate drops from 69% โ†’ 62% on RLBench. Demonstrates that co-training with autoregressive reasoning loss is not just a regularizer โ€” it's load-bearing for preserving VLM capabilities.
  • Without heterogeneous inputs: S1 performance degrades significantly. Robot state and 3D point clouds each contribute substantially.

6.4 Modality ablations

  • Robot state provides internal status โ€” can't be reconstructed from vision alone.
  • 3D point clouds enhance geometric understanding โ€” particularly on tasks with occluded or cluttered objects.
  • 2D images alone are insufficient for the hardest contact-rich tasks.

6.5 Generalization probe

Under object / background / lighting variations, FiS-VLA drops roughly 19โ€“31% โ€” indicating meaningful room for improvement on OOD robustness. Crucially, FiS-VLA degrades less than ฯ€โ‚€ on every axis (ฯ€โ‚€ drops ~27โ€“46%), so the robustness gap is the takeaway, not the absolute drop.


7. Limitations (authors' own + reviewer concerns)

7.1 Authors' stated limitations

  1. Statically-configured sharing. The number of shared blocks (2) and the frequency ratio (1:4) are fixed hyperparameters, not adaptive to task difficulty. Authors identify dynamic adaptation as future work.
  2. OOD performance drops 19โ€“29% under visual / scene perturbations.
  3. Still a cascade at the inter-tier level โ€” S2's latent condition has to be passed forward to S1, which reintroduces a synchronization boundary (just inside one network instead of between two).

7.2 Reviewer's concerns

  • No head-to-head with ฯ€0.6 or ฯ€0.7. FiS-VLA's real-world baselines are CogACT and ฯ€โ‚€ (not ฯ€0.5 / ฯ€0.6). The ฯ€0.6 + Knowledge Insulation recipe (same NeurIPS 2025!) is arguably a stronger baseline for the "same VLM handles both reasoning and action" thesis.
  • No LIBERO numbers. Standard benchmark is conspicuously absent โ€” only RLBench. Makes comparison with ICLR 2026 VLAs (which report LIBERO heavily) awkward.
  • Fixed sharing depth raises transferability questions. "Last 2 blocks are S1" works for the Prismatic / LLaMA2-7B 32-layer LLM. What about Gemma3-4B (ฯ€-series) or different architectures? Untested.
  • Reasoning-preservation claim is tested only on action-task performance. The paper shows โ„’_slow helps action accuracy, but doesn't show S2's reasoning outputs remain as good as the un-finetuned VLM on language-only benchmarks (e.g., MMMU, MMBench). VLM4VLA-style analysis would strengthen the case.
  • Open weights โ€” released post-publication (correction to earlier "closed" claim). Training/inference code is public (github.com/CHEN-H01/Fast-in-Slow, MIT) and a pretrained large-scale checkpoint was released on HuggingFace (haosad/fisvla) on 2025-07-08, so reproducibility is no longer the limitation it was at submission (project page: fast-in-slow.github.io).

8. Significance โ€” why FiS-VLA matters

FiS-VLA defines the "embedded dual-system" architectural pattern. Along with ChatVLA-2 (MoE-routed) and ThinkAct (RL-latent-plan), it establishes three concrete alternatives to the GR00T-style cascade that ICLR 2026 dual-system papers (WholeBodyVLA, HiMoE-VLA, AdaMoE) draw from.

Key structural contribution: parameter sharing between reasoning and action. This contrasts with two competing positions in the 2025 dual-system literature:

  • Separate networks + hand-designed interface (GR00T N1, RoboDual, Hi-Robot)
  • Sparse MoE within one network (ChatVLA-2, AdaMoE)

FiS-VLA's dense-sharing approach may generalize better to small-backbone regimes where MoE capacity is wasteful, while MoE may win at scale.

Efficiency contribution: 117.7 Hz control is production-relevant. Matches the ฯ€0.6 latency target (63 ms/chunk โ‰ˆ 16 Hz chunk rate = comparable) while producing longer chunks per forward pass.

Placement (see Review-VLA-Architecture ยง5.F): FiS-VLA is Category F (Hierarchical / Dual-system / MoE), sub-pattern "embedded."


9. Relation to this wiki

10. Links

โ† Back to NeurIPS-2025-Fast-in-Slow ยท NeurIPS-2025 ยท Home

โš ๏ธ **GitHub.com Fallback** โš ๏ธ