Review DuoCore FS - Heungwoo/research GitHub Wiki
In-Depth Review — DuoCore-FS: Asynchronous Fast-Slow VLA Policies for Whole-Body Robotic Manipulation
Paper: Asynchronous Fast-Slow Vision-Language-Action Policies for Whole-Body Robotic Manipulation Authors: Astribot Team — Core contributors: Teqiang Zou, Hongliang Zeng · Contributors: Yuxuan Nong, Yifan Li, Kehui Liu, Haotian Yang, Xinyang Ling, Xin Li · Project lead: Lianyang Ma Affiliation: Astribot (Shenzhen-based commercial whole-body manipulation startup; see
Astribot Suite) arXiv: 2512.20188 · v1, Dec 23 2025 Code / weights: None public. Abstract states implementation "provided to commercial users by Astribot as part of the Astribot robotic platform."
This page sits inside the VLA architectures review §5 Category F (hierarchical / dual-system / MoE) as the whole-body instance of the parallel-async fast-slow pattern, and pairs with Fast-in-Slow (intra-stack via shared parameters), AsyncVLA (inter-system cloud-edge), RTC (intra-model chunk-boundary), and π0.7 (subgoal-image async refresh) on the broader async-VLA axis.
Astribot's DuoCore-FS is the commercial-team disclosure of a truly parallel fast-slow VLA stack for whole-body manipulation. Four punchy verdicts:
- 3B PaliGemma slow pathway + Pi0-small-style flow-matching fast pathway, jointly trained end-to-end. The slow side is instantiated with π0-FAST built on the 3B PaliGemma backbone, autoregressively emitting RVQ-VAE action tokens; the fast side is a Transformer-based diffusion-policy decoder running flow-matching action chunks of length T=32.
- The interface between them is a written-and-read latent buffer, not shared parameters (Fast-in-Slow), not a re-conditioned action embedding (AsyncVLA), and not a subgoal image (π0.7). The buffer stores the VLM's instruction embeddings + learnable fusion-query embeddings + (optionally) discrete reasoning/action tokens, refreshed at 1–3 Hz by the slow side and consumed at 25–30 Hz by the fast side.
- 30 Hz whole-body chunk rate at a 3B-VLM scale — on a single NVIDIA RTX 4090, DuoCore-FS reaches 32.3 Hz, vs 12.5 Hz for the synchronous π0 baseline at comparable parameter size and vs 3.27 Hz for the slow-side-only baseline. The headline number is ≈3× the chunk frequency of synchronous fast-slow VLAs.
- Whole-body action tokenizer: residual-VQ-VAE with three parallel stream-specific codebooks (position / SO(3) rotation / gripper), each of size 1024, over a 29-dimensional whole-body action on a 25-DoF mobile dual-arm platform (Astribot S1). The position+rotation+gripper factorization with per-stream RVQ produces a fixed token sequence of length 36 vs FAST's average 81 / max 205 — directly enabling the slow-side autoregressive inference at 3.27 Hz.
Real-world: a single long-horizon popcorn-scooping task in a commercial popcorn-kiosk setting, 1,780 trajectories / 10.22 h of teleop data. In-distribution overall success 90% (DuoCore-FS) vs 85% (π0). Out-of-distribution overall success 50% vs 10%. Language-following on a co-collected beverage-cabinet-closing task: 42.9% vs 14.3%.
The paper is honest that the win over π0 is not in absolute success rate but in inference rate, OOD robustness, and language following. The architecture's central contribution is getting a 3B VLM to drive whole-body manipulation at 30 Hz without throttling the fast loop, not raising peak in-distribution success.
Why "truly async" matters specifically for whole-body manipulation. The paper makes a structural argument that synchronous dual-system VLAs hit a stiffer wall on whole-body than on tabletop:
- Joint count. The Astribot S1 platform exposes 25 DoF: 2×7-DoF arms + 4-DoF articulated torso + 2-DoF head + 3-DoF omnidirectional mobile base. The action tokenizer uses 29 dimensions at the controllable level (9 torso delta-pose + 9 left arm + 9 right arm + 2 grippers). Compared to a 7-DoF tabletop arm, the action space is ~4× larger and the action chunk inherits proportionally more dimensions to denoise.
- Motion-space breadth. Whole-body tasks involve torso rotation, bending, reaching across obstacles, and base-shift coordination — geometrically far broader than tabletop reach envelopes.
- Dynamic viewpoint. Head and torso rotation move the cameras; the visual stream is non-stationary across a single episode. A slow VLM cannot afford to drop frames at multi-second cadence.
In the synchronous dual-system pattern (DP-VLA, HiRT, RoboDual, G0-VLA, π0.5 as described in the paper's §2.2), the fast head must wait on the slow VLM's update. As the VLM scales toward 3B–7B, the locked-step frequency drops from "fast enough" to "unusable." Existing async work (FiS-VLA, OpenHelix, Helix, Hume) is critiqued for one of three reasons:
| Prior async work | Why DuoCore-FS authors reject it for whole-body |
|---|---|
| FiS-VLA / OpenHelix | Not truly parallel — fixed scheduling ratio (e.g., 1:4) couples fast to slow's discrete update cycle. |
| Helix (Figure AI) | Closest in spirit to parallel fast-slow, but not open-sourced; public materials describe only latent representations from slow, no published mechanism for reasoning or action-token generation at inference. |
| LCB (Shentu et al., IROS 2024) | Also "conceptually closer to fully parallel" per the paper's §2.2, but no open-source implementation; its slow subsystem is primarily a conversational/language-guidance interface and the paper reports no real-world robot evaluation. |
| Hume | Cascaded slow-fast, not end-to-end trainable; cascading prevents joint optimization of high-level reasoning and real-time control. |
DuoCore-FS positions itself as: parallel like Helix, end-to-end-trained like FiS, action-aware on the slow side like Hume, but without any of the three's individual limitation.
flowchart TB
subgraph SLOW["Slow pathway · 1-3 Hz"]
direction TB
Img["Multi-view images I_t<br/>(head + L-hand + R-hand,<br/>224x224)"] --> VLM["3B VLM<br/>(π0-FAST on PaliGemma-3B;<br/>Qwen2.5-VL-3B/7B drop-in)"]
Prop["Proprio q_t"] --> VLM
Lang["Task instruction l"] --> VLM
VLM --> SemOut["Reasoning outputs r_t<br/>- chain-of-thought traces<br/>- bbox predictions<br/>- subtask indicators<br/>- coarse RVQ action tokens"]
FQ["Learnable fusion queries q_psi"] --> VLM
VLM --> FuseEmb["Fusion-query embeddings"]
VLM --> InstrEmb["Instruction embeddings"]
end
subgraph BUF["Bridge buffer B_t<br/>(decouples generation vs consumption)"]
direction TB
InstrEmb --> BB["B_t = {instr_emb, fusion_emb}<br/>refreshed at 1-3 Hz<br/>consumed at 25-30 Hz"]
FuseEmb --> BB
SemOut -. optional .-> BB
end
subgraph FAST["Fast pathway · 25-30 Hz"]
direction TB
ImgHF["Multi-view images (high freq)"] --> Enc["Shared embedding projection"]
PropHF["Proprio q_t (high freq)"] --> Enc
BB --> Enc
InstrRaw["Raw instruction embeddings<br/>(also passed in directly)"] --> Enc
Noise["Gaussian noise epsilon"] --> Enc
Enc --> Tx["Multi-layer transformer encoder<br/>(Pi0-small style)"]
Tx --> DP["Flow-matching diffusion-policy<br/>velocity-field head v_theta"]
DP --> ChunkA["Action chunk A_{t:t+T-1}<br/>T=32, 29-dim whole-body"]
end
ChunkA --> Robot["Astribot S1 · 25 DoF<br/>(2x7 arms + 4 torso + 2 head + 3 base)"]
classDef slow fill:#bbdefb,stroke:#1565c0,color:#000
classDef buf fill:#fff9c4,stroke:#f57f17,color:#000
classDef fast fill:#c8e6c9,stroke:#2e7d32,color:#000
classDef rob fill:#f8bbd0,stroke:#ad1457,color:#000
class Img,Prop,Lang,VLM,SemOut,FQ,FuseEmb,InstrEmb slow
class BB buf
class ImgHF,PropHF,Enc,Tx,DP,ChunkA,InstrRaw,Noise fast
class Robot rob
The paper is explicit (§3.1 Slow system): the slow pathway "can be instantiated with models such as PaliGemma-3B, Qwen2.5-VL-3B, or Qwen2.5-VL-7B, depending on the computational budget and task requirements."
In the experimental section (§4.1 Training Details) the actual instantiation used in the reported results is given concretely:
"We instantiate the slow system with open-sourced π0-FAST, which is built on the 3B PaliGemma backbone and provides pre-trained weights for visuomotor control."
So the 3B VLM used in all reported numbers is π0-FAST on PaliGemma-3B, not Qwen2.5-VL. The Qwen2.5-VL variants are mentioned as drop-in options but no Qwen2.5-VL numbers are reported. This matters when comparing to peers — DuoCore-FS's slow pathway is the same backbone family as π0 / π0-FAST itself, which makes the head-to-head comparison fair and is presumably why π0 / π0-FAST were chosen as baselines.
The slow side is trained with a standard negative-log-likelihood objective:
where
The bridge buffer
-
Instruction embeddings — the VLM's hidden representation of the task instruction
$l$ . These are largely time-invariant (the instruction does not change across an episode) and serve as a stable linguistic prior. -
Fusion-query embeddings — outputs of learnable fusion queries
$q_\psi$ that attend across visual observations, proprioception, the task instruction, and the VLM's internal semantic space. These produce structured, manipulation-oriented latents.
Crucially, the parameters
Optionally, the buffer can also carry the slow side's discrete reasoning / coarse-action tokens. These are used in the reported experiments — the slow side autoregressively predicts RVQ-VAE action tokens that act as a "coarse-action prior" for the fast side.
Why also pass the raw instruction embedding directly into the fast side. §3.1 Bridge buffer: even though fusion-query embeddings aggregate task-relevant semantics, "the task instruction remains constant throughout execution, whereas relying solely on the task-related semantic information from the fusion queries may result in the loss of crucial task details or cause the representation to overly focus on patterns directly related to action generation." Hence: buffer holds the dynamic semantic latents; raw instruction is wired directly in. This is structurally similar to π0.7's per-component prompt dropout and to AsyncVLA's "both observations to the edge model" trick — they all stem from the same principle that a stale, condensed latent loses information that the original signal still has.
The fast pathway is described as a "Pi0-small style diffusion-policy network" (§3.1 Fast system). At each fast step
- All conditioning modalities — current visual features, proprioceptive state, latest fusion-query embeddings from
$B_{t-\Delta t}$ , raw instruction embedding — are projected into a shared embedding space and concatenated into a unified token sequence. - A multi-layer transformer encoder processes this sequence and emits a conditioning signal for the diffusion-policy head.
- A continuous-time flow-matching head
$v_\theta$ denoises a noisy action chunk via the velocity-field formulation$X_\tau = \tau \varepsilon + (1-\tau), A_{t:t+T-1}$ with objective $\mathcal{L}{\text{fast}}(\theta, \psi) = \mathbb{E}{A,\varepsilon,\tau,\Delta t}\bigl| \varepsilon - A_{t:t+T-1} - v_\theta(X_\tau, o_t, B_{t-\Delta t}) \bigr|_2^2$.
Note that
The action chunk length is T=32, the chunk's per-step dimension is 29 (whole-body, see §4 of this review), and the chunk frequency is 25–30 Hz in practice.
The tokenizer (§3.2) is a geometry-aware residual-VQ-VAE, inspired by RDT2's design, that produces the discrete action tokens consumed by the slow side. Architecturally:
- Cartesian delta-pose control for all major kinematic units (torso, left arm, right arm). The torso has a 9-dimensional delta pose (3 translational + 6 SO(3) continuous orientation). Both arms use the same 9-dim parameterization.
- Absolute opening-width for the grippers (2 dim).
- Total per-step dimension:
$3 \times 9 + 2 = 29$ .
An action chunk
| Stream | Dim per step | Tensor shape | Encoder | Reconstruction loss |
|---|---|---|---|---|
| Position |
9 | 1D conv encoder–decoder | ||
| Rotation |
18 (3 × 6D SO(3)) | 1D conv encoder–decoder | geodesic SO(3) loss | |
| Gripper |
2 | 1D conv encoder–decoder |
Each stream is then discretized through stream-specific RVQ with codebook size 1024. Each RVQ stack maintains its own codebook (i.e., three parallel codebooks, not a shared one).
Rotation-branch reconstruction uses geodesic distance on SO(3), not naive MSE on the 6D continuous representation:
The full tokenizer loss is $\mathcal{L}{\text{token}} = \mathcal{L}^{\text{pos}}{\text{rec}} + \mathcal{L}^{\text{rot}}{\text{rec}} + \mathcal{L}^{\text{grip}}{\text{rec}} + \mathcal{L}{\text{vq}}$, where $\mathcal{L}{\text{vq}}$ is the standard commitment + codebook objective summed over all quantizers.
After training, the RVQ-VAE encoders convert continuous action chunks into discrete token sequences of length 36 (fixed). This compactness is the load-bearing property — see the ablation in §8 below.
gantt
title DuoCore-FS · 1 second of execution
dateFormat X
axisFormat %S
section Slow VLM @ 1-3 Hz
VLM forward A : 0, 333
VLM forward B : 333, 666
VLM forward C : 666, 999
section Bridge buffer refresh
Buffer write A : milestone, 333, 0
Buffer write B : milestone, 666, 0
Buffer write C : milestone, 999, 0
section Fast diffusion @ 30 Hz
chunk 1 : 0, 33
chunk 2 : 33, 66
chunk 3 : 66, 99
chunk 4 : 99, 133
chunk 5 : 133, 166
chunk 6 : 166, 200
chunk 7 : 200, 233
chunk 8 : 233, 266
chunk 9 : 266, 300
chunk 10 : 300, 333
chunk 11 : 333, 366
chunk 12 : 366, 400
chunk 13 : 400, 433
chunk 14 : 433, 466
chunk 15 : 466, 500
chunk 16 : 500, 533
chunk 17 : 533, 566
chunk 18 : 566, 600
chunk 19 : 600, 633
chunk 20 : 633, 666
chunk 21 : 666, 700
chunk 22 : 700, 733
chunk 23 : 733, 766
chunk 24 : 766, 800
chunk 25 : 800, 833
chunk 26 : 833, 866
chunk 27 : 866, 900
chunk 28 : 900, 933
chunk 29 : 933, 966
chunk 30 : 966, 1000
Under a slow rate of 3 Hz and a fast rate of 30 Hz, the fast pathway consumes ~10 chunks per VLM update. The buffer is read every fast step; the buffer's content only changes at slow-pathway completion. The fast pathway operates on
The fast pathway "operates smoothly using the previously refreshed instruction and bridge embeddings" (§3.4). There is no inpainting at the chunk boundary (RTC), no rollback, no async re-conditioning. The fast pathway simply uses whatever is in the buffer when it executes, and trusts that the buffer's latents — which represent task-level intent rather than instantaneous geometry — remain valid across the 333 ms–1000 ms refresh window.
This is the architectural bet: the slow side's outputs are stable enough at sub-3-Hz cadence that the fast side does not need to re-cond on the stale observation. The bet is supported by:
- The slow side emits semantic and intent features (subtask indicator, bbox of the target, coarse RVQ-VAE action prior), which change at human-scale timescales.
- The fast side has direct access to the current observation
$o_t$ and proprioception$q_t$ for fine-grained motor control. - The fast side also gets the raw instruction embedding directly (not via the buffer), giving it a stable linguistic anchor independent of slow-side staleness.
Stage 2 of training simulates the deployed asynchronous timing. Given a slow-side observation
The maximum delay is computed from the slow and fast update periods
In the reported experiments the frame-level delay is sampled from
This is in the same spirit as π0.7's 240 ms RTC-style delay simulation during training, and as AsyncVLA's stage-2 joint fine-tune where the cloud model is trained on edge-delayed inference paths — but the three differ in what they simulate:
| Paper | What is delayed at train time |
|---|---|
| DuoCore-FS | Fast observation |
| π0.7 | Action chunk execution boundary (240 ms RTC delay simulation) |
| AsyncVLA | Cloud↔edge WiFi delay (0.28–6 s) baked into stage-2 joint loss |
| Fast-in-Slow | Asynchronous 1:4 sampling — S2 sees one frame, S1 sees 4 separate noised actions |
§3.4: a Jacobi-style parallel decoding strategy inspired by Consistency Large Language Models (CLLMs) is used on the slow side. This decodes multiple iterative refinements in parallel, reducing the effective per-step cost of the slow VLM without changing its semantic capacity. No numbers are given for the speedup from Jacobi decoding alone; the 3.27 Hz reported for the slow side is the combined number (PaliGemma-3B + RVQ-VAE tokens + Jacobi decoding).
Stage 1: slow-only. The slow VLM is trained independently to autoregressively predict the RVQ-VAE action tokens encoded by the action tokenizer.
- 30 epochs, 24 NVIDIA H100 GPUs, batch size 25 per GPU (effective batch size 600).
- Cosine-decay learning-rate schedule, peak LR
$1 \times 10^{-4}$ , minimum LR$3 \times 10^{-6}$ . - BF16 throughout.
Stage 2: fast + slow joint. Both pathways are jointly optimized under the cross-timescale sampling scheme (§4.3 above) with the unified loss
- 12 epochs, same hardware.
- Cosine schedule, peak LR
$5 \times 10^{-5}$ , minimum LR$3 \times 10^{-6}$ .
The 10:1 weight ratio favors the fast side because the slow side is already pretrained in stage 1 and the fast controller "requires stronger supervision to master precise whole-body action under asynchronous semantic guidance" (§3.3).
| Hyperparameter | Value |
|---|---|
| Slow backbone | π0-FAST on PaliGemma-3B |
| Action expert | Pi0-small-style flow-matching transformer |
| Input image resolution | 224 × 224 |
| Number of camera views | 3 (head, left-hand, right-hand) |
| Action chunk length |
32 |
| RVQ codebook size (per stream) | 1024 |
| Number of RVQ streams | 3 (position, rotation, gripper) |
| Slow rate | 1–3 Hz (reported 3.27 Hz) |
| Fast rate | 25–30 Hz (reported 32.3 Hz) |
| Frame-level delay |
|
| Loss weighting |
|
| Stage-1 epochs | 30 |
| Stage-1 peak LR | |
| Stage-2 epochs | 12 |
| Stage-2 peak LR | |
| Min LR (both stages) | |
| Precision | BF16 |
| Compute | 24× NVIDIA H100, BS=25/GPU |
| Inference compute | 1× NVIDIA RTX 4090 |
| Fast-pathway inference precision | BF16 (TensorRT-compiled) |
The paper does not state:
- Total training time (wall-clock or GPU-hours).
- Number of fast-pathway transformer layers or parameter count.
- Number of flow-matching denoising steps used at deployment.
- Total number of trainable parameters (only "3B VLM" is given for the slow side).
- License / weight release terms.
- Robot: Astribot S1 — mobile dual-arm manipulator. 25 DoF total: 2×7-DoF arms with parallel-jaw grippers, 4-DoF articulated torso, 2-DoF head, 3-DoF omnidirectional mobile base.
- Action dimensionality at the tokenizer: 29 (3×9 delta-pose + 2 gripper).
- Data collection: teleoperation.
- Dataset: 1,780 demonstration trajectories, 10.22 hours total, collected in a commercial popcorn-kiosk scenario.
- Tasks: one long-horizon popcorn-scooping task (4 sub-tasks) + one short-horizon beverage-cabinet-door-closing task (~1% of the data).
The long-horizon task is decomposed into four sequential sub-tasks executed under a single coarse instruction (e.g., "pick up the paper cup and scoop popcorn"). Subsequent stages become unreachable on earlier failures.
- Pick up the cup. Reach toward and grasp an empty cup on the serving table.
- Grab the scoop. Rotate the upper body while carrying the cup from the serving area to the popcorn machine, then grasp the scoop.
- Scoop the popcorn. Scoop popcorn with the right arm and pour into the cup.
- Place the cup on table. Turn back to the serving area and place the filled cup back on the table.
Each sub-task is evaluated as a conditional success rate — success is only computed over trials where the sub-task became reachable.
| Method | Pick up cup | Grab scoop | Scoop popcorn | Place cup | Overall S.R. ↑ | Inf. Hz |
|---|---|---|---|---|---|---|
| π0 | 95% (19/20) | 89.5% (17/19) | 100% (17/17) | 100% (17/17) | 85% (17/20) | 12.5 |
| DuoCore-FS-slow | 100% (20/20) | 75% (15/20) | 73.3% (11/15) | 100% (11/11) | 55% (11/20) | 3.27 |
| DuoCore-FS | 100% (20/20) | 90% (18/20) | 100% (18/18) | 100% (18/18) | 90% (18/20) | 32.3 |
Reading the numbers:
- The slow-side alone (3.27 Hz, autoregressive RVQ-VAE token prediction) underperforms π0 on fine-grained subtasks (scoop the popcorn, grab the scoop). The paper attributes this to RVQ-VAE tokens being a "lossy abstraction of continuous actions" — discretization hurts on alignment-heavy operations like grabbing the scoop handle.
- The full DuoCore-FS matches or slightly exceeds π0 in absolute success (90% vs 85%) while running at 2.6× the inference frequency (32.3 Hz vs 12.5 Hz).
- The headline "3× faster" claim is the abstract's rounding of the single π0 ratio (32.3/12.5 = 2.58×) to "approximately three times as fast as prior VLA models with comparable model sizes"; the conclusion phrases it as "over three times the inference frequency of prior VLA models [π0, π0.5]." It is not a geometric mean over multiple baselines — only π0 is timed head-to-head.
Cup placed at unseen locations (e.g., near the table edge):
| Method | Pick up cup | Grab scoop | Scoop popcorn | Place cup | Overall S.R. ↑ |
|---|---|---|---|---|---|
| π0 | 20% (2/10) | 50% (1/2) | 100% (1/1) | 100% (1/1) | 10% (1/10) |
| DuoCore-FS-slow | 70% (7/10) | 57.1% (4/7) | 25% (1/4) | 100% (1/1) | 10% (1/10) |
| DuoCore-FS | 50% (5/10) | 100% (5/5) | 100% (5/5) | 100% (5/5) | 50% (5/10) |
Notes:
- DuoCore-FS-slow outperforms π0 on the OOD "Pick up the cup" stage (70% vs 20%) — the paper attributes this to the autoregressive token reasoning generalizing better than the flow-matching baseline.
- DuoCore-FS combines the strengths of slow and fast — the slow side's reasoning bumps the OOD pickup, while the fast pathway preserves the fine-grained scoop accuracy that the slow side lacks. Overall OOD success rises from 10% (both π0 and slow-only) to 50%.
- 5× the OOD success rate of π0 is by far the largest single delta in the paper.
The "anomaly" tests evaluate whether the robot detects degenerate cup states ("cup fell over," "cup is upside down," "cup was taken away") and returns to home pose:
| Method | Cup fell over | Cup upside down | Cup taken away | Overall S.R. ↑ |
|---|---|---|---|---|
| π0 | 87.5% (7/8) | 87.5% (7/8) | 100% (8/8) | 91.7% (22/24) |
| DuoCore-FS | 100% (8/8) | 87.5% (7/8) | 100% (8/8) | 95.8% (23/24) |
Small but consistent improvement (95.8% vs 91.7%). The delta here is modest because the failure mode is high-level — both systems mostly detect the anomaly.
The popcorn-scooping and beverage-cabinet-closing tasks were collected with matched initial visual states. At test time, the same visual observation is paired with one of the two instructions, and the policy is scored on whether it executes the instructed behavior (not whichever behavior is more frequent in the data):
| Method | Close beverage cabinet door |
|---|---|
| π0 | 14.3% (1/7) |
| DuoCore-FS | 42.9% (3/7) |
3× the language-following success of π0. The paper attributes this to the slow side's autoregressive training objective encouraging consistent token-by-token reasoning over the instruction. The number is far from saturation (42.9% leaves 4 of 7 trials still failing) and the paper attributes the cap to extreme data imbalance — the beverage-cabinet task is only 1% of the training set and no imbalance-mitigation was applied. The typical failure mode is instruction confusion: the model executes "pick up the cup" when asked to "close the beverage cabinet door."
The 14.3% baseline for π0 is striking — π0 essentially ignores the cabinet instruction and defaults to the dominant popcorn-scooping behavior.
Reported on a single NVIDIA RTX 4090. π0 inference uses the official JAX implementation; DuoCore-FS uses PyTorch with the fast system compiled to TensorRT in BF16:
| System | Frequency (Hz) | Latency per chunk (ms) |
|---|---|---|
| π0 (synchronous baseline) | 12.5 | ~80 |
| DuoCore-FS-slow (autoregressive RVQ tokens, Jacobi decoded) | 3.27 | ~306 |
| DuoCore-FS (parallel async, fast pathway dominates) | 32.3 | ~31 |
So:
- The fast pathway emits whole-body 32-step action chunks every ~31 ms in steady state.
- The slow pathway updates the bridge buffer every ~306 ms (≈3× the fast period × 3, since one slow forward takes ~9 fast-loop steps' worth of wall-clock).
- The system-level chunk rate of 32.3 Hz is governed entirely by the fast pathway in steady state — the slow pathway runs in parallel and refreshes the buffer asynchronously.
This is "approximately three times as fast as prior VLA models with comparable model sizes" (abstract; the conclusion says "over three times … [π0, π0.5]"), where the specific timed peer is π0 (12.5 Hz). Ratio: 32.3 / 12.5 = 2.58×.
The paper does not report:
- Action chunk denoising-step count (number of Euler steps for the flow-matching head).
- Latency breakdown between vision encoder, fast transformer, and flow-matching head.
- Memory footprint of the bridge buffer.
- Behavior at sub-1 Hz slow-side update rates (e.g., if Jacobi decoding stalls).
- Performance under VLM swap to Qwen2.5-VL-3B or 7B.
The paper has one explicit ablation: action tokenizer (FAST vs the proposed RVQ-VAE). Both tokenizers are evaluated as the slow-side action representation on the popcorn-scooping task.
| Tokenizer | Avg token length | Max token length | Success rate | Inference Hz |
|---|---|---|---|---|
| FAST | 81.12 | 205 | 0% | 0.95 |
| RVQ-VAE | 36 | 36 | 83.3% | 3.27 |
Reading:
- FAST tokenizes the whole body so verbosely that the slow side cannot inference faster than 0.95 Hz — the autoregressive sequence is on average 2× longer than RVQ-VAE and up to 5× longer in the worst case.
- To make FAST tractable, the authors had to tokenize the three SO(3) components separately and predict left/right/torso action tokens sequentially. This sequential prediction induces compounding errors. In deployment FAST "frequently generates incorrect tokens, leading to erratic motions and no successful task executions (0% success)."
- RVQ-VAE's fixed 36-token length, per-stream codebooks, and geodesic SO(3) loss together deliver 83.3% success at 3.27 Hz — a 3.4× speedup and a step function in success.
This is the single most decisive number in the paper: it shows that the action tokenizer choice is the load-bearing piece for whole-body slow-side autoregressive inference. The FAST tokenizer, which works well for tabletop 7-DoF arms, simply does not scale to 25 DoF.
The paper does not ablate any of:
- Buffer presence. No "DuoCore-FS without fusion queries" baseline — i.e., what is the contribution of the latent buffer vs simply passing instruction-embedding + raw observation to the fast pathway?
- Asynchrony. No "DuoCore-FS at 1:1 frequency ratio" baseline — i.e., does the speedup come from the parallel inference or just from a faster fast pathway?
- End-to-end vs separate training. No "stage 2 only with fast pathway, slow frozen" baseline — i.e., do gradients flowing into the slow side via the fusion queries actually matter for final task success?
- Number of fusion queries. No sweep on how many learnable queries are needed.
- Slow VLM size. The paper claims Qwen2.5-VL-3B and 7B are drop-in options but reports no numbers for them.
- Slow rate. Range "1–3 Hz" is reported but no ablation at the 1 Hz vs 3 Hz end of the range.
-
Delay distribution.
$\Delta \sim \mathcal{U}[0, 25]$ is fixed; no sweep on$\Delta_{\max}$ .
These gaps are material. The paper's core architectural claim — that the buffer + parallel async + end-to-end joint training are each necessary — is not directly tested. The win over π0 in Table 1 (90% vs 85%) is small enough that any one of these ablations could plausibly close the gap.
The async / fast-slow VLA design space by April 2026 has at least five distinct architectural patterns, each attacking a different latency or staleness problem:
| Dimension | Where async lives | What is the interface? | How is staleness handled? |
|---|---|---|---|
| Intra-model (chunk-boundary) | Within a single flow-matching model | Frozen tail of executed chunk | Inpainting (RTC) |
| Intra-stack (shared parameters) | Inside one transformer | Shared blocks (last 2 of 32) | Fixed 1:4 ratio scheduling (FiS-VLA) |
| Intra-policy (parallel two-network) | Two co-located networks running in parallel | Latent buffer + raw instruction | Train-time delay augmentation (DuoCore-FS) |
| Inter-system (cloud-edge) | Across a network | Action-token embedding + dual observation | Two-stage joint fine-tune (AsyncVLA) |
| Inter-frame (subgoal cadence) | Across multi-second world-model refresh | Subgoal image + prompt-dropout | 4 s async refresh + CFG (π0.7) |
DuoCore-FS occupies the intra-policy parallel slot — two co-located networks (slow VLM + fast diffusion expert) running concurrently, decoupled by a written-and-read latent buffer, trained jointly end-to-end. Closest in spirit to Helix (Figure AI's claimed parallel slow-fast, but closed) and to Hume (cascaded slow-fast, but not end-to-end). DuoCore-FS is the first open-architecture-disclosed parallel + end-to-end instance.
| Axis | DuoCore-FS | AsyncVLA | Fast-in-Slow | RTC | π0.7 dual-async | GR00T N1.x | Helix-02 | π0.5 / Hi Robot |
|---|---|---|---|---|---|---|---|---|
| Where async lives | Intra-policy (parallel two-network) | Inter-system (cloud↔edge) | Intra-stack (shared blocks 31–32 of 32) | Intra-model (chunk boundary) | Inter-frame (subgoal refresh) | Intra-model (DiT cross-attn) | Intra-stack (S0 1 kHz prior + S1 + S2) | Intra-machine (NL subtask) |
| VLM update rate | 1–3 Hz | 5 Hz (remote 4090) | full VLM at S2 rate; S1 at 117.7 Hz | n/a (no async VLM split) | 0.25 Hz BAGEL subgoal refresh + main 5 Hz | ~5–10 Hz | ~5 Hz | ~5–10 Hz |
| Action update rate | 25–30 Hz chunk emission | 8 Hz edge adapter + 10 Hz PD | 117.7 Hz @ chunk=8 (4090) | up to one chunk per period | 50 Hz / 20 Hz | 30–50 Hz | 200 Hz S1, 1 kHz S0 | 50–100 Hz |
| Slow-to-fast interface | Latent buffer (fusion queries + instruction embed) + optional RVQ action tokens | Stale action-token embedding from cloud, re-conditioned with |
Shared parameters (blocks 31–32) + S2 latent condition | Frozen tail of previous chunk for inpainting | Subgoal image + NL subtask + metadata, each with dropout | DiT cross-attention into VLM hidden states (layer 12 mid-tap in N1; vlln + AlternateVLDiT in N1.6; vl_self_attention in N1.7) | Latent vector from S2; 1 kHz neural prior from S0 | NL subtask string |
| Backbone identity | π0-FAST on PaliGemma-3B (3B) | OmniVLA 8.26B (SigLIP+DINOv2+LLaMA-2-7B) | LLaVA-class 32-layer LLM | n/a | Gemma3-4B (~4B) + BAGEL-14B world model | Eagle-2.5 (N1.5) / Cosmos-Reason2-2B (N1.7) | Closed | π0/π0.5 PaliGemma backbone |
| Action expert identity / size | Pi0-small-style flow-matching transformer (size not stated) | Edge Adapter 76M | Last 2 of 32 LLM blocks (shared) + diffusion head | Whichever underlying flow VLA | 860M flow-matching | DiT ~1B (N1.7) | Closed | π0/π0.5 action expert |
| Whole-body vs single-arm | Whole-body 25 DoF (29-dim action) | Single ground-robot (Vizbot, 3-DoF base) | Single-arm Franka + dual-arm AgileX/AlphaBot | Backbone-dependent | Bimanual UR5e / mobile manipulation | Humanoid 35+ DoF | Humanoid (Figure 02) | Bimanual mobile manipulation |
| Trained jointly or separately | Joint end-to-end (stage 1 slow → stage 2 joint) | Joint end-to-end (stage 1 edge → stage 2 joint) | Joint end-to-end (dual-aware co-training) | Plug-and-play, no retraining required | Joint end-to-end on prompt components | Joint end-to-end | Closed | Joint end-to-end |
| Handles network / latency delay | Train-time delay sampling |
Train-time WiFi delay 0.28–6 s, two-stage fine-tune | 1:4 fixed scheduling, asynchronous sampling at train time | Async chunk inpainting at boundary | 240 ms RTC delay simulation at train; 4 s subgoal refresh | None explicit | Closed | None explicit |
| Empirical headroom vs sync | 2.6×–3× chunk Hz vs π0; +5 pp ID, +40 pp OOD, +28.6 pp lang-following | +40 pp success vs cloud-only / edge-only on Vizbot nav under 0.28–6 s WiFi | +8% sim, +11% real vs CogACT/π0; 117.7 Hz | Smooth boundaries at no accuracy cost | Out-of-the-box matches RL specialists; first compositional generalization | Per release; N1.7 hits 30+ Hz on Jetson Thor | Closed | Hi Robot enables open-ended instruction following |
| Hardware target | Single RTX 4090 | Remote 4090 + Jetson Orin 30W | Single 4090 | Backbone-dependent | Workstation-class | Jetson Thor / Jetson Orin | On-board Figure 02 compute | Workstation-class |
| Open-source posture | Closed weights; commercial licensing via Astribot | Closed implementation; paper-documented | Closed weights; project page | Open algorithm | Closed model | N1.7 weights Apache-2.0 (first commercial-grade open generalist humanoid) | Fully closed | Closed |
The architecture is the first publicly-disclosed-mechanism instance of three properties simultaneously:
- Truly parallel slow-fast inference, not fixed-ratio scheduled.
- End-to-end joint training of both pathways with gradients flowing through the buffer via the fusion queries.
- Whole-body 25-DoF target with a tokenizer co-designed for the joint count.
Each property exists individually in prior work:
- Truly parallel — Helix (closed) and Hume (cascaded, not E2E) are the antecedents.
- Joint E2E — Fast-in-Slow (shared params, fixed 1:4) and π0.5 (NL subtask string, intra-machine).
- Whole-body action tokenizer — RDT2's residual VQ-VAE design is the explicit inspiration.
DuoCore-FS's contribution is putting all three together in a working stack, with the latent buffer as the load-bearing piece. The buffer is what enables the parallelism (it decouples generation from consumption), and the fusion-query gradients are what enables the joint training (they back-propagate fast-side action loss into the slow side's intermediate representations).
Architecturally closest peer: Fast-in-Slow in terms of the intra-stack vs intra-policy distinction, and Helix in terms of the parallel slow-fast pattern. Hume is the closest peer in terms of action-aware slow side. None of the three combines all of "parallel + end-to-end + whole-body + open architecture description."
The paper is not the one-async-pattern-to-rule-them-all. It explicitly does not address:
- Network-scale latency (WiFi seconds, cellular jitter). That is AsyncVLA's space. DuoCore-FS assumes both pathways live on the same machine — the 3.27 Hz vs 32.3 Hz decoupling is meaningful only if both processes share memory.
- Chunk-boundary smoothness. The fast pathway emits 32-step chunks at 30 Hz; whether the boundary between chunks is smooth or jittery is not addressed. RTC's inpainting trick is orthogonal and stackable.
- Long-horizon foundation-world-model planning. DuoCore-FS has no subgoal-image refresh, no future-frame prediction, no Goal world model. That is π0.7's space (BAGEL-14B at 4 s cadence).
- System-0 reflex layer. No 1 kHz balance or actuator prior beneath the fast pathway. Whole-body humanoids that need such a layer (Figure 02, Sharpa CraftNet) would need to stack one underneath.
- Tactile/force loop. The paper's own future-work section flags this — incorporating high-frequency force and tactile signals is identified as a future direction. The current architecture is vision + proprioception only.
Each async-VLA paper attacks a distinct sub-problem. DuoCore-FS's contribution is intra-stack frequency decoupling for whole-body manipulation.
Genuinely novel:
- The explicit latent buffer as a differentiable interface that the fast side reads at high frequency and the slow side writes at low frequency, with the fusion queries trained through the fast-side action loss. This is structurally different from Fast-in-Slow's parameter sharing and from AsyncVLA's stale-embedding-plus-re-conditioning. Whether it is materially better than these alternatives is not tested by ablation.
- The whole-body action tokenizer with three parallel stream-specific RVQ codebooks and geodesic SO(3) loss. This is the single most data-supported architectural contribution in the paper — the FAST vs RVQ-VAE ablation shows a 0% → 83.3% step function in slow-side success and a 3.4× speedup. For 25-DoF whole-body manipulation specifically, this matters.
Incremental improvements:
-
End-to-end joint training of fast and slow pathways is not new — Fast-in-Slow's dual-aware loss does the same thing through different architecture, π0.5 / Hi Robot trains both pathways together through the NL-subtask interface, GR00T N1.x co-trains the VLM and DiT. DuoCore-FS's
$\lambda_{\text{slow}}=1$ ,$\lambda_{\text{fast}}=10$ weighting is the only specific contribution here, and it is reported without ablation. - Two-stage training (independent slow → joint) is a standard recipe (Knowledge Insulation, Qwen-VLA T2A→CPT, AsyncVLA stage 1→2).
- Jacobi-style parallel decoding on the slow side is taken directly from CLLMs without modification.
Where evidence is strongest:
- The 3× inference speedup vs π0 (12.5 Hz → 32.3 Hz on the same RTX 4090) — directly measured, comparable platform, similar parameter scale on the VLM side.
- The FAST vs RVQ-VAE comparison (0% success at 0.95 Hz vs 83.3% at 3.27 Hz on the slow side) — clearest empirical signal in the paper.
- The OOD improvement (10% → 50% overall) — large enough delta that it is unlikely to be noise.
Where evidence is weakest or absent:
- No buffer ablation. The "latent buffer" is the central architectural claim, but no experiment compares DuoCore-FS to (a) DuoCore-FS without the fusion queries (instruction embed only), (b) DuoCore-FS without the buffer's bidirectional gradient (frozen fusion queries), or (c) DuoCore-FS with stale raw observations instead of latent embeddings. Any of these would isolate the buffer's contribution.
- No asynchrony ablation. Whether the speedup comes from parallelism, from the fast-pathway's specialized lightweight architecture, or from TensorRT compilation is not isolated. A "DuoCore-FS run synchronously at 1:1" baseline is missing.
- No baseline against FiS-VLA, Hume, or OpenHelix on the same task. Only π0 / π0-FAST are reported as baselines. The paper differentiates itself from FiS-VLA / Hume / Helix at length in §2.2, but does not run any of them on the popcorn-scooping task.
- Single task, single platform. All numbers come from one long-horizon kiosk task (with the beverage cabinet as a 1% side task). The whole-body generalization argument depends on the platform's joint count, not the diversity of tasks.
- No reported compute budget. Training cost (H100-hours, wall-clock) is not stated.
- No weight release. Implementation "provided to commercial users by Astribot" — reproduction is gated.
Mapping DuoCore-FS onto the S0/S1/S2 taxonomy:
- System 2 (slow, deliberative, 1–3 Hz): the π0-FAST 3B VLM with multi-faceted reasoning outputs (CoT, bbox, subtask indicator, RVQ action tokens) and fusion-query embeddings.
- System 1 (fast, intuitive, 25–30 Hz): the Pi0-small-style flow-matching diffusion-policy decoder, emitting 32-step whole-body action chunks at 30 Hz.
- System 0: not present. DuoCore-FS has no separate reflex / balance / safety / tactile layer beneath the fast pathway. The 30 Hz chunk emission is the lowest tier.
This places DuoCore-FS in the explicit-S1/S2-only column of the System 0/1/2 review, alongside Figure Helix v1 and the GR00T N1 → N1.7 lineage. The 4-DoF articulated torso + 3-DoF mobile base in Astribot S1 would benefit from an S0 balance prior for highly dynamic motion — but the kiosk-scenario task evaluated in the paper is essentially quasi-static (kiosk operation, no walking), so this gap is not stressed.
Suggested System-0/1/2 review row:
| Stack | S2 | S1 | S0 | Posture |
|---|---|---|---|---|
| DuoCore-FS (Astribot) | π0-FAST 3B VLM @ 1–3 Hz | Pi0-small flow-matching, 25–30 Hz, whole-body 25 DoF | None | Explicit S1/S2; commercial closed implementation |
The paper's §5 Conclusion explicitly identifies four future directions, which double as stated limitations:
- Slow-pathway capability is incomplete. Future work: multi-task foundation-model pretraining (subtask prediction, object grounding, affordance reasoning), more diverse real-world manipulation data, chain-of-thought reasoning at inference time.
- Fast pathway can be accelerated further. Reducing the number of flow-matching denoising steps or moving toward single-step generation would push the chunk rate higher.
- Bridging mechanism is rudimentary. "More effective and sophisticated bridging mechanisms between the slow and fast pathways" are identified as future work — implicit admission that the current latent buffer is a starting point, not a fully-optimized interface.
- No tactile / force. "Incorporating high-frequency force and tactile signals could improve responsiveness in contact-rich tasks."
- Task coverage. "Evaluating the framework on a broader set of real-world tasks, including dynamic scenarios and long-horizon manipulation tasks, will be important to fully validate its robustness and generalization."
Additional concerns not explicitly flagged by the authors:
- Single-task evaluation. The headline numbers come from one popcorn-kiosk task, evaluated on 20 in-distribution and 10 out-of-distribution trials. Per-cell statistical confidence is low.
- No ablation of central claims. As detailed in §8.2 above — the latent buffer, asynchrony, joint training, and slow VLM choice are not directly tested.
- No head-to-head against the cited differentiators (FiS-VLA, Hume). The introduction's careful positioning vs Fast-in-Slow / Helix / Hume is rhetorical, not empirical.
- Closed implementation. The implementation is "provided to commercial users" — reproduction is gated, and the architectural details cannot be verified by independent re-implementation.
-
Frame-level delay range is task-specific.
$\Delta \sim \mathcal{U}[0,25]$ frames is tied to the camera frame rate and slow:fast period ratio used. The recipe for choosing$\Delta_{\max}$ for a different platform is not given. - No latency analysis under VLM swap. The paper claims Qwen2.5-VL-7B is a drop-in replacement, but a 7B VLM at the same 3.27 Hz slow-rate would require either a much bigger compute budget or longer per-step latency — neither is quantified.
- No multi-task training results. All 1,780 trajectories are from the popcorn kiosk; multi-task scaling is not exercised.
- No comparison to π0.5 / π0.6 / π0.7. The π0 baseline is one model generation behind the current PI production stack. A π0.5 or π0.6 baseline on the same task would test whether the gain over "comparable model size" holds against current SOTA.
DuoCore-FS is Astribot's public reveal of the VLA stack architecture that drives the Astribot S1 platform. The lineage from Astribot Suite (the platform paper, July 2025) is direct: the suite paper describes the hardware, teleop interface, and deployment workflow; DuoCore-FS describes the VLA policy that runs on top.
The release pattern is comparable to:
- NVIDIA GR00T N1.x — open weights (N1.7 Apache-2.0) for a commercial-grade humanoid generalist.
- Figure Helix-02 — architectural blog post, no weights, closed deployment.
- Genesis AI GENE-26.5 — no weights, no arXiv, blog post only, closed.
- Physical Intelligence π0.x — open code (openpi), open weights (π0, π0-FAST, π0.5), closed for the production models (π0.6, π0.7).
DuoCore-FS sits between Helix-02 (architectural disclosure, no weights, no implementation) and π0 (full open release, paper + weights + code). Astribot publishes the architecture and training recipe with enough detail to understand the design, but gates the implementation to commercial users. The 3B-VLM-at-30-Hz-whole-body number is the empirical contribution; the latent-buffer + whole-body-tokenizer combination is the architectural contribution.
The paper's central novel claim — that a truly parallel fast-slow architecture trained end-to-end with a written-and-read latent buffer can drive 25-DoF whole-body manipulation at 30 Hz with a 3B VLM — is plausible and consistent with the reported numbers, but is not isolated from confounders by the ablations included.
For the field, the most actionable contribution is the whole-body RVQ-VAE action tokenizer with stream-specific codebooks. The FAST-vs-RVQ-VAE comparison directly motivates designing per-stream codebooks for any platform beyond 14 DoF.
- arXiv: 2512.20188 (v1, Dec 23 2025)
- Astribot Suite (platform paper): arXiv 2507.17141
- Astribot website: astribot.com (commercial)
- Project email: [email protected]
- Code / weights: none public; commercial licensing only
- AsyncVLA — Hirose/Levine cloud↔edge async at WiFi seconds scale
- Fast-in-Slow — embedded dual-system via shared transformer blocks (the intra-stack alternative)
- Real-Time Chunking — async chunk inpainting (the intra-model alternative)
- π0.7 long-form — subgoal-image async refresh at 4-second cadence (the inter-frame alternative)
- GR00T series (N1 → N1.7) — NVIDIA cascaded dual-system humanoid lineage
- System 0 / 1 / 2 for Humanoids — the cognitive-trichotomy framing
- VLA Architectures §5 Category F — hierarchical / dual-system / MoE family
- π0.5 — hierarchical subtask interface (NL subtask string as S2/S1 interface)
- Knowledge Insulation — the gradient-routing alternative for VLM preservation under action training
← Back to Home · Review-VLA-Architecture · Review-System-0-1-2