Review AsyncVLA - Heungwoo/research GitHub Wiki
Paper: AsyncVLA: An Asynchronous VLA for Fast and Robust Navigation on the Edge Authors: Noriaki Hirose · Catherine Glossop · Dhruv Shah · Sergey Levine Affiliations: UC Berkeley (Hirose, Glossop, Levine) · Toyota Motor North America (Hirose) · Princeton (Shah) arXiv: 2602.13476 · Feb 13, 2026
This page sits alongside Review-System-0-1-2 (hierarchical-control taxonomy) and Review-VLA-Architecture (dual-system category) as the navigation instance of cloud-edge VLA decomposition, and complements RTC (intra-model latency) on the inter-model latency axis.
Hirose, Glossop, Shah, and Levine (Berkeley + Toyota + Princeton) build a two-policy navigation stack in which:
- An 8.26 B-parameter OmniVLA foundation model (ICRA 2026) runs on a remote workstation at 5 Hz.
- A 76 M-parameter Edge Adapter runs on the robot (Jetson Orin, 30 W) at 8 Hz.
- The two communicate over WiFi with up to 6 seconds of latency.
The Edge Adapter receives stale action-token embeddings from the cloud, re-conditions them on the most recent observation plus the same delayed observation the cloud model used, and emits a refined action chunk (8 steps at a 3 Hz control rate) that the PD controller (10 Hz velocity commands) executes.
End-to-end fine-tuning teaches the cloud model to emit guidance that is robust to being interpreted by the small edge model under delay, and a trajectory-re-weighting strategy emphasizes the dynamic-interaction samples (collision avoidance, pedestrian yielding) where the async correction matters most.
Result on a Vizbot ground robot in dynamic indoor + outdoor environments: 85% success vs. 45% (cloud-only) and 25% (edge-only) — a +40 pp absolute improvement over the best single-system baseline (cloud-only, 45%), with >10× fewer dynamic collisions (0.10 vs. 1.05) and faster goal reaching (59.18 s vs. 70.73 s cloud-only / 80.07 s edge-only).
The paper's headline contribution is the first hierarchical VLA explicitly designed for seconds-scale latency — every prior dual-system VLA (Fast-in-Slow, π0.5, Hi Robot, GR00T) assumes the S2↔S1 hand-off is < 300 ms because both systems live on the same machine.
By Q1 2026, "dual-system" had become a default architecture pattern. The publicly defined stacks are:
| Stack | S2 ↔ S1 | Latency assumption |
|---|---|---|
| Fast-in-Slow | shared transformer (last 2 of 32 blocks repurposed) | sub-ms (same forward pass) |
| π0.5 / Hi Robot | NL subtask string | < 300 ms |
| GR00T N1.x | DiT cross-attention into LLM features | < 100 ms |
| Figure Helix-02 | 1 kHz neural prior | < 1 ms |
| Sharpa CraftNet | 100 Hz tactile loop | < 10 ms |
All of these assume co-location. AsyncVLA is the first to publicly target the regime where S2 lives in the cloud and S1 lives on the robot, separated by a flaky WiFi link. That is the realistic deployment scenario for:
- Mobile robots (delivery, security, last-mile) that cannot carry a 4090.
- Humanoid teleop fleets where the operator's "guidance VLA" runs in a datacenter.
- Multi-robot fleets that share a single big VLA across many edge bodies.
If that deployment model becomes standard — and the trend lines for ICLR-2026-Human-Video-Pretraining, ICLR-2026-RoboArena, and the increasingly large LBM backbones suggest it will — AsyncVLA's recipe is the canonical reference.
flowchart TB
subgraph Remote["Remote workstation · NVIDIA RTX 4090"]
direction TB
G["Goal: 2D pose OR language"] --> BVLA
HX["Past observation I(t−k)"] --> BVLA["OmniVLA · 8.26B<br/>SigLIP + DINOv2 + LLaMA-2-7B<br/>5 Hz"]
BVLA --> Z["Action-token embedding z(t−k)"]
end
Z -- "WiFi · 0.28–6.0 s" --> NET(("stale<br/>delivery"))
subgraph Edge["Robot · Jetson Orin 30W"]
direction TB
NET --> EA["Edge Adapter · 76M params · 8 Hz"]
OBSC["Recent observation I(t)"] --> EA
HXC["Same delayed I(t−k) the cloud used"] --> EA
EA --> AC["Refined action chunk a(t) … a(t+H)"]
AC --> PD["PD controller · 10 Hz"]
PD --> ROB["Vizbot"]
end
classDef remote fill:#bbdefb,stroke:#1565c0,color:#000
classDef edge fill:#c8e6c9,stroke:#2e7d32,color:#000
classDef net fill:#fff9c4,stroke:#f57f17,color:#000
class BVLA,Z,G,HX remote
class EA,OBSC,HXC,AC,PD,ROB edge
class NET net
This is the structural insight. A naive design would feed the Edge Adapter only the most recent observation
By giving the edge model both
- Translate the cloud's guidance from the past frame into a goal-direction in the current frame.
- Detect what has changed in
$\Delta = I_t - I_{t-k}$ and react to it (a pedestrian stepped in, a door closed). - Maintain semantic intent (cloud-level) while updating spatial detail (edge-level).
That separation is exactly the System-2/System-1 division of labor — but distributed over a network rather than over a transformer.
| Component | Params | Hardware | Frequency |
|---|---|---|---|
| OmniVLA base VLA | 8.26 B | Remote RTX 4090 | 5 Hz |
| Edge Adapter | 76 M | Onboard Jetson Orin (30 W) | 8 Hz |
| PD controller | — | Onboard | 10 Hz |
| WiFi link | — | — | variable, 0.28–6.0 s |
| OmniVLA-edge baseline | 108 M | Onboard | 6 Hz |
Note that the Edge Adapter (76 M) is smaller than the OmniVLA-edge baseline (108 M) — the gain is not from a bigger edge model, but from the decomposition that lets the edge model focus on reactive correction while the cloud handles semantics.
Architecturally the Edge Adapter is a lightweight ViT with a token projector that compresses the cloud's stale action-token embedding from 4×4096 → 1024 dimensions before fusing it with the edge observations.
flowchart LR
D["Training trajectories"] --> CLF{"chunk Δpose > d_th = 1.0 m ?"}
CLF -- yes --> UP["Up-weight in loss<br/>(reactive sample)"]
CLF -- no --> NORM["Normal weight"]
UP --> L["Imitation + smoothing loss"]
NORM --> L
classDef d fill:#e1bee7,stroke:#6a1b9a,color:#000
classDef step fill:#fff9c4,stroke:#f57f17,color:#000
class D d
class CLF,UP,NORM,L step
A chunk is flagged as "reactive" if the final-pose distance from the start of the chunk exceeds 1.0 m. These flagged samples — which contain large path corrections, dynamic-obstacle responses, or pedestrian yielding — are up-weighted in the loss. The training therefore concentrates capacity on the cases where the Edge Adapter's correction over the cloud guidance is exactly what's needed.
This is conceptually similar to the hard-example mining tradition in detection / classification, but applied at chunk level for navigation.
| Stage | Trainable | Frozen | Goal |
|---|---|---|---|
| 1 | Edge Adapter (θ, φ) | OmniVLA (ψ) | Edge model learns to interpret stale embeddings |
| 2 | Edge Adapter + OmniVLA (ψ) | nothing | Cloud model learns to emit guidance robust to being consumed by the small edge model under delay |
Both stages optimize the same composite loss:
where
Why stage 2 matters. Without it, the cloud model is optimized only for "produce the correct action chunk if executed immediately" — but at deployment the chunk will be interpreted by a small edge model after up to 6 seconds of delay. Stage 2 bakes the asymmetric inference path into the cloud model's training distribution. This is a navigation-domain version of Knowledge Insulation's thesis that the deployment data path must be reflected in the training data path.
GNM (an aggregate of six public navigation datasets) + LeLaN + SACSoN (HuRoN, for pedestrian scenarios) — the Levine-lab navigation dataset stack. Indoor + outdoor, language and pose goals, real human-co-presence trajectories. Training used five H200 GPUs at an effective batch size of 180.
Vizbot ground robot:
- Fisheye camera (forward-facing).
- Tracking camera (visual-inertial pose).
- Wheel-velocity sensors.
- Onboard compute: NVIDIA Jetson Orin in 30 W mode.
flowchart LR
subgraph T1[Task 1: 2D-pose-conditioned navigation]
direction TB
T1a[10 cluttered environments<br/>indoor + outdoor]
T1b[12 to 30 m goals]
T1c[Pedestrians + dynamic obstacles]
end
subgraph T2[Task 2: Language-conditioned navigation]
direction TB
T2a[12 environments<br/>offices · kitchens · halls]
T2b[5 to 20 m goals]
T2c[Out-of-distribution prompts included]
end
| Baseline | Notes |
|---|---|
| OmniVLA (8.26 B remote, 5 Hz) | Cloud-only, the obvious "use the big model" baseline. |
| OmniVLA-edge (108 M onboard, 6 Hz) | Edge-only, the obvious "use a fast small model" baseline. |
| Ours w/o E2E | AsyncVLA architecture but without stage-2 joint fine-tuning — isolates the contribution of stage 2. |
| Ours (workstation) | AsyncVLA architecture but both halves on the workstation — isolates the value of cloud-edge separation vs. just having two heads. |
To stress-test, WiFi latency is varied from 0.28 seconds (clean lab) to 6.0 seconds (worst-case packet loss).
| Method | Success rate | Static collisions | Dynamic collisions | Time to goal (s) |
|---|---|---|---|---|
| Method | SR↑ | Static col.↓ | Dynamic col.↓ | Time to goal (s)↓ |
| --- | --- | --- | --- | --- |
| OmniVLA-edge (108 M onboard, 6 Hz) | 25% | 0.60 | 1.00 | 80.07 |
| OmniVLA (8.26 B remote, 5 Hz) | 45% | 0.30 | 1.05 | 70.73 |
| Ours w/o E2E (8.27 B + 76 M) | 25% | 0.60 | 1.05 | 82.78 |
| Ours (workstation) (8.27 B + 76 M, WS only) | 30% | 0.70 | 0.50 | 89.79 |
| AsyncVLA | 85% | 0.10 | 0.10 | 59.18 |
+40 pp absolute success gain over the best single-system baseline (cloud-only, 45%). ~10× reduction in dynamic collisions (0.10 vs. 1.00–1.05) and a 3–7× reduction in static collisions. Faster goal reaching (59.18 s vs. 70.73 s cloud-only).
Note the striking ablation result: neither the no-E2E variant (25% SR) nor the workstation-only variant (30% SR) comes close to the full system (85%) — both the stage-2 joint fine-tuning and the actual cloud-edge asynchronous split are individually necessary. Removing either collapses success to baseline levels, even though those variants still have the same two-model architecture.
The success rate stays high through the full 0.28–6.0 s latency range. Both single-system baselines collapse:
- OmniVLA: success drops as latency grows (chunks are increasingly stale).
- OmniVLA-edge: success is flat but low (no semantic guidance to use).
AsyncVLA's success curve is flat at the 85% range — the system is genuinely robust to multi-second delay.
- Stage-2 end-to-end fine-tuning — decisive: removing it (Ours w/o E2E) collapses SR from 85% to 25%, the level of the edge-only baseline. The cloud model genuinely must be trained to emit edge-friendly embeddings.
-
Both observations in the edge model (
$I_t$ +$I_{t-k}$ ) — required; ablating$I_{t-k}$ breaks the embedding interpretation. - Trajectory re-weighting — material on dynamic-obstacle subset, smaller on static.
- Cloud-edge vs. workstation-only — Ours (workstation) (both halves co-located) reaches only 30% SR, vs. 85% for the deployed async split. The asynchronous high-frequency edge refinement is itself part of the win, not just having two heads.
flowchart TB
subgraph LAT[Latency-handling axis]
direction LR
INTRA[Intra-model latency<br/>chunk boundary] --- INTER[Inter-system latency<br/>cloud ↔ edge]
end
RTC[RTC<br/>NeurIPS 2025] -.-> INTRA
PI07[π0.7 async subgoal<br/>4s refresh] -.-> INTER
ASYNC[AsyncVLA<br/>this paper] -.-> INTER
FIS[Fast-in-Slow] -.-> INTRA
HEL[Helix-02 1kHz] -.-> INTRA
classDef paper fill:#fff9c4,stroke:#f57f17,color:#000
classDef this fill:#c8e6c9,stroke:#2e7d32,color:#000
class RTC,PI07,FIS,HEL paper
class ASYNC this
| Paper | Latency axis | S2 ↔ S1 separation | Latency tolerance |
|---|---|---|---|
| Fast-in-Slow | intra-model | shared transformer blocks | sub-ms |
| Helix-02 | intra-stack | dedicated 1 kHz prior | < 1 ms |
| π0.5 / Hi Robot | intra-machine | NL subtask | < 300 ms |
| GR00T N1.x | intra-model | cross-attn at layer 12 | < 100 ms |
| π0.7 async subgoal | inter-system (sort of) | subgoal image every ~4 s | ~4 s |
| RTC | intra-model | async inpainting at chunk boundary | one chunk |
| AsyncVLA (this paper) | inter-system (cloud ↔ edge) | full model split | up to 6 s |
AsyncVLA is the only publicly demonstrated VLA that handles the cloud-edge case at seconds scale.
The closest cousin is π0.7's async subgoal-image refresh — but in π0.7 the action expert and the BAGEL-14B subgoal generator are still co-located; only the subgoal image is refreshed asynchronously. AsyncVLA splits the action policy itself across the network.
The closest manipulation analog would be a hypothetical "remote π0.7 + on-robot Edge Adapter for π0.7" — which doesn't yet exist publicly and would be the obvious follow-up paper for a manipulation team.
The wiki is mostly manipulation, so the natural question is: does the AsyncVLA recipe transfer?
The structural arguments suggest yes, with caveats:
- The cloud-edge bottleneck is the same. A π0.5-class system at 8 B parameters cannot run at 50 Hz on a Jetson; either you ship a 4090 with each robot, or you accept WiFi delay.
-
The two-observation re-conditioning trick generalizes. Manipulation has even more visual structure (gripper occlusion, bimanual coordination) — the edge model would benefit at least as much from
$I_t$ +$I_{t-k}$ as the navigation model does. - The trajectory-re-weighting trick has a manipulation analog: up-weight chunks where the gripper makes contact, where the trajectory has high curvature, or where bimanual coordination shifts.
The caveats:
- Manipulation is higher-frequency than navigation (50–200 Hz typical for π0.5-class action expert, vs. 8 Hz for AsyncVLA's edge). The PD controller's 10 Hz is fine for navigation but insufficient for fine manipulation. A manipulation port would need a higher-rate edge model.
- Manipulation has tighter precision — multi-second cloud delay during a contact-rich grasp is a different failure mode than a multi-second delay while crossing a hallway.
- Bimanual coordination would need synchronization across two edge adapters or a single adapter with a wider observation footprint.
The most likely manipulation follow-up paper: "AsyncVLA-Manip" — π0.5/π0.6 in the cloud, a 100 M edge adapter on the robot, evaluated under datacenter-to-robot WiFi latency on tabletop tasks. As of the AsyncVLA preprint date this paper does not exist publicly. It is the obvious next step.
If you are building a real-world VLA deployment and cannot ship a workstation per robot:
- Pick a base VLA you trust (NaVILA / OmniVLA / π0.5 / GR00T-N1.7) and run it remotely.
- Train a small (50–100 M) edge adapter that consumes the same delayed observation the cloud sees + the current observation, plus the cloud's action-token embedding. Do not just give it the current frame — that breaks the interpretation of stale guidance.
- Two-stage fine-tune: edge first, then joint. Joint stage is mandatory if you care about robustness to delay variance.
- Re-weight your training data toward chunks with high motion, contact events, or dynamic-obstacle interactions — the cases where the edge correction matters most.
- Test under realistic latency, not lab WiFi. AsyncVLA validates 0.28–6 s; production WiFi/4G/Starlink can be worse.
- Keep the safety-critical reflex layer fully on-edge (collision avoidance, e-stop). The AsyncVLA Edge Adapter is not a safety layer — that's still PD + a hard collision check.
The paper is honest about its scope:
- Navigation only. No manipulation, no contact-rich tasks.
- Vizbot only. A single ground-robot platform; no quadruped, no humanoid, no aerial.
- Dataset is Levine-lab navigation stack (GNM/LeLaN/SACSoN). Generalization to non-Berkeley datasets is an open question.
- 5 Hz cloud + 8 Hz edge — the choice is informed by Vizbot's dynamics. Manipulation would need higher rates.
- PD controller is the safety layer. No formal safety analysis under cloud-edge delay; collision-free operation is empirical, not certified.
- Network model is WiFi-only. No cellular, no Starlink, no mesh; jitter and packet-loss patterns differ across these.
- Cloud-vs-workstation ablation lacks mechanistic detail. The "ours (workstation)" variant (30% SR) shows co-location underperforms the async split, but why the asynchronous high-frequency refinement helps so much — beyond reactive latency hiding — is not unpacked in detail.
What the field still needs:
- A manipulation port at π0.5/π0.6 scale — see §8.
- Multi-robot fleet experiments where one cloud VLA serves N edge adapters concurrently.
- Heterogeneous-edge experiments (different hardware tiers consuming the same cloud model).
- Safety-certified version with formal guarantees under bounded latency.
- A negative-result companion: tasks where the AsyncVLA pattern fails — likely contact-rich high-frequency manipulation.
- Per-paper page: AsyncVLA
- RTC — intra-model latency complement
- Fast-in-Slow — embedded ms-scale dual-system
- π0.7 — async subgoal-image refresh
- Steerable Policies — richer S2→S1 vocabulary
- VLA Architectures — dual-system category
- System 0/1/2 Review — hierarchical-control taxonomy
- Knowledge Insulation — analogous gradient-routing thesis
- Cross-Embodiment Review — adjacent on multi-platform deployment