Review AsyncVLA - Heungwoo/research GitHub Wiki

In-Depth Review — AsyncVLA: Asynchronous VLA for Fast & Robust Edge Navigation

Paper: AsyncVLA: An Asynchronous VLA for Fast and Robust Navigation on the Edge Authors: Noriaki Hirose · Catherine Glossop · Dhruv Shah · Sergey Levine Affiliations: UC Berkeley (Hirose, Glossop, Levine) · Toyota Motor North America (Hirose) · Princeton (Shah) arXiv: 2602.13476 · Feb 13, 2026

This page sits alongside Review-System-0-1-2 (hierarchical-control taxonomy) and Review-VLA-Architecture (dual-system category) as the navigation instance of cloud-edge VLA decomposition, and complements RTC (intra-model latency) on the inter-model latency axis.


1. TL;DR

Hirose, Glossop, Shah, and Levine (Berkeley + Toyota + Princeton) build a two-policy navigation stack in which:

  • An 8.26 B-parameter OmniVLA foundation model (ICRA 2026) runs on a remote workstation at 5 Hz.
  • A 76 M-parameter Edge Adapter runs on the robot (Jetson Orin, 30 W) at 8 Hz.
  • The two communicate over WiFi with up to 6 seconds of latency.

The Edge Adapter receives stale action-token embeddings from the cloud, re-conditions them on the most recent observation plus the same delayed observation the cloud model used, and emits a refined action chunk (8 steps at a 3 Hz control rate) that the PD controller (10 Hz velocity commands) executes.

End-to-end fine-tuning teaches the cloud model to emit guidance that is robust to being interpreted by the small edge model under delay, and a trajectory-re-weighting strategy emphasizes the dynamic-interaction samples (collision avoidance, pedestrian yielding) where the async correction matters most.

Result on a Vizbot ground robot in dynamic indoor + outdoor environments: 85% success vs. 45% (cloud-only) and 25% (edge-only) — a +40 pp absolute improvement over the best single-system baseline (cloud-only, 45%), with >10× fewer dynamic collisions (0.10 vs. 1.05) and faster goal reaching (59.18 s vs. 70.73 s cloud-only / 80.07 s edge-only).

The paper's headline contribution is the first hierarchical VLA explicitly designed for seconds-scale latency — every prior dual-system VLA (Fast-in-Slow, π0.5, Hi Robot, GR00T) assumes the S2↔S1 hand-off is < 300 ms because both systems live on the same machine.


2. Why this paper matters in the 2026 landscape

By Q1 2026, "dual-system" had become a default architecture pattern. The publicly defined stacks are:

Stack S2 ↔ S1 Latency assumption
Fast-in-Slow shared transformer (last 2 of 32 blocks repurposed) sub-ms (same forward pass)
π0.5 / Hi Robot NL subtask string < 300 ms
GR00T N1.x DiT cross-attention into LLM features < 100 ms
Figure Helix-02 1 kHz neural prior < 1 ms
Sharpa CraftNet 100 Hz tactile loop < 10 ms

All of these assume co-location. AsyncVLA is the first to publicly target the regime where S2 lives in the cloud and S1 lives on the robot, separated by a flaky WiFi link. That is the realistic deployment scenario for:

  • Mobile robots (delivery, security, last-mile) that cannot carry a 4090.
  • Humanoid teleop fleets where the operator's "guidance VLA" runs in a datacenter.
  • Multi-robot fleets that share a single big VLA across many edge bodies.

If that deployment model becomes standard — and the trend lines for ICLR-2026-Human-Video-Pretraining, ICLR-2026-RoboArena, and the increasingly large LBM backbones suggest it will — AsyncVLA's recipe is the canonical reference.


3. Architecture

3.1 Overall topology

flowchart TB
  subgraph Remote["Remote workstation · NVIDIA RTX 4090"]
    direction TB
    G["Goal: 2D pose OR language"] --> BVLA
    HX["Past observation I(t−k)"] --> BVLA["OmniVLA · 8.26B<br/>SigLIP + DINOv2 + LLaMA-2-7B<br/>5 Hz"]
    BVLA --> Z["Action-token embedding z(t−k)"]
  end

  Z -- "WiFi · 0.28–6.0 s" --> NET(("stale<br/>delivery"))

  subgraph Edge["Robot · Jetson Orin 30W"]
    direction TB
    NET --> EA["Edge Adapter · 76M params · 8 Hz"]
    OBSC["Recent observation I(t)"] --> EA
    HXC["Same delayed I(t−k) the cloud used"] --> EA
    EA --> AC["Refined action chunk a(t) … a(t+H)"]
    AC --> PD["PD controller · 10 Hz"]
    PD --> ROB["Vizbot"]
  end

  classDef remote fill:#bbdefb,stroke:#1565c0,color:#000
  classDef edge fill:#c8e6c9,stroke:#2e7d32,color:#000
  classDef net fill:#fff9c4,stroke:#f57f17,color:#000
  class BVLA,Z,G,HX remote
  class EA,OBSC,HXC,AC,PD,ROB edge
  class NET net
Loading

3.2 Why both observations enter the edge model

This is the structural insight. A naive design would feed the Edge Adapter only the most recent observation $I_t$ and the cloud guidance $z_{t-k}$. That fails because the embedding $z_{t-k}$ encodes "what to do given world state at time $t-k$", and the edge model has no way to relate that to its current view.

By giving the edge model both $I_{t-k}$ (the same view the cloud model saw) and $I_t$ (the current view), the Edge Adapter can:

  • Translate the cloud's guidance from the past frame into a goal-direction in the current frame.
  • Detect what has changed in $\Delta = I_t - I_{t-k}$ and react to it (a pedestrian stepped in, a door closed).
  • Maintain semantic intent (cloud-level) while updating spatial detail (edge-level).

That separation is exactly the System-2/System-1 division of labor — but distributed over a network rather than over a transformer.

3.3 Component sizes

Component Params Hardware Frequency
OmniVLA base VLA 8.26 B Remote RTX 4090 5 Hz
Edge Adapter 76 M Onboard Jetson Orin (30 W) 8 Hz
PD controller — Onboard 10 Hz
WiFi link — — variable, 0.28–6.0 s
OmniVLA-edge baseline 108 M Onboard 6 Hz

Note that the Edge Adapter (76 M) is smaller than the OmniVLA-edge baseline (108 M) — the gain is not from a bigger edge model, but from the decomposition that lets the edge model focus on reactive correction while the cloud handles semantics.

Architecturally the Edge Adapter is a lightweight ViT with a token projector that compresses the cloud's stale action-token embedding from 4×4096 → 1024 dimensions before fusing it with the edge observations.


4. Training pipeline

4.1 Trajectory re-weighting

flowchart LR
  D["Training trajectories"] --> CLF{"chunk Δpose > d_th = 1.0 m ?"}
  CLF -- yes --> UP["Up-weight in loss<br/>(reactive sample)"]
  CLF -- no --> NORM["Normal weight"]
  UP --> L["Imitation + smoothing loss"]
  NORM --> L

  classDef d fill:#e1bee7,stroke:#6a1b9a,color:#000
  classDef step fill:#fff9c4,stroke:#f57f17,color:#000
  class D d
  class CLF,UP,NORM,L step
Loading

A chunk is flagged as "reactive" if the final-pose distance from the start of the chunk exceeds 1.0 m. These flagged samples — which contain large path corrections, dynamic-obstacle responses, or pedestrian yielding — are up-weighted in the loss. The training therefore concentrates capacity on the cases where the Edge Adapter's correction over the cloud guidance is exactly what's needed.

This is conceptually similar to the hard-example mining tradition in detection / classification, but applied at chunk level for navigation.

4.2 Two-stage end-to-end fine-tuning

Stage Trainable Frozen Goal
1 Edge Adapter (θ, φ) OmniVLA (ψ) Edge model learns to interpret stale embeddings
2 Edge Adapter + OmniVLA (ψ) nothing Cloud model learns to emit guidance robust to being consumed by the small edge model under delay

Both stages optimize the same composite loss:

$$ \mathcal{L} = \mathcal{L}_\text{imitation} + \lambda \cdot \mathcal{L}_\text{smooth} $$

where $\mathcal{L}_\text{smooth}$ penalizes large step-to-step action changes (analogous to the smoothing terms in DSRL / Streaming Flow Policy).

Why stage 2 matters. Without it, the cloud model is optimized only for "produce the correct action chunk if executed immediately" — but at deployment the chunk will be interpreted by a small edge model after up to 6 seconds of delay. Stage 2 bakes the asymmetric inference path into the cloud model's training distribution. This is a navigation-domain version of Knowledge Insulation's thesis that the deployment data path must be reflected in the training data path.

4.3 Training data

GNM (an aggregate of six public navigation datasets) + LeLaN + SACSoN (HuRoN, for pedestrian scenarios) — the Levine-lab navigation dataset stack. Indoor + outdoor, language and pose goals, real human-co-presence trajectories. Training used five H200 GPUs at an effective batch size of 180.


5. Experimental setup

5.1 Robot

Vizbot ground robot:

  • Fisheye camera (forward-facing).
  • Tracking camera (visual-inertial pose).
  • Wheel-velocity sensors.
  • Onboard compute: NVIDIA Jetson Orin in 30 W mode.

5.2 Tasks

flowchart LR
  subgraph T1[Task 1: 2D-pose-conditioned navigation]
    direction TB
    T1a[10 cluttered environments<br/>indoor + outdoor]
    T1b[12 to 30 m goals]
    T1c[Pedestrians + dynamic obstacles]
  end
  subgraph T2[Task 2: Language-conditioned navigation]
    direction TB
    T2a[12 environments<br/>offices · kitchens · halls]
    T2b[5 to 20 m goals]
    T2c[Out-of-distribution prompts included]
  end
Loading

5.3 Baselines

Baseline Notes
OmniVLA (8.26 B remote, 5 Hz) Cloud-only, the obvious "use the big model" baseline.
OmniVLA-edge (108 M onboard, 6 Hz) Edge-only, the obvious "use a fast small model" baseline.
Ours w/o E2E AsyncVLA architecture but without stage-2 joint fine-tuning — isolates the contribution of stage 2.
Ours (workstation) AsyncVLA architecture but both halves on the workstation — isolates the value of cloud-edge separation vs. just having two heads.

5.4 Latency injection

To stress-test, WiFi latency is varied from 0.28 seconds (clean lab) to 6.0 seconds (worst-case packet loss).


6. Headline numbers

6.1 Aggregate

Method Success rate Static collisions Dynamic collisions Time to goal (s)
Method SR↑ Static col.↓ Dynamic col.↓ Time to goal (s)↓
--- --- --- --- ---
OmniVLA-edge (108 M onboard, 6 Hz) 25% 0.60 1.00 80.07
OmniVLA (8.26 B remote, 5 Hz) 45% 0.30 1.05 70.73
Ours w/o E2E (8.27 B + 76 M) 25% 0.60 1.05 82.78
Ours (workstation) (8.27 B + 76 M, WS only) 30% 0.70 0.50 89.79
AsyncVLA 85% 0.10 0.10 59.18

+40 pp absolute success gain over the best single-system baseline (cloud-only, 45%). ~10× reduction in dynamic collisions (0.10 vs. 1.00–1.05) and a 3–7× reduction in static collisions. Faster goal reaching (59.18 s vs. 70.73 s cloud-only).

Note the striking ablation result: neither the no-E2E variant (25% SR) nor the workstation-only variant (30% SR) comes close to the full system (85%) — both the stage-2 joint fine-tuning and the actual cloud-edge asynchronous split are individually necessary. Removing either collapses success to baseline levels, even though those variants still have the same two-model architecture.

6.2 Robustness to WiFi latency

The success rate stays high through the full 0.28–6.0 s latency range. Both single-system baselines collapse:

  • OmniVLA: success drops as latency grows (chunks are increasingly stale).
  • OmniVLA-edge: success is flat but low (no semantic guidance to use).

AsyncVLA's success curve is flat at the 85% range — the system is genuinely robust to multi-second delay.

6.3 Ablation reads

  • Stage-2 end-to-end fine-tuning — decisive: removing it (Ours w/o E2E) collapses SR from 85% to 25%, the level of the edge-only baseline. The cloud model genuinely must be trained to emit edge-friendly embeddings.
  • Both observations in the edge model ($I_t$ + $I_{t-k}$) — required; ablating $I_{t-k}$ breaks the embedding interpretation.
  • Trajectory re-weighting — material on dynamic-obstacle subset, smaller on static.
  • Cloud-edge vs. workstation-only — Ours (workstation) (both halves co-located) reaches only 30% SR, vs. 85% for the deployed async split. The asynchronous high-frequency edge refinement is itself part of the win, not just having two heads.

7. How AsyncVLA differs from neighboring 2025–2026 work

flowchart TB
  subgraph LAT[Latency-handling axis]
    direction LR
    INTRA[Intra-model latency<br/>chunk boundary] --- INTER[Inter-system latency<br/>cloud ↔ edge]
  end

  RTC[RTC<br/>NeurIPS 2025] -.-> INTRA
  PI07[π0.7 async subgoal<br/>4s refresh] -.-> INTER
  ASYNC[AsyncVLA<br/>this paper] -.-> INTER
  FIS[Fast-in-Slow] -.-> INTRA
  HEL[Helix-02 1kHz] -.-> INTRA

  classDef paper fill:#fff9c4,stroke:#f57f17,color:#000
  classDef this fill:#c8e6c9,stroke:#2e7d32,color:#000
  class RTC,PI07,FIS,HEL paper
  class ASYNC this
Loading
Paper Latency axis S2 ↔ S1 separation Latency tolerance
Fast-in-Slow intra-model shared transformer blocks sub-ms
Helix-02 intra-stack dedicated 1 kHz prior < 1 ms
π0.5 / Hi Robot intra-machine NL subtask < 300 ms
GR00T N1.x intra-model cross-attn at layer 12 < 100 ms
π0.7 async subgoal inter-system (sort of) subgoal image every ~4 s ~4 s
RTC intra-model async inpainting at chunk boundary one chunk
AsyncVLA (this paper) inter-system (cloud ↔ edge) full model split up to 6 s

AsyncVLA is the only publicly demonstrated VLA that handles the cloud-edge case at seconds scale.

The closest cousin is π0.7's async subgoal-image refresh — but in π0.7 the action expert and the BAGEL-14B subgoal generator are still co-located; only the subgoal image is refreshed asynchronously. AsyncVLA splits the action policy itself across the network.

The closest manipulation analog would be a hypothetical "remote π0.7 + on-robot Edge Adapter for π0.7" — which doesn't yet exist publicly and would be the obvious follow-up paper for a manipulation team.


8. Why this matters for manipulation, even though the paper is navigation

The wiki is mostly manipulation, so the natural question is: does the AsyncVLA recipe transfer?

The structural arguments suggest yes, with caveats:

  1. The cloud-edge bottleneck is the same. A π0.5-class system at 8 B parameters cannot run at 50 Hz on a Jetson; either you ship a 4090 with each robot, or you accept WiFi delay.
  2. The two-observation re-conditioning trick generalizes. Manipulation has even more visual structure (gripper occlusion, bimanual coordination) — the edge model would benefit at least as much from $I_t$ + $I_{t-k}$ as the navigation model does.
  3. The trajectory-re-weighting trick has a manipulation analog: up-weight chunks where the gripper makes contact, where the trajectory has high curvature, or where bimanual coordination shifts.

The caveats:

  • Manipulation is higher-frequency than navigation (50–200 Hz typical for π0.5-class action expert, vs. 8 Hz for AsyncVLA's edge). The PD controller's 10 Hz is fine for navigation but insufficient for fine manipulation. A manipulation port would need a higher-rate edge model.
  • Manipulation has tighter precision — multi-second cloud delay during a contact-rich grasp is a different failure mode than a multi-second delay while crossing a hallway.
  • Bimanual coordination would need synchronization across two edge adapters or a single adapter with a wider observation footprint.

The most likely manipulation follow-up paper: "AsyncVLA-Manip" — π0.5/π0.6 in the cloud, a 100 M edge adapter on the robot, evaluated under datacenter-to-robot WiFi latency on tabletop tasks. As of the AsyncVLA preprint date this paper does not exist publicly. It is the obvious next step.


9. Recipe for practitioners

If you are building a real-world VLA deployment and cannot ship a workstation per robot:

  1. Pick a base VLA you trust (NaVILA / OmniVLA / π0.5 / GR00T-N1.7) and run it remotely.
  2. Train a small (50–100 M) edge adapter that consumes the same delayed observation the cloud sees + the current observation, plus the cloud's action-token embedding. Do not just give it the current frame — that breaks the interpretation of stale guidance.
  3. Two-stage fine-tune: edge first, then joint. Joint stage is mandatory if you care about robustness to delay variance.
  4. Re-weight your training data toward chunks with high motion, contact events, or dynamic-obstacle interactions — the cases where the edge correction matters most.
  5. Test under realistic latency, not lab WiFi. AsyncVLA validates 0.28–6 s; production WiFi/4G/Starlink can be worse.
  6. Keep the safety-critical reflex layer fully on-edge (collision avoidance, e-stop). The AsyncVLA Edge Adapter is not a safety layer — that's still PD + a hard collision check.

10. Limitations and open questions

The paper is honest about its scope:

  • Navigation only. No manipulation, no contact-rich tasks.
  • Vizbot only. A single ground-robot platform; no quadruped, no humanoid, no aerial.
  • Dataset is Levine-lab navigation stack (GNM/LeLaN/SACSoN). Generalization to non-Berkeley datasets is an open question.
  • 5 Hz cloud + 8 Hz edge — the choice is informed by Vizbot's dynamics. Manipulation would need higher rates.
  • PD controller is the safety layer. No formal safety analysis under cloud-edge delay; collision-free operation is empirical, not certified.
  • Network model is WiFi-only. No cellular, no Starlink, no mesh; jitter and packet-loss patterns differ across these.
  • Cloud-vs-workstation ablation lacks mechanistic detail. The "ours (workstation)" variant (30% SR) shows co-location underperforms the async split, but why the asynchronous high-frequency refinement helps so much — beyond reactive latency hiding — is not unpacked in detail.

What the field still needs:

  • A manipulation port at π0.5/π0.6 scale — see §8.
  • Multi-robot fleet experiments where one cloud VLA serves N edge adapters concurrently.
  • Heterogeneous-edge experiments (different hardware tiers consuming the same cloud model).
  • Safety-certified version with formal guarantees under bounded latency.
  • A negative-result companion: tasks where the AsyncVLA pattern fails — likely contact-rich high-frequency manipulation.

11. Pointers

← Back to Home · ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️