ICLR 2026 AsyncVLA - Heungwoo/research GitHub Wiki
Authors: Noriaki Hirose · Catherine Glossop · Dhruv Shah · Sergey Levine Affiliations: UC Berkeley · Toyota Motor North America (Hirose) · Princeton (Shah) Venue: arXiv preprint · Feb 13, 2026 · arXiv 2602.13476 Category: VLA Architecture (dual-system / hierarchical) · Navigation Trend tag: Trend 1 (system architecture) · Trend 7 (real-time deployment)
flowchart LR
subgraph Remote[Remote workstation · RTX 4090]
OBS1["Past observation I_{t-k}"] --> BVLA
INST[Language / 2D-pose goal] --> BVLA[OmniVLA<br/>8.26B params · 5 Hz]
BVLA -- delayed action token<br/>embeddings via WiFi --> NET((WiFi · 0.28–6.0 s))
end
subgraph Edge[Robot · Jetson Orin 30W]
NET --> EA[Edge Adapter<br/>76M params · 8 Hz]
OBS2[Recent observation I_t] --> EA
EA --> ACT[Refined action chunk]
ACT --> PD[PD controller · 10 Hz]
PD --> ROB[Vizbot ground robot]
end
classDef remote fill:#bbdefb,stroke:#1565c0,color:#000
classDef edge fill:#c8e6c9,stroke:#2e7d32,color:#000
classDef net fill:#fff9c4,stroke:#f57f17,color:#000
class BVLA,OBS1,INST remote
class EA,OBS2,ACT,PD,ROB edge
class NET net
Robotic foundation models (OmniVLA, NaVILA-class) achieve strong open-world generalization by leveraging internet-scale vision-language pretraining, but their 8B-parameter inference cost creates a fundamental bottleneck for real-time deployment: a single forward pass can take hundreds of milliseconds on a workstation GPU, and multi-second WiFi round-trips push end-to-end latency to 0.28–6 seconds in realistic mobile-robot settings. With a moving robot in a dynamic environment (pedestrians, doors, other vehicles), the foundation model's "current" action chunk is already stale by the time it lands on the edge — collisions follow.
Two prior responses both fall short:
- Run the small model only (OmniVLA-edge, 108 M params, 6 Hz onboard) — fast and reactive, but limited semantic capability; struggles with language-conditioned navigation and out-of-distribution prompts.
- Run the big model only — full semantic capability, but at 5 Hz on workstation + WiFi delay it cannot react to dynamic obstacles.
Prior dual-system work (Fast-in-Slow, π0.5, Hi Robot) uses synchronous coupling with sub-300 ms hand-offs — none has been demonstrated under multi-second edge-to-cloud latency.
Two policies, asynchronously coupled.
-
Base VLA (OmniVLA, 8.26 B): SigLIP + DINOv2 vision encoders + LLaMA-2-7B language model. Runs on a remote RTX 4090 workstation at 5 Hz. Emits action-token embeddings conditioned on a (delayed) past observation
$I_{t-k}$ and the goal (language or 2D pose). -
Edge Adapter (76 M params): Lightweight onboard policy on Jetson Orin (30 W). Runs at 8 Hz, consumes the most recent observation
$I_t$ , the delayed past observation$I_{t-k}$ , and the delayed action-token embeddings from the base VLA, and emits a refined action chunk that the PD controller (10 Hz) executes.
Why both observations. The Edge Adapter receives the same
Trajectory re-weighting during training. Sequences with large intra-chunk pose changes (final-pose distance >
Two-stage end-to-end finetuning.
- Stage 1: Train Edge Adapter (θ, φ) with the base VLA frozen.
- Stage 2: Fine-tune the entire system (including base VLA ψ) jointly so the two policies align their representations across the WiFi delay.
Both stages use an imitation loss + an action-smoothing loss.
Vizbot ground robot. Two task families. Strong dynamic-obstacle setting (pedestrians, cluttered indoor + outdoor).
| AsyncVLA | OmniVLA (8.26 B remote, 5 Hz) | OmniVLA-edge (108 M onboard, 6 Hz) | |
|---|---|---|---|
| Success rate | 85% | 45% | 25% |
| Static collisions | 0.10 | 0.30 | 0.60 |
| Dynamic collisions | 0.10 | 1.05 | 1.00 |
| Time to goal (s) | 59.18 | 70.73 | 80.07 |
→ +40 pp success vs. SOTA baselines (85% vs. 45% for remote OmniVLA, 25% for OmniVLA-edge), >10× fewer dynamic collisions (0.10 vs. ~1.0–1.05), faster goal-reaching (59.18 s vs. 70.73–80.07 s). Holds with WiFi latency injected up to 6 seconds. Note: the strong "Ours (workstation)" ablation (no edge, 89.79 s, 0.50 static / 0.67 dynamic collisions) underperforms the full async system — the edge adapter is what delivers the reactive gains.
Robust to:
- 2D-pose-conditioned navigation (12–30 m, 10 environments).
- Language-conditioned navigation (5–20 m, 12 environments — offices, kitchens, halls), including out-of-distribution prompts.
The first hierarchical VLA designed for seconds-scale latency, not milliseconds. Prior dual-system VLAs (Fast-in-Slow, π0.5/Hi Robot, GR00T, ChatVLA-2) all assume the System-2 ↔ System-1 hand-off is < 300 ms — i.e., both systems are co-located on the same machine or rack. AsyncVLA breaks that assumption: System 2 lives in the cloud, System 1 lives on the robot, and they communicate over WiFi.
Two structural ideas are new:
-
The edge model conditions on both observations (
$I_t$ +$I_{t-k}$ ), not just the recent one — that's what lets it correctly interpret stale guidance. - End-to-end fine-tuning across the WiFi gap — both policies are jointly optimized despite being non-co-located at deployment, so the foundation model learns to emit guidance that is robust to being interpreted by a 76 M edge model under delay.
Compared to neighboring 2025–2026 latency work:
- Real-Time Chunking — handles chunk-boundary latency within one model via async inpainting. AsyncVLA handles inter-system latency between two models. Complementary, not overlapping.
- π0.7 — uses async subgoal-image refresh (every ~4 s) but the action expert and VLM are co-located. AsyncVLA is the WiFi-separated cousin.
- Fast-in-Slow — ms-scale embedded dual-system. AsyncVLA is the s-scale physically separated dual-system.
- Steerable Policies — adds a richer S2→S1 vocabulary; AsyncVLA decouples where S2 and S1 run.
Cross-domain note: this is a navigation paper, not manipulation. But the architecture template — large remote VLA + small edge adapter, end-to-end fine-tuned, observation-aware re-conditioning of stale embeddings — is the most plausible deployment story for any robot fleet that wants the semantic strength of an 8 B-class VLA without a workstation onboard. Direct read-across to mobile-manipulation π0.5-class systems and to humanoid stacks where the S0/S1 layer must remain on-board.
- arXiv: 2602.13476
- HTML: https://arxiv.org/html/2602.13476
- OmniVLA (base model, 2024): arXiv 2406.04823
- Lineage: ViNT, NoMaD, LeLaN, GNM, SACSoN (Levine lab navigation series)
- In-depth review: AsyncVLA (long-form)
- Real-Time Chunking — intra-model async inpainting
- Fast-in-Slow — embedded ms-scale dual-system
- π0.7 — async subgoal-image refresh
- Steerable Policies — richer S2→S1 vocabulary
- VLA Architectures Review — dual-system category
- System 0/1/2 Review — hierarchical-control taxonomy
- Survey: VLA & Manipulation
← Back to ICLR-2026