ICLR 2026 AsyncVLA - Heungwoo/research GitHub Wiki

AsyncVLA — Asynchronous Foundation-Model + Edge Adapter for Navigation Under 6-Second Latency

Authors: Noriaki Hirose · Catherine Glossop · Dhruv Shah · Sergey Levine Affiliations: UC Berkeley · Toyota Motor North America (Hirose) · Princeton (Shah) Venue: arXiv preprint · Feb 13, 2026 · arXiv 2602.13476 Category: VLA Architecture (dual-system / hierarchical) · Navigation Trend tag: Trend 1 (system architecture) · Trend 7 (real-time deployment)

Approach diagram

flowchart LR
  subgraph Remote[Remote workstation · RTX 4090]
    OBS1["Past observation I_{t-k}"] --> BVLA
    INST[Language / 2D-pose goal] --> BVLA[OmniVLA<br/>8.26B params · 5 Hz]
    BVLA -- delayed action token<br/>embeddings via WiFi --> NET((WiFi · 0.28–6.0 s))
  end

  subgraph Edge[Robot · Jetson Orin 30W]
    NET --> EA[Edge Adapter<br/>76M params · 8 Hz]
    OBS2[Recent observation I_t] --> EA
    EA --> ACT[Refined action chunk]
    ACT --> PD[PD controller · 10 Hz]
    PD --> ROB[Vizbot ground robot]
  end

  classDef remote fill:#bbdefb,stroke:#1565c0,color:#000
  classDef edge fill:#c8e6c9,stroke:#2e7d32,color:#000
  classDef net fill:#fff9c4,stroke:#f57f17,color:#000
  class BVLA,OBS1,INST remote
  class EA,OBS2,ACT,PD,ROB edge
  class NET net
Loading

Problem

Robotic foundation models (OmniVLA, NaVILA-class) achieve strong open-world generalization by leveraging internet-scale vision-language pretraining, but their 8B-parameter inference cost creates a fundamental bottleneck for real-time deployment: a single forward pass can take hundreds of milliseconds on a workstation GPU, and multi-second WiFi round-trips push end-to-end latency to 0.28–6 seconds in realistic mobile-robot settings. With a moving robot in a dynamic environment (pedestrians, doors, other vehicles), the foundation model's "current" action chunk is already stale by the time it lands on the edge — collisions follow.

Two prior responses both fall short:

  • Run the small model only (OmniVLA-edge, 108 M params, 6 Hz onboard) — fast and reactive, but limited semantic capability; struggles with language-conditioned navigation and out-of-distribution prompts.
  • Run the big model only — full semantic capability, but at 5 Hz on workstation + WiFi delay it cannot react to dynamic obstacles.

Prior dual-system work (Fast-in-Slow, π0.5, Hi Robot) uses synchronous coupling with sub-300 ms hand-offs — none has been demonstrated under multi-second edge-to-cloud latency.

Method

Two policies, asynchronously coupled.

  1. Base VLA (OmniVLA, 8.26 B): SigLIP + DINOv2 vision encoders + LLaMA-2-7B language model. Runs on a remote RTX 4090 workstation at 5 Hz. Emits action-token embeddings conditioned on a (delayed) past observation $I_{t-k}$ and the goal (language or 2D pose).
  2. Edge Adapter (76 M params): Lightweight onboard policy on Jetson Orin (30 W). Runs at 8 Hz, consumes the most recent observation $I_t$, the delayed past observation $I_{t-k}$, and the delayed action-token embeddings from the base VLA, and emits a refined action chunk that the PD controller (10 Hz) executes.

Why both observations. The Edge Adapter receives the same $I_{t-k}$ that the base VLA used so it can interpret the (delayed) embedding correctly, then conditions its refinement on the current $I_t$ to react to anything that has changed since.

Trajectory re-weighting during training. Sequences with large intra-chunk pose changes (final-pose distance > $d_\text{th} = 1.0$ m) are up-weighted, prioritizing samples that exhibit reactive behavior (collision avoidance, yielding to pedestrians) — i.e., training emphasizes exactly the cases where async correction matters.

Two-stage end-to-end finetuning.

  • Stage 1: Train Edge Adapter (θ, φ) with the base VLA frozen.
  • Stage 2: Fine-tune the entire system (including base VLA ψ) jointly so the two policies align their representations across the WiFi delay.

Both stages use an imitation loss + an action-smoothing loss.

Results

Vizbot ground robot. Two task families. Strong dynamic-obstacle setting (pedestrians, cluttered indoor + outdoor).

AsyncVLA OmniVLA (8.26 B remote, 5 Hz) OmniVLA-edge (108 M onboard, 6 Hz)
Success rate 85% 45% 25%
Static collisions 0.10 0.30 0.60
Dynamic collisions 0.10 1.05 1.00
Time to goal (s) 59.18 70.73 80.07

→ +40 pp success vs. SOTA baselines (85% vs. 45% for remote OmniVLA, 25% for OmniVLA-edge), >10× fewer dynamic collisions (0.10 vs. ~1.0–1.05), faster goal-reaching (59.18 s vs. 70.73–80.07 s). Holds with WiFi latency injected up to 6 seconds. Note: the strong "Ours (workstation)" ablation (no edge, 89.79 s, 0.50 static / 0.67 dynamic collisions) underperforms the full async system — the edge adapter is what delivers the reactive gains.

Robust to:

  • 2D-pose-conditioned navigation (12–30 m, 10 environments).
  • Language-conditioned navigation (5–20 m, 12 environments — offices, kitchens, halls), including out-of-distribution prompts.

Significance

The first hierarchical VLA designed for seconds-scale latency, not milliseconds. Prior dual-system VLAs (Fast-in-Slow, π0.5/Hi Robot, GR00T, ChatVLA-2) all assume the System-2 ↔ System-1 hand-off is < 300 ms — i.e., both systems are co-located on the same machine or rack. AsyncVLA breaks that assumption: System 2 lives in the cloud, System 1 lives on the robot, and they communicate over WiFi.

Two structural ideas are new:

  1. The edge model conditions on both observations ($I_t$ + $I_{t-k}$), not just the recent one — that's what lets it correctly interpret stale guidance.
  2. End-to-end fine-tuning across the WiFi gap — both policies are jointly optimized despite being non-co-located at deployment, so the foundation model learns to emit guidance that is robust to being interpreted by a 76 M edge model under delay.

Compared to neighboring 2025–2026 latency work:

  • Real-Time Chunking — handles chunk-boundary latency within one model via async inpainting. AsyncVLA handles inter-system latency between two models. Complementary, not overlapping.
  • π0.7 — uses async subgoal-image refresh (every ~4 s) but the action expert and VLM are co-located. AsyncVLA is the WiFi-separated cousin.
  • Fast-in-Slow — ms-scale embedded dual-system. AsyncVLA is the s-scale physically separated dual-system.
  • Steerable Policies — adds a richer S2→S1 vocabulary; AsyncVLA decouples where S2 and S1 run.

Cross-domain note: this is a navigation paper, not manipulation. But the architecture template — large remote VLA + small edge adapter, end-to-end fine-tuned, observation-aware re-conditioning of stale embeddings — is the most plausible deployment story for any robot fleet that wants the semantic strength of an 8 B-class VLA without a workstation onboard. Direct read-across to mobile-manipulation π0.5-class systems and to humanoid stacks where the S0/S1 layer must remain on-board.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️