ICRA 2026 OmniVLA Nav - Heungwoo/research GitHub Wiki
Venue: ICRA 2026 ยท Authors: Noriaki Hirose, Catherine Glossop, Dhruv Shah, Sergey Levine โ UC Berkeley (Hirose also Toyota Motor North America; Shah also Princeton) ยท arXiv: 2509.19480 Category: VLA for navigation (omni-modal goal conditioning) Trend tag: Multi-modal goal specification
flowchart LR
POSE[2D goal pose] --> FUSE
IMG[Egocentric goal image] --> FUSE
LANG[Natural-language instruction] --> FUSE
FUSE[Randomized modality fusion<br/>any subset / combination] --> VLA[High-capacity VLA backbone<br/>OpenVLA-based, ~8.26B]
OBS[Current egocentric RGB] --> VLA
VLA --> ACT[Navigation actions<br/>waypoint / velocity trajectory]
Humans flexibly compose goal specifications โ a coordinate, a photo of the destination, or a spoken instruction โ when navigating. Most learned navigation policies are trained on a single goal modality, which limits adaptability: a pose-only policy cannot follow language, a language-only policy cannot exploit a precise coordinate, and each modality is bottlenecked by whatever single-modality data exists. OmniVLA targets a single backbone that accepts any of these goal modalities (or combinations), unlocking a much larger pool of heterogeneous navigation datasets.
- Backbone: a high-capacity VLA built on OpenVLA (~8.26B params per the AsyncVLA system spec), conditioned on the current egocentric RGB observation plus a goal.
- Goal modalities: three primary forms โ 2D goal poses, egocentric goal images, and natural-language instructions โ plus their combinations.
- Randomized modality fusion: during training, goal modalities are randomly dropped/combined per sample, forcing one shared policy to handle any available modality and to build richer geometric, semantic, and visual representations. This also lets datasets with only one goal type all contribute.
- Training data: ~9,500 hours across 10 platforms, including GO Stanford, SACSoN/HuRoN, RECON, SCAND, CoryHall, TartanDrive, LeLaN, FrodoBots-2k (synthetic actions via MBRA), and BDD-V (driving data reannotated for robots).
- Foundation-model behaviors: efficient fine-tuning to a new goal modality (satellite imagery), adaptation to new environments/language domains (e.g. CAST) with limited data, and zero-shot cross-embodiment transfer to Unitree Go1 and VizBot.
- Headline: OmniVLA "outperforms specialist baselines across modalities" while a single model handles pose, image, and language goals โ and shows strong generalization to unseen environments plus robustness when a modality is scarce or missing.
- Follows novel natural-language instructions zero-shot, and serves as a flexible base for fine-tuning to new modalities/tasks. Checkpoints and training code are released. (Per-task success-rate tables are in the paper/project page; no numbers are reproduced here unverified.)
OmniVLA reframes navigation goal specification as omni-modal rather than single-modal, making one backbone usable across the fragmented landscape of navigation datasets and goal interfaces. It is the cloud/base model in AsyncVLA review โ AsyncVLA runs this same ~8.26B OmniVLA on a remote workstation and pairs it with a 76M onboard Edge Adapter for low-latency reactive control. Its zero-shot transfer to Go1 and VizBot makes it a strong data point for Cross-Embodiment review.
Name collision: a different ICRA 2026 paper, "OmniVLA: Physically-Grounded Multimodal VLA with Unified Multi-Sensor Perception for Robotic Manipulation" (2511.01210, Princeton/UCLA/MSRA), shares the name but addresses manipulation with fused IR/radar/audio sensing โ unrelated to this navigation model.
- arXiv: 2509.19480
- Project page: omnivla-nav.github.io
โ Back to ICRA-2026