ICLR 2026 Ground Slow Move Fast - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 ยท Authors: Meng Wei, Chenyang Wan, Jiaqi Peng, Xiqian Yu, Yuqiang Yang, Delin Feng, Wenzhe Cai, Chenming Zhu, Tai Wang, Jiangmiao Pang, Xihui Liu ยท arXiv: 2512.08186 ยท Category: Embodied navigation / VLN ยท Trend tag: Dual-system (System 1 / System 2) navigation foundation models.
flowchart LR
Obs[Egocentric RGB stream] --> S2
Inst[Language instruction] --> S2
subgraph S2["System 2 โ Global Planner (slow, ~2 Hz)"]
VLM[7B pretrained VLM<br/>image-grounded reasoning] --> Goal[Mid-term pixel-goal waypoint<br/>+ latent features]
end
Goal --> S1
Obs --> S1
subgraph S1["System 1 โ Local Policy (fast, ~30 Hz)"]
DiT[Lightweight multi-modal<br/>conditioning Diffusion Transformer] --> Traj[Smooth low-level trajectory]
end
Traj --> Robot[Robot / continuous control]
Vision-and-Language Navigation (VLN) systems face a tension: VLM-based planners reason well about language-grounded goals but are too slow and coarse for closed-loop control, while end-to-end policies move fast but generalize poorly to unseen instructions and environments. The paper argues prior work forces a single model to do both, sacrificing either reasoning generalization or real-time reactivity.
DualVLN is presented as the first dual-system VLN foundation model that decouples deliberation from execution:
- System 2 (Ground Slow): a VLM-based global planner that performs image-grounded reasoning to predict mid-term waypoint goals, expressed as explicit pixel goals plus latent features. Runs at the slower deliberative rate.
- System 1 (Move Fast): a lightweight, multi-modal conditioning Diffusion Transformer policy that consumes both the explicit pixel goal and the latent features from System 2 to generate smooth, accurate low-level trajectories at high frequency.
The decoupled training scheme preserves the VLM's pretrained generalization while letting the local policy specialize in reactive control. Conditioning System 1 on both a discrete pixel goal and continuous latent features is the key interface design that lets the fast policy stay aligned with the slow planner's intent.
The paper reports that DualVLN outperforms prior methods across the evaluated VLN benchmarks, and that real-world experiments demonstrate robust long-horizon planning together with real-time adaptability in dynamic environments. (Specific per-benchmark SR / SPL / NE figures are omitted here pending confirmation from the camera-ready tables.)
DualVLN brings the System 1 / System 2 split โ now common in manipulation VLAs โ into VLN as a foundation-model recipe. The pixel-goal-plus-latent interface between a slow VLM planner and a fast diffusion policy is a transferable pattern for any embodied agent that must combine language-grounded deliberation with high-rate closed-loop control.
- arXiv: https://arxiv.org/abs/2512.08186
- OpenReview: https://openreview.net/forum?id=GK4rznYwhn
- Hugging Face: https://huggingface.co/papers/2512.08186
- ICLR 2026 Survey
- OmniVLA (navigation)
- NavFoM (navigation foundation model)
โ Back to ICLR-2026