ICLR 2026 CE Nav - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 ยท arXiv: 2509.23203 ยท Category: Embodied navigation (cross-embodiment local navigation) ยท Trend tag: IL-then-RL decoupling of geometric reasoning from embodiment dynamics.
flowchart LR
Planner[Classical planner<br/>large-scale offline data] --> IL
subgraph IL["Stage 1 โ General Expert (offline IL)"]
VelFlow[VelFlow<br/>conditional normalizing flow<br/>distribution of feasible velocities]
end
VelFlow -->|frozen prior| RL
subgraph RL["Stage 2 โ Dynamics-Aware Refiner (online RL)"]
Refiner[Lightweight refiner<br/>compensates target dynamics]
end
Refiner --> Cmd[Body velocity command<br/>vx, vy, vyaw]
Cmd --> Robots[Quadruped / biped / quadrotor]
Local navigation policies are usually embodiment-specific: a policy tuned for a quadruped does not transfer to a biped or a quadrotor without costly retraining on each robot's dynamics. Imitation from a single planner path also suffers from multi-modality (many valid actions per state). CE-Nav targets low-cost transfer across embodiments without sacrificing performance.
CE-Nav is a two-stage IL-then-RL framework that decouples universal geometric reasoning from embodiment-specific dynamic adaptation:
- Stage 1 โ General Expert (offline IL): VelFlow, a conditional normalizing-flow model, learns the full distribution of kinematically sound actions from a large dataset generated by a classical planner. Using a flow resolves the multi-modality problem and avoids any real-robot data.
- Stage 2 โ Dynamics-Aware Refiner (online RL): for a new robot the expert is frozen as a guiding prior, and a lightweight refiner is trained by online RL to compensate for that robot's specific dynamics and controller imperfections with minimal interaction.
The high-level interface is a universal body-velocity action space (vx, vy, vyaw), shared across many mobile robots.
Experiments span quadrupeds, bipeds, and quadrotors, where CE-Nav reports state-of-the-art performance while drastically reducing per-embodiment adaptation cost (the frozen flow prior means only a small refiner is learned online). Exact per-robot success/cost numbers are omitted here pending the camera-ready tables.
CE-Nav shows that geometric "what to do" can be learned once from a classical planner via a normalizing-flow expert, while only the cheap, embodiment-specific "how to move" refiner needs online RL per robot. This is a scalable template for cross-embodiment deployment of local navigation policies.
- ICLR 2026 Survey
- OmniVLA (navigation)
- NavFoM (cross-embodiment navigation foundation model)
โ Back to ICLR-2026