ICLR 2026 CE Nav - Heungwoo/research GitHub Wiki

CE-Nav โ€” flow-guided RL refinement for cross-embodiment local navigation

Venue: ICLR 2026 ยท arXiv: 2509.23203 ยท Category: Embodied navigation (cross-embodiment local navigation) ยท Trend tag: IL-then-RL decoupling of geometric reasoning from embodiment dynamics.

Approach diagram

flowchart LR
  Planner[Classical planner<br/>large-scale offline data] --> IL
  subgraph IL["Stage 1 โ€” General Expert (offline IL)"]
    VelFlow[VelFlow<br/>conditional normalizing flow<br/>distribution of feasible velocities]
  end
  VelFlow -->|frozen prior| RL
  subgraph RL["Stage 2 โ€” Dynamics-Aware Refiner (online RL)"]
    Refiner[Lightweight refiner<br/>compensates target dynamics]
  end
  Refiner --> Cmd[Body velocity command<br/>vx, vy, vyaw]
  Cmd --> Robots[Quadruped / biped / quadrotor]
Loading

Problem

Local navigation policies are usually embodiment-specific: a policy tuned for a quadruped does not transfer to a biped or a quadrotor without costly retraining on each robot's dynamics. Imitation from a single planner path also suffers from multi-modality (many valid actions per state). CE-Nav targets low-cost transfer across embodiments without sacrificing performance.

Method

CE-Nav is a two-stage IL-then-RL framework that decouples universal geometric reasoning from embodiment-specific dynamic adaptation:

  • Stage 1 โ€” General Expert (offline IL): VelFlow, a conditional normalizing-flow model, learns the full distribution of kinematically sound actions from a large dataset generated by a classical planner. Using a flow resolves the multi-modality problem and avoids any real-robot data.
  • Stage 2 โ€” Dynamics-Aware Refiner (online RL): for a new robot the expert is frozen as a guiding prior, and a lightweight refiner is trained by online RL to compensate for that robot's specific dynamics and controller imperfections with minimal interaction.

The high-level interface is a universal body-velocity action space (vx, vy, vyaw), shared across many mobile robots.

Results

Experiments span quadrupeds, bipeds, and quadrotors, where CE-Nav reports state-of-the-art performance while drastically reducing per-embodiment adaptation cost (the frozen flow prior means only a small refiner is learned online). Exact per-robot success/cost numbers are omitted here pending the camera-ready tables.

Significance

CE-Nav shows that geometric "what to do" can be learned once from a classical planner via a normalizing-flow expert, while only the cheap, embodiment-specific "how to move" refiner needs online RL per robot. This is a scalable template for cross-embodiment deployment of local navigation policies.

Links

Related pages

โ† Back to ICLR-2026

โš ๏ธ **GitHub.com Fallback** โš ๏ธ