ICML 2026 N2M - Heungwoo/research GitHub Wiki

N2M — Bridging navigation and manipulation by learning pose preference from rollout

Venue: ICML 2026 (Poster) Category: VLA Architecture / RL for VLA Affiliations: Kaixin Chai, Hyunjun Lee, Joseph J. Lim — KAIST; CUHK; Seoul National University Traction (2026-06): 1 citation (arXiv)

System overview: the transition from the navigation end pose to a manipulation initial pose preferred by the policy (Figure 1 from Chai et al., 2026)

Problem

In mobile manipulation, the manipulation policy has strong preferences about the initial pose from which it is executed — the same task can succeed or fail depending on where the base is positioned. But the navigation module only aims to reach the task area; it does not consider which initial pose is preferable for downstream manipulation. This misalignment between "arriving" and "arriving in a good pose" causes poor task success even when navigation and manipulation each work in isolation. The paper targets a lightweight transition module that, after navigation reaches the area, repositions the robot to a policy-preferred initial pose using only onboard observations — without pre-built scene reconstruction or global/historical information.

Method

N2M is a transition module run between navigation and manipulation in four steps: (1) at the navigation end pose, capture an RGB point cloud from the onboard RGB-D camera; (2) the N2M network predicts a distribution of preferable initial poses; (3) a collision-free pose is sampled from that distribution; (4) the robot navigates to the selected pose and executes the pre-trained manipulation policy.

The N2M network takes egocentric RGB point clouds and predicts a distribution of preferable initial poses (Figure 3 from Chai et al., 2026)

Because preferable poses are inherently multi-modal (e.g., a drawer can be closed from either side), the network models the pose distribution as a Gaussian Mixture Model, predicting the weights, means, and covariances {αₖ, μₖ, Σₖ} of K kernels. The egocentric RGB point cloud is encoded by Point-BERT into a fixed-length latent, which an MLP maps to the GMM parameters. The network is trained by minimizing a negative log-likelihood loss over observation–pose pairs, fine-tuning a pre-trained Point-BERT.

Training data comes directly from policy rollouts, not human labels. Raw data collection stitches multiple RGB point-cloud frames (via odometry or point-cloud registration) to reconstruct a local scene Sᵢ; a candidate pose is sampled, the manipulation policy is rolled out, and if it succeeds the pose is recorded as a preferable pose p_{π,i}. A viewpoint augmentation step renders each scene from diverse viewpoints to make predictions viewpoint-robust and data-efficient. Everything is transformed into the robot's body coordinate frame.

Results

Evaluated extensively in RoboCasa simulation and the real world, across multiple tasks, policies, and robots, against a reachability-based baseline and an oracle baseline.

  • Broad applicability (4 tasks × 3 policies, 300 trials each): N2M consistently beats the reachability baseline and is comparable to — sometimes better than — the oracle. In the headline PnPCounterToCab task, average success rises from 3% (reachability baseline) to 54%. On PnPCounterToCab and the DP policy, N2M even surpasses the oracle, showing that the policy's true preference need not match its training distribution.
  • Data efficiency: On PnPCounterToCab, N2M matches the oracle with only ~10 rollouts and surpasses it at ~20.
  • Generalizability: Trained on rollouts from 0–5 distinct scenes (10 rollouts each), N2M reliably estimates pose preferences in unseen scenes; performance improves as more scenes are added, under both furniture-texture and furniture-layout variation.
  • Real-world cases: Across Lamp Retrieval, Open Microwave, Use Laptop, Push Chair, and Toybox Handover, N2M predicts preferable poses from very few rollouts (3–15). The Toybox Handover task gives reliable predictions in unseen environments with only 15 samples.
  • Ablation: Removing viewpoint augmentation drops performance below the oracle, confirming its role in data efficiency and viewpoint robustness; even without it, N2M still beats the reachability baseline. Encoder visualizations show attention focused on task-relevant regions (the lamp; the person and toybox) even in unseen scenes.

Method diagram

flowchart LR
    A[Navigation end pose] --> B[Onboard RGB-D capture: RGB point cloud]
    B --> C[Point-BERT encoder]
    C --> D[MLP to GMM params]
    D --> E[Sample collision-free preferable pose]
    E --> F[Move to pose, run manipulation policy]

Significance

N2M isolates a frequently overlooked failure mode in mobile manipulation — base placement — and solves it with a small, rollout-trained, GMM-based module that needs only egocentric sensing and a handful of demonstrations. Its policy-aware, viewpoint-robust pose prediction transfers across tasks, policies, and hardware, offering a practical, data-efficient bridge between navigation and manipulation.

Links

← Back to ICML-2026