ICLR 2026 Coarse To Fine Keypoints - Heungwoo/research GitHub Wiki

CLAP — coarse-to-fine manipulation via language-aligned 3D keypoints

Venue: ICLR 2026 Authors: Jianshu Hu, Lidi Wang, Shujia Li, Yunpeng Jiang, Xiao Li, Paul Weng, Yutong Ban arXiv: 2509.23575 Category: Policy learning — 3D / language-grounded manipulation Trend tag: Coarse-to-fine policy / VLM 3D keypoint grounding

Approach diagram

flowchart LR
  Inst[Language instruction] --> Dec[Task decomposition]
  RGBD[RGB-D / point cloud] --> Rep[3D-aware representation]
  Dec --> VLM[VLM fine-tuned for<br/>3D keypoint prediction]
  Rep --> VLM
  VLM --> KP[Language-aligned 3D keypoints<br/>= region of interest]
  KP --> Coarse[Coarse branch:<br/>predict region of interest]
  Coarse --> Fine[Fine branch:<br/>fine-grained action predictor]
  Rep --> Fine
  Fine --> Act[Manipulation actions]
Loading

Problem

Generalizing learned manipulation to novel instructions and environments is hard: methods that regress actions directly are data-hungry and brittle, while purely 2D grounding loses the geometry needed for precise control. The paper asks how to ground language into 3D structure so a policy can transfer with few demonstrations.

Method

CLAP (Coarse-to-fine Language-Aligned manipulation Policy) integrates three components:

  1. Task decomposition of the language instruction into subgoals.
  2. VLM fine-tuning for 3D keypoint prediction — the VLM emits language-aligned 3D keypoints that define a region of interest.
  3. 3D-aware representation feeding a hierarchical policy.

A coarse branch predicts the region of interest (from the keypoints), which guides a fine-grained action predictor for precise 3D manipulation.

Results

  • On the GemBench benchmark, CLAP achieves a 12% higher average success rate than the SOTA method while using only 1/5 of the training trajectories.
  • In real-world experiments, a policy trained on only 10 demonstrations generalizes to novel instructions and environments.

Significance

CLAP shows that aligning language to 3D keypoints as an intermediate, geometry-grounded representation yields strong sample efficiency and generalization, combining the interpretability of keypoint targets with a coarse-to-fine action hierarchy. It sits alongside other keypoint- and 3D-grounded manipulation work (e.g. ReKep) but adds VLM fine-tuning for keypoint prediction within an end-to-end policy.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️