ICLR 2026 Coarse To Fine Keypoints - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Authors: Jianshu Hu, Lidi Wang, Shujia Li, Yunpeng Jiang, Xiao Li, Paul Weng, Yutong Ban arXiv: 2509.23575 Category: Policy learning — 3D / language-grounded manipulation Trend tag: Coarse-to-fine policy / VLM 3D keypoint grounding
flowchart LR
Inst[Language instruction] --> Dec[Task decomposition]
RGBD[RGB-D / point cloud] --> Rep[3D-aware representation]
Dec --> VLM[VLM fine-tuned for<br/>3D keypoint prediction]
Rep --> VLM
VLM --> KP[Language-aligned 3D keypoints<br/>= region of interest]
KP --> Coarse[Coarse branch:<br/>predict region of interest]
Coarse --> Fine[Fine branch:<br/>fine-grained action predictor]
Rep --> Fine
Fine --> Act[Manipulation actions]
Generalizing learned manipulation to novel instructions and environments is hard: methods that regress actions directly are data-hungry and brittle, while purely 2D grounding loses the geometry needed for precise control. The paper asks how to ground language into 3D structure so a policy can transfer with few demonstrations.
CLAP (Coarse-to-fine Language-Aligned manipulation Policy) integrates three components:
- Task decomposition of the language instruction into subgoals.
- VLM fine-tuning for 3D keypoint prediction — the VLM emits language-aligned 3D keypoints that define a region of interest.
- 3D-aware representation feeding a hierarchical policy.
A coarse branch predicts the region of interest (from the keypoints), which guides a fine-grained action predictor for precise 3D manipulation.
- On the GemBench benchmark, CLAP achieves a 12% higher average success rate than the SOTA method while using only 1/5 of the training trajectories.
- In real-world experiments, a policy trained on only 10 demonstrations generalizes to novel instructions and environments.
CLAP shows that aligning language to 3D keypoints as an intermediate, geometry-grounded representation yields strong sample efficiency and generalization, combining the interpretability of keypoint targets with a coarse-to-fine action hierarchy. It sits alongside other keypoint- and 3D-grounded manipulation work (e.g. ReKep) but adds VLM fine-tuning for keypoint prediction within an end-to-end policy.
- arXiv: https://arxiv.org/abs/2509.23575
- OpenReview / PDF: https://openreview.net/pdf/c917563473da6d5f8455d72ba42222b9722824de.pdf
← Back to ICLR-2026