RSS 2026 Contact Anchored Policies - Heungwoo/research GitHub Wiki

Contact-Anchored Policies: Contact Conditioning Creates Strong Robot Utility Models

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Imitation learning 2 · paper #141 Authors: Zichen Jeff Cui, Omar Rayyan, Haritheja Etukuru, Bowen Tan, Zavier Andrianarivo, Zicheng Teng, Yihang Zhou, Krish Mehta, Nicholas Wojno, Kevin Yuanbo Wu, Manan H. Anjaria, Ziyuan Wu, Manrong Mao, Guangxun Zhang, Binit Shah, Yejin Kim, Soumith Chintala, Lerrel Pinto, Nur Muhammad Mahi Shafiullah arXiv: 2602.09017 · program page

Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Contact-Anchored Policies overview (Figure 1 of arXiv 2602.09017, © the authors)

Left: instead of an ambiguous language prompt, the policy is conditioned on a 3D contact anchor (x, y, z) plus vision. Middle: zero-shot roll-outs of the Pick / Open / Close utility models in unseen environments. Right: aggregate success rates — CAP first-try/retry (83%/90% Pick, 81%/91% Open, 96%/98% Close) versus π0.5-DROID (25%), AnyGrasp (47%), and Stretch-Open (58%).

Problem

Language prompts are often too abstract to specify the concrete physical interaction needed for robust manipulation (e.g., which of several similar objects to grab, and where). The NYU-led team replaces runtime language conditioning with an explicit 3D contact point and swaps the monolithic-generalist framing for a library of modular "utility models."

Method

CAP conditions a VQ-BeT policy (ResNet-50 vision backbone MoCo-pretrained on their own data; context k = 3; delta-EE-pose + gripper actions) on a contact anchor: during training, anchors are hindsight-labeled at the gripper-contact frame and back-propagated through camera odometry; at inference the anchor comes from a clicked pixel or a VLM point query (e.g., Gemini Robotics-ER 1.5) deprojected through depth, then tracked via forward kinematics. Data: 23 hours of demonstrations via an iPhone-based 3D-printable handheld gripper — Pick 14,606 demos/289 environments, Open 3,690/87, Close 2,069/48. EgoGym, a lightweight MuJoCo simulation with procedurally generated scenes, closes a real-to-sim iteration loop for failure-mode discovery across four checkpoint iterations.

Results

Zero-shot on unseen scenes/objects (Stretch 3): 83% Pick, 81% Open, 96% Close single-try; 90/91/98% with one retry. Baselines: AnyGrasp 47% (Pick), π0.5-DROID 25% (Pick, Franka), Stretch-Open 58% (Open) — CAP wins by 23-56%. The same Pick checkpoint transfers across Stretch, Franka FR3, XArm 6, and UR3e, and external collaborators (including Hello Robot) reproduced evaluations off-site. An RGB-only ablation on Close collapses from 96% to 58%, isolating the value of the contact anchor. Everything (checkpoints, code, hardware, sim, data) is slated for open-sourcing.

Significance

A pointed counterargument to language-conditioned generalist VLAs: a 3D contact point is a cheaper, less ambiguous interface that generalizes across embodiments with orders of magnitude less data — continuing the NYU "Robot Utility Models" line and relevant to Review-LBM-Cotraining and grounding debates in Review-VLA-Architecture.

← Back to RSS 2026 survey · RSS-2026-Papers · Home