RSS 2026 Semantic Contact Fields for Category Level - Heungwoo/research GitHub Wiki
Semantic Contact Fields for Category-Level Generalizable Tool Manipulation
Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Manipulation 1 · paper #4 Authors: Kevin Yuchen Ma, Heng Zhang, Weisi Lin, Mike Zheng Shou, Yan Wu arXiv: 2602.13833 · program page
Summary compiled from the arXiv paper (v2, titled "Semantic-Contact Fields for Category-Level Generalizable Tactile Tool Manipulation"); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Three-stage overview: (1) multimodal inputs — RGB-D point cloud plus GelSight tactile force arrays; (2) SCFields generation, fusing a dense semantic field (heatmap over the tool) with a contact field of contact-probability and force vectors concentrated at the peeler blade; (3) a diffusion policy executes and generalizes the invariant physical interaction from a train tool to unseen test peelers.
Problem
Contact-rich tool use needs both semantic planning (where to hold/apply a tool) and precise force regulation, but vision-centric generalist policies are physically naive while tactile policies are instance-specific. Diverse real tactile data is prohibitive to collect at scale, and zero-shot sim-to-real transfer is blocked by the nonlinear deformation of soft tactile sensors.
Method
Semantic-Contact Fields (SCFields) are a unified 3D representation: each point of the tool point cloud carries dense semantic features plus extrinsic-contact estimates (contact probability and force vector). A two-stage Sim-to-Real Contact Learning Pipeline first pre-trains a Tactile-as-PointCloud contact estimator (Focal Loss) on large-scale simulation of a Franka Panda with simulated GelSight data, then fine-tunes on a small real dataset pseudo-labeled via geometric heuristics and force optimization to align real sensor characteristics. SCFields then condition a 3D diffusion policy. Hardware: Franka Emika Panda, two GelSight Mini fingertip sensors, three RealSense D435 cameras.
Results
On real-world contact estimation, the aligned model reaches F1 0.534 on seen scrapers and 0.657 on crayons unseen in alignment, versus a near-total Sim-Only failure (F1 0.002). On the scraping task, SCFields achieves 79.6% success and 73.5% cleaning efficiency on unseen tools versus 35.1%/25.4% for the Vision-Only (GenDP) baseline; crayon-drawing consistency is 0.78 on unseen crayons (Vision-Only 0.60). On carrot peeling — with the contact model aligned only on scraper data — average peel length is 4.52 cm, about four times the Vision-Only baseline's 1.12 cm.
Significance
A concrete recipe for tactile sim-to-real at the representation level: pre-train contact physics in simulation, align sensor response with a little pseudo-labeled real data, and let the policy consume the resulting force-aware field rather than raw tactile signals. Related wiki threads: Review-Dexterous-Manipulation · Review-Tactile-VLA.
← Back to RSS 2026 survey · RSS-2026-Papers · Home