ICLR 2026 AnyTouch 2 - Heungwoo/research GitHub Wiki

AnyTouch 2 — general optical tactile representation for dynamic, force-aware perception

Venue: ICLR 2026 · Authors: Ruoxuan Feng, Yuxuan Zhou, Siyu Mei, Dongzhan Zhou, Pengwei Wang, Shaowei Cui, Bin Fang, Guocai Yao, Di Hu · arXiv 2602.09617 · Category: tactile sensing · Trend tag: contact-rich / dynamic tactile perception

Approach diagram

flowchart LR
  D[ToucHD dataset<br/>atomic actions · real manip · touch-force] --> P[Multi-sensor optical tactile frames]
  P --> M[AnyTouch 2 encoder]
  M --> L1[Pixel-level: masked video + frame-difference recon]
  M --> L2[Action-level: action matching + multi-modal/cross-sensor align]
  M --> L3[Force-level: predict temporal force variation]
  L1 & L2 & L3 --> R[Unified static + dynamic, force-aware representation]
  R --> T[Object properties · dynamic attributes · dexterous manipulation]
Loading

Problem

Contact-rich manipulation needs robots to perceive temporal tactile feedback, capture subtle surface deformations, and reason about object properties and force dynamics. Optical tactile sensors can supply this rich signal, but existing tactile datasets and models focus mostly on object-level attributes (e.g., material) and largely overlook fine-grained tactile temporal dynamics during physical interaction. The authors argue progress requires a systematic hierarchy of dynamic perception capabilities guiding both data collection and model design.

Method

Two coupled contributions:

  • ToucHD — a large-scale hierarchical tactile dataset spanning tactile atomic actions, real-world manipulations, and touch-force paired data, intended as a tactile dynamic data ecosystem that explicitly supports hierarchical perception capabilities.
  • AnyTouch 2 — a general tactile representation learning framework for diverse optical tactile sensors that unifies object-level understanding with fine-grained, force-aware dynamic perception. Beyond masked video reconstruction, multi-modal alignment, and cross-sensor matching, it adds multi-level modules: frame-difference reconstruction (sensitivity to subtle temporal deformations), action matching (action understanding), and temporal force-variation prediction (physical-property modeling). It captures both pixel-level and action-specific deformations across frames.

Results

Evaluated on benchmarks covering static object properties and dynamic physical attributes, plus real-world manipulation tasks spanning multiple tiers of dynamic perception — from basic object-level understanding to force-aware dexterous manipulation. The paper reports consistent and strong performance across sensors and tasks. (Specific per-benchmark numbers are not stated in the abstract and are omitted here.)

Significance

Extends the original AnyTouch line from unified static-dynamic representation toward an explicit hierarchy of dynamic, force-aware tactile perception, paired with a dataset (ToucHD) designed around that hierarchy. Targets cross-sensor generality, which is a recurring obstacle for optical tactile sensing where each sensor produces different image characteristics.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️