CoRL 2026 FTP 1 - Heungwoo/research GitHub Wiki

CoRL 2026 — FTP-1: A Generalist Foundation Tactile Policy Across Sensors

Venue: CoRL 2026 (Austin, TX, Nov 9–12). Paper: arXiv 2606.13102. Representative of: foundation model for touch — one tactile policy transferring across sensors, hands, embodiments. Companions: Tactile VLA · Dexterous-Hand Data Pyramid · CoRL 2026 survey.

FTP-1 teaser: one tactile policy pretrained across many sensors and embodiments, transferring to unseen tactile setups (figure from the authors, arXiv 2606.13102, © the authors)

1. Problem

Vision-based generalist policies scale, but tactile policies stay tied to a fixed embodiment and a fixed sensor. Tactile signals are highly heterogeneous across hardware — optical/image sensors, taxel arrays, and raw force/torque state all look different — so a policy trained on one sensor rarely transfers to another. FTP-1 asks whether a single policy can be pretrained across many tactile sensors and then transfer, including to sensors never seen in pretraining.

2. Method

FTP-1 unifies the three tactile modalities under a Morphology-aware Tactile Token Space (MTTS): signals are mapped onto 24 functional areas, each emitting one token with shared functional-area embeddings so different sensors can share representation.

  • Heterogeneous encoders feed the shared space: ViT + a T3 Transformer for image-type sensors (e.g. GelSight), a CNN for array-type sensors (e.g. Contactile), and Fourier + MLP encoding for state-type force/torque.
  • The resulting morphology-aware latent tokens are jointly modeled by a shared ~300M-parameter tactile Transformer expert, kept as a separate expert from the vision/language stream so the tactile knowledge can transfer to new sensors.
  • Pretrained on ~3,000 hours of tactile manipulation data aggregated from 26 data sources, spanning human and robot demonstrations across 21 sensors (7 image-type, 5 array-type, 9 state-type); mix is roughly 20% human, 30% dexterous-hand robot, 50% gripper robot.

3. Results

Across downstream finetuning on 5 hardware configurations, FTP-1 improves contact-rich manipulation on seen sensor setups by +17.2%, and — the headline result — transfers to two previously unseen tactile-sensor setups for a +31% success-rate gain over baseline. Reported real-world averages: 62.5% on seen sensors and 46.6% on unseen; simulation (UniVTAC) 66.66% overall.

4. Why it matters

FTP-1 is presented as the first unified foundation baseline for tactile manipulation — a shared, model-level starting point that future tactile policies can finetune from, analogous to what pretrained backbones did for vision. The cross-sensor transfer (especially to unseen sensors) suggests tactile skills, not just per-sensor calibration, can be learned once and reused.

Limitations (reviewer): The authors note FTP-1 targets general tactile perception and does not yet address tactile/force-based servoing; and despite aggregation, the pretraining scale (~3,000 h) is still small relative to vision corpora, so generality claims rest on a modest, sensor-imbalanced dataset.

5. Links

← Back to CoRL 2026 survey · Home