Review FTP 1 - Heungwoo/research GitHub Wiki

FTP-1 β€” A Generalist Foundation Tactile Policy

Title: FTP-1: A Generalist Foundation Tactile Policy Across Tactile Sensors for Contact-Rich Manipulation Venue: arXiv preprint (11 Jun 2026) Β· arXiv: 2606.13102 Β· Project: ftp1-policy.github.io Authors: Chengbo Yuan*, Zicheng Zhang*, Mingjie Zhou*, Wendi Chen, Yi Wang, Zhuoyang Liu, Dantong Niu, Shuo Wang, Hui Zhang, Wenkang Zhang, Yingdong Hu, Yuanqing Gong, Wanli Xing, Chuan Wen, Cewu Lu, Kaifeng Zhang, Yang Gao† Institutions: Tsinghua University Β· Shanghai Qi Zhi Institute Β· Sharpa Β· Shanghai Jiao Tong University Β· UC Berkeley Β· ETH ZΓΌrich Β· Fudan Β· Shanghai Innovation Institute Category: Tactile foundation policy / cross-sensor generalist Trend tag: The "OpenVLA/Ο€0 moment" for touch β€” foundation pretraining across heterogeneous tactile sensors

Approach diagram

flowchart LR
  IMG["image-based tactile<br/>(7 sensors)"] --> EI["image encoder"]
  ARR["array-based tactile<br/>(5 sensors)"] --> EA["array encoder"]
  STT["state-based tactile<br/>(9 sensors)"] --> ES["state encoder"]
  EI --> MTTS
  EA --> MTTS
  ES --> MTTS["Morphology-Aware Tactile Token Space (MTTS)<br/>+ functional-area embeddings"]
  MTTS --> TE["shared tactile Transformer expert"]
  VL["vision-language expert"] --> ACT["action-generation expert"]
  TE --> ACT
  ACT --> A["contact-rich action"]
Loading

Problem

Vision-based generalist policies (OpenVLA, Ο€0.5, GR00T) scaled because RGB is a standardized modality β€” any camera produces compatible pixels. Tactile has no such luck. Signals are highly heterogeneous across hardware: an optical GelSight image, a taxel pressure array, and a 6-axis force/torque reading share no common representation, and the same signal means different things on a parallel-jaw fingertip vs. a dexterous hand vs. a human finger. The consequence is that every prior tactile policy is welded to one sensor on one embodiment, and the field has had no model-level starting point the way ImageNet/CLIP weights are for vision. FTP-1's framing β€” "the first unified foundation baseline for tactile manipulation" β€” is the real contribution claim, more than any single benchmark number.

Method

Multi-expert foundation policy. Three experts share one model: a vision-language expert (instruction + scene grounding), an action-generation expert (policy head), and a shared tactile expert β€” a Transformer that ingests tactile tokens regardless of source sensor.

Morphology-Aware Tactile Token Space (MTTS) β€” the crux. Per-sensor heterogeneous encoders project image-, array-, and state-based signals into a common morphology-aware latent token space that is semantically aligned across sensors. MTTS adds functional-area embeddings so that similar regions of an end-effector (e.g., a "fingertip pad") share tactile semantics even when the underlying hardware differs. This is what lets the shared expert generalize across morphologies rather than memorizing one sensor's statistics.

Downstream fine-tuning protocol. To adapt to a new sensor, only the sensor-specific encoder is trained while the pretrained tactile expert is reused. A new sensor therefore needs just a thin adapter into MTTS, not a new policy β€” the mechanism behind unseen-sensor transfer.

Pretraining data

  • ~3,000 hours of tactile manipulation data, aggregated from 26 data sources.
  • 21 distinct sensors: 7 image-type, 5 array-type, 9 state-type.
  • Embodiment mix: 20% human demonstrations, 30% dexterous-hand, 50% gripper.
  • Language instructions rewritten with GPT-4o for linguistic diversity.

The cross-embodiment-including-human mix mirrors the data philosophy of Yang Gao's prior generalist-policy work and is the most resource-intensive part of the paper.

Results

Setting FTP-1 Gain vs. best baseline
UniVTAC simulation (familiar sensors) 66.7% avg SR β‰ˆ +17.5%
Real-robot contact-rich (familiar sensors) 62.5% avg SR β‰ˆ +17.2%
Unseen-sensor transfer (avg of 3 tasks) 46.6% +31% (vs Ο€0.5 at 15.0%)

Unseen-sensor task breakdown: Insert Hanoi 55%, Insert USB 30%, Wipe Board 55%. Evaluation spans 5 hardware configurations; transfer is shown to 2 previously unseen tactile-sensor setups. Baselines: VITaL, UniVTAC-ACT, Ο€0.5, Tactile-VLA, and an FTP-Ο€0.5 ablation (the FTP recipe grafted onto a Ο€0.5 backbone). Qualitatively, FTP-1 shows contact-reactive behavior β€” e.g., slowing down when insertion misalignment is felt β€” that vision-only policies structurally cannot produce.

That the largest margin appears on unseen sensors is the paper's most important result: direct evidence that the foundation-pretraining hypothesis transfers to touch, not just that more in-distribution data helps.

Significance

FTP-1 introduces a third axis to the field's tactile fault line. Your existing taxonomy frames a sensor-free vs. sensor-in-the-loop split; FTP-1 is orthogonal β€” cross-sensor generalization via foundation pretraining:

  • FD-VLA / HapticVLA β€” remove the sensor (distill a force/tactile token, none needed at deploy).
  • OmniVTA β€” servo on the sensor (60 Hz reflex world-model loop).
  • FTP-1 β€” make the sensor interchangeable (one pretrained tactile expert, swap a thin encoder).

One-liner: FD-VLA removes the sensor, OmniVTA servos on the sensor, FTP-1 makes the sensor fungible. As a released, fine-tunable pretrained tactile expert, FTP-1's lasting value is likely infrastructural β€” a shared model-level checkpoint others can adapt β€” more than any single benchmark. If independent groups reproduce the unseen-sensor transfer on their hardware, this becomes the default tactile pretraining checkpoint, the strongest signal yet that the "foundation model" recipe is crossing from vision into touch.

Critical assessment

  • Margin attribution. Ο€0.5 at 15.0% is a vision-centric generalist not built to ingest arbitrary tactile, so the +31% partly measures "tactile-aware vs. not," not purely "pretrained-tactile vs. from-scratch-tactile." The FTP-Ο€0.5 ablation is the more honest control β€” read its exact numbers in the paper body to isolate the MTTS+pretraining contribution. (Ablation numbers not yet verified from full text.)
  • Thin unseen set. Unseen-sensor evaluation is 3 tasks / 2 setups; USB insertion at 30% is still low in absolute terms. Strong as a proof-of-concept, thin as a generalization claim.
  • "+31%" vs "+31.6%". Abstract states +31%; project page implies +31.6%. Minor, but flag the discrepancy.
  • Reproducibility is strong: paper, code (GitHub), pretrained weights (HuggingFace), and dataset (ModelScope) are all linked β€” important for a "foundation baseline" claim to stick.

Limitations

  • Real-time control rate unspecified β€” unclear whether FTP-1 supports the reflex-rate closed loop that OmniVTA targets.
  • Sample efficiency of the sensor-specific encoder (how many demos to onboard a new sensor) not characterized.
  • State-based sensors (9 of 21) β€” unclear whether they carry weight or dilute the tactile signal in MTTS.
  • Sim-vs-real gap on UniVTAC not separated from the real-robot numbers.
  • No formal limitations section in the abstract/page; the above are open questions pending full-text reading.

Links

Related pages

← Back to Home

⚠️ **GitHub.com Fallback** ⚠️