Review FTP 1 - Heungwoo/research GitHub Wiki
Title: FTP-1: A Generalist Foundation Tactile Policy Across Tactile Sensors for Contact-Rich Manipulation Venue: arXiv preprint (11 Jun 2026) Β· arXiv: 2606.13102 Β· Project: ftp1-policy.github.io Authors: Chengbo Yuan*, Zicheng Zhang*, Mingjie Zhou*, Wendi Chen, Yi Wang, Zhuoyang Liu, Dantong Niu, Shuo Wang, Hui Zhang, Wenkang Zhang, Yingdong Hu, Yuanqing Gong, Wanli Xing, Chuan Wen, Cewu Lu, Kaifeng Zhang, Yang Gaoβ Institutions: Tsinghua University Β· Shanghai Qi Zhi Institute Β· Sharpa Β· Shanghai Jiao Tong University Β· UC Berkeley Β· ETH ZΓΌrich Β· Fudan Β· Shanghai Innovation Institute Category: Tactile foundation policy / cross-sensor generalist Trend tag: The "OpenVLA/Ο0 moment" for touch β foundation pretraining across heterogeneous tactile sensors
flowchart LR
IMG["image-based tactile<br/>(7 sensors)"] --> EI["image encoder"]
ARR["array-based tactile<br/>(5 sensors)"] --> EA["array encoder"]
STT["state-based tactile<br/>(9 sensors)"] --> ES["state encoder"]
EI --> MTTS
EA --> MTTS
ES --> MTTS["Morphology-Aware Tactile Token Space (MTTS)<br/>+ functional-area embeddings"]
MTTS --> TE["shared tactile Transformer expert"]
VL["vision-language expert"] --> ACT["action-generation expert"]
TE --> ACT
ACT --> A["contact-rich action"]
Vision-based generalist policies (OpenVLA, Ο0.5, GR00T) scaled because RGB is a standardized modality β any camera produces compatible pixels. Tactile has no such luck. Signals are highly heterogeneous across hardware: an optical GelSight image, a taxel pressure array, and a 6-axis force/torque reading share no common representation, and the same signal means different things on a parallel-jaw fingertip vs. a dexterous hand vs. a human finger. The consequence is that every prior tactile policy is welded to one sensor on one embodiment, and the field has had no model-level starting point the way ImageNet/CLIP weights are for vision. FTP-1's framing β "the first unified foundation baseline for tactile manipulation" β is the real contribution claim, more than any single benchmark number.
Multi-expert foundation policy. Three experts share one model: a vision-language expert (instruction + scene grounding), an action-generation expert (policy head), and a shared tactile expert β a Transformer that ingests tactile tokens regardless of source sensor.
Morphology-Aware Tactile Token Space (MTTS) β the crux. Per-sensor heterogeneous encoders project image-, array-, and state-based signals into a common morphology-aware latent token space that is semantically aligned across sensors. MTTS adds functional-area embeddings so that similar regions of an end-effector (e.g., a "fingertip pad") share tactile semantics even when the underlying hardware differs. This is what lets the shared expert generalize across morphologies rather than memorizing one sensor's statistics.
Downstream fine-tuning protocol. To adapt to a new sensor, only the sensor-specific encoder is trained while the pretrained tactile expert is reused. A new sensor therefore needs just a thin adapter into MTTS, not a new policy β the mechanism behind unseen-sensor transfer.
- ~3,000 hours of tactile manipulation data, aggregated from 26 data sources.
- 21 distinct sensors: 7 image-type, 5 array-type, 9 state-type.
- Embodiment mix: 20% human demonstrations, 30% dexterous-hand, 50% gripper.
- Language instructions rewritten with GPT-4o for linguistic diversity.
The cross-embodiment-including-human mix mirrors the data philosophy of Yang Gao's prior generalist-policy work and is the most resource-intensive part of the paper.
| Setting | FTP-1 | Gain vs. best baseline |
|---|---|---|
| UniVTAC simulation (familiar sensors) | 66.7% avg SR | β +17.5% |
| Real-robot contact-rich (familiar sensors) | 62.5% avg SR | β +17.2% |
| Unseen-sensor transfer (avg of 3 tasks) | 46.6% | +31% (vs Ο0.5 at 15.0%) |
Unseen-sensor task breakdown: Insert Hanoi 55%, Insert USB 30%, Wipe Board 55%. Evaluation spans 5 hardware configurations; transfer is shown to 2 previously unseen tactile-sensor setups. Baselines: VITaL, UniVTAC-ACT, Ο0.5, Tactile-VLA, and an FTP-Ο0.5 ablation (the FTP recipe grafted onto a Ο0.5 backbone). Qualitatively, FTP-1 shows contact-reactive behavior β e.g., slowing down when insertion misalignment is felt β that vision-only policies structurally cannot produce.
That the largest margin appears on unseen sensors is the paper's most important result: direct evidence that the foundation-pretraining hypothesis transfers to touch, not just that more in-distribution data helps.
FTP-1 introduces a third axis to the field's tactile fault line. Your existing taxonomy frames a sensor-free vs. sensor-in-the-loop split; FTP-1 is orthogonal β cross-sensor generalization via foundation pretraining:
- FD-VLA / HapticVLA β remove the sensor (distill a force/tactile token, none needed at deploy).
- OmniVTA β servo on the sensor (60 Hz reflex world-model loop).
- FTP-1 β make the sensor interchangeable (one pretrained tactile expert, swap a thin encoder).
One-liner: FD-VLA removes the sensor, OmniVTA servos on the sensor, FTP-1 makes the sensor fungible. As a released, fine-tunable pretrained tactile expert, FTP-1's lasting value is likely infrastructural β a shared model-level checkpoint others can adapt β more than any single benchmark. If independent groups reproduce the unseen-sensor transfer on their hardware, this becomes the default tactile pretraining checkpoint, the strongest signal yet that the "foundation model" recipe is crossing from vision into touch.
- Margin attribution. Ο0.5 at 15.0% is a vision-centric generalist not built to ingest arbitrary tactile, so the +31% partly measures "tactile-aware vs. not," not purely "pretrained-tactile vs. from-scratch-tactile." The FTP-Ο0.5 ablation is the more honest control β read its exact numbers in the paper body to isolate the MTTS+pretraining contribution. (Ablation numbers not yet verified from full text.)
- Thin unseen set. Unseen-sensor evaluation is 3 tasks / 2 setups; USB insertion at 30% is still low in absolute terms. Strong as a proof-of-concept, thin as a generalization claim.
- "+31%" vs "+31.6%". Abstract states +31%; project page implies +31.6%. Minor, but flag the discrepancy.
- Reproducibility is strong: paper, code (GitHub), pretrained weights (HuggingFace), and dataset (ModelScope) are all linked β important for a "foundation baseline" claim to stick.
- Real-time control rate unspecified β unclear whether FTP-1 supports the reflex-rate closed loop that OmniVTA targets.
- Sample efficiency of the sensor-specific encoder (how many demos to onboard a new sensor) not characterized.
- State-based sensors (9 of 21) β unclear whether they carry weight or dilute the tactile signal in MTTS.
- Sim-vs-real gap on UniVTAC not separated from the real-robot numbers.
- No formal limitations section in the abstract/page; the above are open questions pending full-text reading.
- arXiv: 2606.13102 Β· Project page: ftp1-policy.github.io
- Tactile VLA cross-paper review (cross-sensor generalist β new third axis)
- OmniVTA (sensor-in-the-loop counterpoint)
- FD-VLA Β· HapticVLA (sensor-free counterpoints)
- Dexterous Manipulation review (Β§ tactile) Β· ICRA 2026 Tactile/Force cluster
- VLA Architectures (Category J) (sensory-augmented category)
β Back to Home