Review Tactile VLA - Heungwoo/research GitHub Wiki
Compiled June 2026 · Focus: how touch and force are integrated into learned manipulation policies and VLAs — the model architecture (how tactile enters the policy), the sensor hardware, the 2026 venue trends (ICRA · ICLR · CVPR), and the open limitations.
This is the touch-centric companion to Dexterous Manipulation (§G tactile cluster), VLA Architectures §Category J (sensory-augmented input), and World Models (OmniVTA). For the full ICRA force/tactile cohort see ICRA 2026 Tactile/Force (92 papers).
Vision tells you where things are; touch tells you what is happening at the contact — force, slip, texture, and the moment of make/break contact. For insertion, tool use, fragile grasping, and in-hand manipulation, the decisive signal is at the fingertip and changes in milliseconds. VLAs inherited a vision-language prior with no tactile channel (π-series, GR00T, OpenVLA are touch-blind), so 2025–2026 produced a wave of work bolting touch onto policies. The central design questions are: (a) how does tactile enter the model, (b) what sensor produces it, (c) is the sensor needed at deployment, and — newly in mid-2026 — (d) does it transfer across sensors/embodiments, or is it welded to one rig? Question (d) is what the first tactile foundation policy, FTP-1, puts on the table.
flowchart TB
Q{How is touch integrated?}
Q --> A["A. Native tactile token<br/>fused into VLM/policy"]
Q --> B["B. Modality expert + router"]
Q --> C["C. Tactile world model<br/>+ reflexive control"]
Q --> D["D. Distilled / sensor-free<br/>at deploy"]
Q --> E["E. Frequency-aware<br/>F/T fusion"]
Q --> F["F. Transient-contact /<br/>minimal sensing"]
Q --> G["G. Cross-sensor foundation policy<br/>pretrain once, fine-tune per sensor"]
| # | Mechanism | How it works | Representative work | Pros / Cons |
|---|---|---|---|---|
| A | Native tactile encoder → token, fused into the VLM/policy | A (often pretrained/SSL) tactile encoder turns the touch image/signal into tokens concatenated or cross-attended alongside vision-language | Tactile-VLA, VLA-Touch (tactile↔language), TaF-VLA (tactile + 6-axis F/T), OmniVTLA (semantic-aligned tactile ViT, 2508.08706), AnyTouch 2 (general optical-tactile representation, 2602.09617), DexMove (R-Tac markers, ICLR 2026) | simplest; reuses VLA stack; language↔touch grounding · vision can drown out sparse touch; sensor-specific encoders |
| B | Per-modality expert + learned router | Separate diffusion/policy experts per modality; a router computes consensus weights so sparse-but-critical touch keeps influence | Multi-Modal Policy Consensus (2509.23468) | touch isn't averaged away; add modalities w/o full retrain · more components; routing must be learned well |
| C | Tactile world model + reflexive control | Predict short-horizon contact evolution, then servo predicted-vs-observed touch at sensor rate | OmniVTA (2603.19201; 60 Hz reflex loop) | true reflex layer; perturbation recovery · sensor in-loop at deploy; not language-conditioned |
| D | Distilled / sensor-free at deployment | Train with touch, then distill it into a token or a vision-only policy so no tactile sensor is needed at inference | FD-VLA (force token distilled into VLM, 2602.02142), HapticVLA (tactile→vision-only, ~87% real) | no sensor cost/durability at deploy; easy fleet rollout · no true reactive correction; bounded by what distillation captured |
| E | Frequency-aware vision + F/T fusion | Embed asynchronous high-rate F/T and low-rate RGB by frequency/modality, fuse via cross-attention in a diffusion policy | ManipForce (2509.19047, Frequency-Aware Multimodal Transformer, 83% on 6 contact tasks), TaF-VLA | respects the rate mismatch; strong on contact-rich assembly · needs F/T hardware + synchronized logging |
| F | Transient-contact / minimal sensing | A cheap sensor captures transient contact dynamics rather than steady-state force | TranTac (2509.16550, single 6-axis IMU in elastomer tip; 79% avg, ~70% unseen) | very low cost; transient signal often decisive · no rich spatial tactile map; limited force magnitude |
| G | Cross-sensor foundation policy | Heterogeneous per-sensor encoders project image-/array-/state-based touch into a shared morphology-aware token space modeled by one pretrained tactile expert; fine-tune by training only the sensor-specific encoder | FTP-1 (2606.13102, MTTS + multi-expert; ~3,000 h / 26 sources / 21 sensors; +31% on unseen sensors) | one checkpoint transfers across sensors & embodiments; the "foundation model" recipe for touch; releasable shared starting point · still a single-sensor-tuned policy after fine-tune; thin unseen-set eval; real-time rate unproven |
The field's central fault lines: the original split is sensor-free at deploy (D — FD-VLA, HapticVLA, distill touch away) vs sensor-in-the-loop (C/A — OmniVTA, native-token VLAs, servo on touch): sensor-free wins on cost/deployability, sensor-in-loop on reactive precision. Mid-2026 adds an orthogonal third axis — cross-sensor generality (G): is the policy welded to one sensor, or is the sensor interchangeable? The one-liner: FD-VLA removes the sensor, OmniVTA servos on the sensor, FTP-1 makes the sensor fungible. G is orthogonal because a foundation policy can itself be deployed sensor-in-loop or distilled sensor-free.
The 2026 message (ICRA especially): vision-based tactile sensors have splintered from a single GelSight clone into a hardware design space.
| Family | Examples (2024–2026) | Signal | Trade-off |
|---|---|---|---|
| Vision-based tactile (VBTS), elastomer+camera | GelSight / GelSight Mini (OmniVTA), DIGIT, Tac3D, 9DTact, Xense, Daimon | high-res deformation image | rich spatial map; bulky, finite life, optics latency |
| VBTS variants / novel optics | UVDtact / TransTac (UV-encoded elastomers), MoiréTac (Moiré amplification), marker-based R-Tac (DexMove, 120 FPS), curved/biomimetic skins | amplified / encoded deformation | better sensitivity or geometry; sensor-specific pipelines |
| Non-optical transduction | magnetic / u-skin, EIT (electrical impedance tomography), acoustic, barometric, capacitive + PVDF spiking (SpikeATac, 2510.27048) | large-area / low-cost / high-rate | cheap & robust; lower spatial resolution |
| Force / torque | 6-axis F/T wrist sensors, fin-ray air-channel (FORTE, 2506.18960; 0–8 N, ~0.2 N err, slip <100 ms), IMU-in-tip (TranTac) | net wrench / slip / transient | direct force; no contact-shape map |
| Soft / compliant & special | liquid-metal skins, suction-fingertip (Suction Leap-Hand), 3D-printed soft sensors | adhesion / compliance-preserving | changes the contact problem itself |
Tactile simulation & data (the enabling layer): 2026's clearest methodological shift is differentiable, physically-calibrated optical-tactile sim — DOT-Sim (2604.27367, Material Point Method + residual-image rendering, calibrates in minutes), differentiable path-tracing Sim2Real, ETac, TacFlex, ConTact (contrastive sim-to-real). On the data side, FreeTacMan (2506.01941) collects visuo-tactile demos via a robot-free wearable gripper (≈+50% over vision-only). The dominant practical lever remains data + sim, not the policy architecture.
The systems venue owns tactile hardware and deployment. Three shifts: (1) VBTS splintered into a design space (UV-encoded, Moiré, magnetic/EIT/acoustic/barometric, curved skins); (2) sim-to-real went differentiable + physically-calibrated (DOT-Sim, path-tracing, ETac, TacFlex, ConTact); (3) touch is being absorbed into learned policies/VLAs — including the headline force-without-a-sensor distillation (FD-VLA), hand-frame tactile anchoring for sub-mm mating (SaTA, 2510.14647, +30%), frequency-aware F/T fusion (ManipForce), transient-contact IMU sensing (TranTac), per-modality routing (Multi-Modal Consensus), and visuo-tactile world modeling (OmniVTA). Also: Symmetry-Aware Vision-Tactile (2602.13689). Trend: contact meets the robot — hardware, sim, and policy integration at volume.
The ML venue's tactile thread is representation-centric: learn a general, transferable tactile representation that survives sensor heterogeneity and captures dynamics. AnyTouch 2 (2602.09617) is the flagship — a general optical-tactile representation unifying object-level understanding with fine-grained, force-aware dynamic perception (ToucHD hierarchical dataset). DexMove brings a native-tactile (R-Tac marker) VLA with optical-flow contact features. Tactile also rides the broader sensory-augmented / unified-diffusion architectures here (it is the "J" modality in Review-VLA-Architecture). Trend: a tactile foundation-representation, sensor-agnostic and dynamics-aware — the missing "DINO for touch."
The vision venue frames touch through perception and deployment economics. HapticVLA (Skoltech) is the emblem: train an action model with safety-aware tactile rewards, then distill tactile knowledge into a vision-only VLA so it acts without a sensor at deploy (~87% real). The tactile/haptic/bimanual bucket sits alongside CVPR's egocentric and reasoning clusters. Trend: distill touch into vision; treat tactile as a training-time teacher, not a deploy-time requirement — the mirror image of ICRA's sensor-in-the-loop work.
Cross-venue arc: ICRA = hardware + sim + sensor-in-loop policies; ICLR = transferable tactile representation; CVPR = distill touch into vision, sensor-free deploy. The three are complementary: ICLR's representations feed ICRA's policies; CVPR's distillation is the deployment escape hatch from ICRA's sensor cost.
The arXiv frontier (mid-2026): the newest move closes the loop between ICLR's representation thread and ICRA's policy thread — a cross-sensor tactile foundation policy. FTP-1 (2606.13102, Tsinghua / Shanghai Qi Zhi / Yang Gao et al.) pretrains one policy on ~3,000 h across 21 sensors and fine-tunes per sensor by swapping only the encoder — the "OpenVLA/π0 moment" for touch. Not yet at a venue, but it is the first attempt to make a tactile policy a reusable checkpoint rather than a one-off.
- Sensor heterogeneity — the universal tactile foundation is now being attempted, not yet settled. Encoders are largely sensor-specific (GelSight ≠ DIGIT ≠ magnetic). AnyTouch 2 / OmniVTLA push toward cross-sensor representations; FTP-1 (Jun 2026) is the first cross-sensor foundation policy, claiming +31% on unseen sensors by mapping image-/array-/state-based touch into a shared token space (MTTS). But it is a proof-of-concept — unseen-sensor eval is 3 tasks / 2 setups, and the gain partly reflects "tactile-aware vs. not" rather than purely "pretrained vs. from-scratch tactile." A robustly validated touch foundation model transferable across VBTS and force-based sensors is in progress, not established.
- Sim-to-real gap. Tactile has a notoriously large sim-to-real gap (soft-body contact, optics). 2026's differentiable/calibrated sims (DOT-Sim, ETac) help but are mostly per-sensor and short-horizon.
- Vision drowns out touch. Naive feature concatenation lets the high-dimensional vision stream dominate; routing (Multi-Modal Consensus) and frame-anchoring (SaTA) are partial fixes, not settled practice.
- Sensor cost / durability / deploy. Fingertip sensors wear, add latency, and complicate fleets — the economic pressure behind the sensor-free distillation camp (FD-VLA, HapticVLA), which in turn cannot do true reactive correction.
- Evaluation. Tactile work is largely real-robot-only with bespoke task suites — no standard tactile-VLA benchmark (LIBERO/SIMPLER are touch-blind), so cross-paper comparison is weak.
- Language↔touch grounding is thin. Few systems connect tactile to language/instructions (VLA-Touch, OmniVTLA are early); "grasp it gently" / "until you feel a click" remains largely unsolved.
- Reflex vs chunk mismatch. Chunk-level diffusion/VLA policies are too slow for contact reflexes; OmniVTA's 60 Hz loop is one answer, but reflexive correction is "largely missing from current visuo-tactile policies."
| Paper | Venue | Tactile enters via | Sensor HW | Note / number |
|---|---|---|---|---|
| FTP-1 | arXiv 2026 | G — cross-sensor foundation policy | 21 sensors (image/array/state) | ~3,000 h pretrain; +31% unseen-sensor, +17% seen |
| OmniVTA | arXiv 2026 | C — world model + 60 Hz reflex | GelSight Mini (VBTS) | OmniViTac 21k+ traj; closed-loop ≫ open-loop |
| FD-VLA | ICRA 2026 | D — distilled force token | none at deploy (F/T at train) | distilled token reportedly beats real sensor |
| HapticVLA | CVPR 2026 | D — distill into vision-only | none at deploy | ~87% real; safety-aware tactile rewards |
| AnyTouch 2 | ICLR 2026 | A — general tactile representation | optical (VBTS, multi) | ToucHD dataset; dynamic/force-aware |
| OmniVTLA | 2025 | A — semantic-aligned tactile ViT | vision- & force-based (dual-path) | ObjTac dataset (56 objects) |
| SaTA | ICRA 2026 | A — hand-frame-anchored tactile | VBTS | +30% sub-mm USB-C mating |
| ManipForce | ICRA 2026 | E — freq-aware F/T fusion | 6-axis F/T + RGB | 83% on 6 contact tasks |
| TranTac | ICRA 2026 | F — transient contact | IMU-in-tip | 79% avg; beats vision-only & F/T |
| Multi-Modal Consensus | ICRA 2026 | B — modality experts + router | vision + touch | sparse touch retains influence |
| Tactile-VLA / VLA-Touch / TaF-VLA | 2025 | A — native tactile/F-T token | VBTS / F/T | the founding native-tactile-VLA cluster |
- Need true reflexive contact correction (insertion, slip recovery)? → C: visuo-tactile world model + high-rate reflex (OmniVTA); or E if you have F/T (ManipForce).
- Cost-sensitive fleet, can't put sensors on every robot? → D: distill at train, sensor-free at deploy (FD-VLA, HapticVLA).
- Cheapest possible touch upgrade to a vision policy? → F: an IMU in the fingertip (TranTac).
- Vision keeps overpowering touch? → B: per-modality experts + router (Multi-Modal Consensus), or A with frame-anchoring (SaTA).
- Building a multi-sensor / transferable stack? → A with a general tactile representation (AnyTouch 2, OmniVTLA), or a pretrained cross-sensor foundation policy you fine-tune per sensor (FTP-1 — swap only the sensor encoder, reuse the shared tactile expert).
- Blocked on data / sim? → FreeTacMan (wearable data) + a differentiable tactile sim (DOT-Sim/ETac) before touching the policy.
- Companion reviews: Dexterous Manipulation · VLA Architectures (Cat J) · World Models · VLM↔Action
- Venue cohorts: ICRA 2026 Tactile/Force (92) · ICRA 2026 survey · CVPR 2026 survey · ICLR 2026 survey
- Per-paper: FTP-1 · OmniVTA · FD-VLA · HapticVLA · AnyTouch 2
← Back to Home