Review Tactile VLA - Heungwoo/research GitHub Wiki

In-Depth Review — Tactile VLA (touch- & force-grounded manipulation policies)

Compiled June 2026 · Focus: how touch and force are integrated into learned manipulation policies and VLAs — the model architecture (how tactile enters the policy), the sensor hardware, the 2026 venue trends (ICRA · ICLR · CVPR), and the open limitations.

This is the touch-centric companion to Dexterous Manipulation (§G tactile cluster), VLA Architectures §Category J (sensory-augmented input), and World Models (OmniVTA). For the full ICRA force/tactile cohort see ICRA 2026 Tactile/Force (92 papers).


1. Why tactile for VLAs

Vision tells you where things are; touch tells you what is happening at the contact — force, slip, texture, and the moment of make/break contact. For insertion, tool use, fragile grasping, and in-hand manipulation, the decisive signal is at the fingertip and changes in milliseconds. VLAs inherited a vision-language prior with no tactile channel (π-series, GR00T, OpenVLA are touch-blind), so 2025–2026 produced a wave of work bolting touch onto policies. The central design questions are: (a) how does tactile enter the model, (b) what sensor produces it, (c) is the sensor needed at deployment, and — newly in mid-2026 — (d) does it transfer across sensors/embodiments, or is it welded to one rig? Question (d) is what the first tactile foundation policy, FTP-1, puts on the table.


2. Architecture taxonomy — how tactile enters the policy

flowchart TB
  Q{How is touch integrated?}
  Q --> A["A. Native tactile token<br/>fused into VLM/policy"]
  Q --> B["B. Modality expert + router"]
  Q --> C["C. Tactile world model<br/>+ reflexive control"]
  Q --> D["D. Distilled / sensor-free<br/>at deploy"]
  Q --> E["E. Frequency-aware<br/>F/T fusion"]
  Q --> F["F. Transient-contact /<br/>minimal sensing"]
  Q --> G["G. Cross-sensor foundation policy<br/>pretrain once, fine-tune per sensor"]
Loading
# Mechanism How it works Representative work Pros / Cons
A Native tactile encoder → token, fused into the VLM/policy A (often pretrained/SSL) tactile encoder turns the touch image/signal into tokens concatenated or cross-attended alongside vision-language Tactile-VLA, VLA-Touch (tactile↔language), TaF-VLA (tactile + 6-axis F/T), OmniVTLA (semantic-aligned tactile ViT, 2508.08706), AnyTouch 2 (general optical-tactile representation, 2602.09617), DexMove (R-Tac markers, ICLR 2026) simplest; reuses VLA stack; language↔touch grounding · vision can drown out sparse touch; sensor-specific encoders
B Per-modality expert + learned router Separate diffusion/policy experts per modality; a router computes consensus weights so sparse-but-critical touch keeps influence Multi-Modal Policy Consensus (2509.23468) touch isn't averaged away; add modalities w/o full retrain · more components; routing must be learned well
C Tactile world model + reflexive control Predict short-horizon contact evolution, then servo predicted-vs-observed touch at sensor rate OmniVTA (2603.19201; 60 Hz reflex loop) true reflex layer; perturbation recovery · sensor in-loop at deploy; not language-conditioned
D Distilled / sensor-free at deployment Train with touch, then distill it into a token or a vision-only policy so no tactile sensor is needed at inference FD-VLA (force token distilled into VLM, 2602.02142), HapticVLA (tactile→vision-only, ~87% real) no sensor cost/durability at deploy; easy fleet rollout · no true reactive correction; bounded by what distillation captured
E Frequency-aware vision + F/T fusion Embed asynchronous high-rate F/T and low-rate RGB by frequency/modality, fuse via cross-attention in a diffusion policy ManipForce (2509.19047, Frequency-Aware Multimodal Transformer, 83% on 6 contact tasks), TaF-VLA respects the rate mismatch; strong on contact-rich assembly · needs F/T hardware + synchronized logging
F Transient-contact / minimal sensing A cheap sensor captures transient contact dynamics rather than steady-state force TranTac (2509.16550, single 6-axis IMU in elastomer tip; 79% avg, ~70% unseen) very low cost; transient signal often decisive · no rich spatial tactile map; limited force magnitude
G Cross-sensor foundation policy Heterogeneous per-sensor encoders project image-/array-/state-based touch into a shared morphology-aware token space modeled by one pretrained tactile expert; fine-tune by training only the sensor-specific encoder FTP-1 (2606.13102, MTTS + multi-expert; ~3,000 h / 26 sources / 21 sensors; +31% on unseen sensors) one checkpoint transfers across sensors & embodiments; the "foundation model" recipe for touch; releasable shared starting point · still a single-sensor-tuned policy after fine-tune; thin unseen-set eval; real-time rate unproven

The field's central fault lines: the original split is sensor-free at deploy (D — FD-VLA, HapticVLA, distill touch away) vs sensor-in-the-loop (C/A — OmniVTA, native-token VLAs, servo on touch): sensor-free wins on cost/deployability, sensor-in-loop on reactive precision. Mid-2026 adds an orthogonal third axis — cross-sensor generality (G): is the policy welded to one sensor, or is the sensor interchangeable? The one-liner: FD-VLA removes the sensor, OmniVTA servos on the sensor, FTP-1 makes the sensor fungible. G is orthogonal because a foundation policy can itself be deployed sensor-in-loop or distilled sensor-free.


3. Tactile sensor hardware

The 2026 message (ICRA especially): vision-based tactile sensors have splintered from a single GelSight clone into a hardware design space.

Family Examples (2024–2026) Signal Trade-off
Vision-based tactile (VBTS), elastomer+camera GelSight / GelSight Mini (OmniVTA), DIGIT, Tac3D, 9DTact, Xense, Daimon high-res deformation image rich spatial map; bulky, finite life, optics latency
VBTS variants / novel optics UVDtact / TransTac (UV-encoded elastomers), MoiréTac (Moiré amplification), marker-based R-Tac (DexMove, 120 FPS), curved/biomimetic skins amplified / encoded deformation better sensitivity or geometry; sensor-specific pipelines
Non-optical transduction magnetic / u-skin, EIT (electrical impedance tomography), acoustic, barometric, capacitive + PVDF spiking (SpikeATac, 2510.27048) large-area / low-cost / high-rate cheap & robust; lower spatial resolution
Force / torque 6-axis F/T wrist sensors, fin-ray air-channel (FORTE, 2506.18960; 0–8 N, ~0.2 N err, slip <100 ms), IMU-in-tip (TranTac) net wrench / slip / transient direct force; no contact-shape map
Soft / compliant & special liquid-metal skins, suction-fingertip (Suction Leap-Hand), 3D-printed soft sensors adhesion / compliance-preserving changes the contact problem itself

Tactile simulation & data (the enabling layer): 2026's clearest methodological shift is differentiable, physically-calibrated optical-tactile sim — DOT-Sim (2604.27367, Material Point Method + residual-image rendering, calibrates in minutes), differentiable path-tracing Sim2Real, ETac, TacFlex, ConTact (contrastive sim-to-real). On the data side, FreeTacMan (2506.01941) collects visuo-tactile demos via a robot-free wearable gripper (≈+50% over vision-only). The dominant practical lever remains data + sim, not the policy architecture.


4. Per-venue research trends (2026)

ICRA 2026 — the hardware/deployment center (~92 force-tactile papers)

The systems venue owns tactile hardware and deployment. Three shifts: (1) VBTS splintered into a design space (UV-encoded, Moiré, magnetic/EIT/acoustic/barometric, curved skins); (2) sim-to-real went differentiable + physically-calibrated (DOT-Sim, path-tracing, ETac, TacFlex, ConTact); (3) touch is being absorbed into learned policies/VLAs — including the headline force-without-a-sensor distillation (FD-VLA), hand-frame tactile anchoring for sub-mm mating (SaTA, 2510.14647, +30%), frequency-aware F/T fusion (ManipForce), transient-contact IMU sensing (TranTac), per-modality routing (Multi-Modal Consensus), and visuo-tactile world modeling (OmniVTA). Also: Symmetry-Aware Vision-Tactile (2602.13689). Trend: contact meets the robot — hardware, sim, and policy integration at volume.

ICLR 2026 — tactile representation learning

The ML venue's tactile thread is representation-centric: learn a general, transferable tactile representation that survives sensor heterogeneity and captures dynamics. AnyTouch 2 (2602.09617) is the flagship — a general optical-tactile representation unifying object-level understanding with fine-grained, force-aware dynamic perception (ToucHD hierarchical dataset). DexMove brings a native-tactile (R-Tac marker) VLA with optical-flow contact features. Tactile also rides the broader sensory-augmented / unified-diffusion architectures here (it is the "J" modality in Review-VLA-Architecture). Trend: a tactile foundation-representation, sensor-agnostic and dynamics-aware — the missing "DINO for touch."

CVPR 2026 — perception-grounded touch & "act without the sensor"

The vision venue frames touch through perception and deployment economics. HapticVLA (Skoltech) is the emblem: train an action model with safety-aware tactile rewards, then distill tactile knowledge into a vision-only VLA so it acts without a sensor at deploy (~87% real). The tactile/haptic/bimanual bucket sits alongside CVPR's egocentric and reasoning clusters. Trend: distill touch into vision; treat tactile as a training-time teacher, not a deploy-time requirement — the mirror image of ICRA's sensor-in-the-loop work.

Cross-venue arc: ICRA = hardware + sim + sensor-in-loop policies; ICLR = transferable tactile representation; CVPR = distill touch into vision, sensor-free deploy. The three are complementary: ICLR's representations feed ICRA's policies; CVPR's distillation is the deployment escape hatch from ICRA's sensor cost.

The arXiv frontier (mid-2026): the newest move closes the loop between ICLR's representation thread and ICRA's policy thread — a cross-sensor tactile foundation policy. FTP-1 (2606.13102, Tsinghua / Shanghai Qi Zhi / Yang Gao et al.) pretrains one policy on ~3,000 h across 21 sensors and fine-tunes per sensor by swapping only the encoder — the "OpenVLA/π0 moment" for touch. Not yet at a venue, but it is the first attempt to make a tactile policy a reusable checkpoint rather than a one-off.


5. Limitations & open problems

  1. Sensor heterogeneity — the universal tactile foundation is now being attempted, not yet settled. Encoders are largely sensor-specific (GelSight ≠ DIGIT ≠ magnetic). AnyTouch 2 / OmniVTLA push toward cross-sensor representations; FTP-1 (Jun 2026) is the first cross-sensor foundation policy, claiming +31% on unseen sensors by mapping image-/array-/state-based touch into a shared token space (MTTS). But it is a proof-of-concept — unseen-sensor eval is 3 tasks / 2 setups, and the gain partly reflects "tactile-aware vs. not" rather than purely "pretrained vs. from-scratch tactile." A robustly validated touch foundation model transferable across VBTS and force-based sensors is in progress, not established.
  2. Sim-to-real gap. Tactile has a notoriously large sim-to-real gap (soft-body contact, optics). 2026's differentiable/calibrated sims (DOT-Sim, ETac) help but are mostly per-sensor and short-horizon.
  3. Vision drowns out touch. Naive feature concatenation lets the high-dimensional vision stream dominate; routing (Multi-Modal Consensus) and frame-anchoring (SaTA) are partial fixes, not settled practice.
  4. Sensor cost / durability / deploy. Fingertip sensors wear, add latency, and complicate fleets — the economic pressure behind the sensor-free distillation camp (FD-VLA, HapticVLA), which in turn cannot do true reactive correction.
  5. Evaluation. Tactile work is largely real-robot-only with bespoke task suites — no standard tactile-VLA benchmark (LIBERO/SIMPLER are touch-blind), so cross-paper comparison is weak.
  6. Language↔touch grounding is thin. Few systems connect tactile to language/instructions (VLA-Touch, OmniVTLA are early); "grasp it gently" / "until you feel a click" remains largely unsolved.
  7. Reflex vs chunk mismatch. Chunk-level diffusion/VLA policies are too slow for contact reflexes; OmniVTA's 60 Hz loop is one answer, but reflexive correction is "largely missing from current visuo-tactile policies."

6. Comparison table

Paper Venue Tactile enters via Sensor HW Note / number
FTP-1 arXiv 2026 G — cross-sensor foundation policy 21 sensors (image/array/state) ~3,000 h pretrain; +31% unseen-sensor, +17% seen
OmniVTA arXiv 2026 C — world model + 60 Hz reflex GelSight Mini (VBTS) OmniViTac 21k+ traj; closed-loop ≫ open-loop
FD-VLA ICRA 2026 D — distilled force token none at deploy (F/T at train) distilled token reportedly beats real sensor
HapticVLA CVPR 2026 D — distill into vision-only none at deploy ~87% real; safety-aware tactile rewards
AnyTouch 2 ICLR 2026 A — general tactile representation optical (VBTS, multi) ToucHD dataset; dynamic/force-aware
OmniVTLA 2025 A — semantic-aligned tactile ViT vision- & force-based (dual-path) ObjTac dataset (56 objects)
SaTA ICRA 2026 A — hand-frame-anchored tactile VBTS +30% sub-mm USB-C mating
ManipForce ICRA 2026 E — freq-aware F/T fusion 6-axis F/T + RGB 83% on 6 contact tasks
TranTac ICRA 2026 F — transient contact IMU-in-tip 79% avg; beats vision-only & F/T
Multi-Modal Consensus ICRA 2026 B — modality experts + router vision + touch sparse touch retains influence
Tactile-VLA / VLA-Touch / TaF-VLA 2025 A — native tactile/F-T token VBTS / F/T the founding native-tactile-VLA cluster

7. Decision guide

  1. Need true reflexive contact correction (insertion, slip recovery)? → C: visuo-tactile world model + high-rate reflex (OmniVTA); or E if you have F/T (ManipForce).
  2. Cost-sensitive fleet, can't put sensors on every robot? → D: distill at train, sensor-free at deploy (FD-VLA, HapticVLA).
  3. Cheapest possible touch upgrade to a vision policy? → F: an IMU in the fingertip (TranTac).
  4. Vision keeps overpowering touch? → B: per-modality experts + router (Multi-Modal Consensus), or A with frame-anchoring (SaTA).
  5. Building a multi-sensor / transferable stack? → A with a general tactile representation (AnyTouch 2, OmniVTLA), or a pretrained cross-sensor foundation policy you fine-tune per sensor (FTP-1 — swap only the sensor encoder, reuse the shared tactile expert).
  6. Blocked on data / sim? → FreeTacMan (wearable data) + a differentiable tactile sim (DOT-Sim/ETac) before touching the policy.

8. Links

← Back to Home

⚠️ **GitHub.com Fallback** ⚠️