RSS 2026 TouchGuide - Heungwoo/research GitHub Wiki
TouchGuide: Inference-Time Steering of Visuomotor Policies via Touch Guidance
Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Imitation learning 1 · paper #78 Authors: Zhemeng Zhang, Jiahua Ma, Xincheng Yang, Xin Wen, Yuzhi Zhang, Boyan Li, Yiran Qin, Jin Liu, Can Zhao, Li Kang, Haoqin Hong, Zhenfei Yin, Philip Torr, Hao Su, Ruimao Zhang, Daolin Ma arXiv: 2601.20239 · program page
Summary compiled from the arXiv paper (v6); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Figure 1: left, the TacUMI handheld gripper (Vive-tracker localization, rigid fingertips with Xense tactile sensors) giving the operator direct tactile feedback; right, TouchGuide steering — during denoising/flow-matching, sampled actions are steered by touch guidance so the final action succeeds (green, shoelace threaded) where the unguided base action fails (red).
Problem
Fine-grained contact-rich skills (shoe lacing, handing over fragile chips) need tight visuo-tactile synergy, but existing fusion strategies fall short: feature-level concatenation lets vision dominate sparse tactile features, and policy-level composition trains separate single-modality policies that miss cross-modal correlations. The authors (SJTU, Xense Robotics, Sun Yat-sen, Oxford, and others) instead fuse modalities in the low-dimensional action space at inference time.
Method
TouchGuide steers a pre-trained diffusion or flow-matching policy (DP, or the flow-matching action expert of π0.5) without retraining. Early sampling steps generate a coarse, visually plausible action from vision alone; in the later steps, a task-specific Contact Physical Model (CPM) injects touch guidance via classifier-guidance-style gradients (the paper derives the flow-matching variant with a t/(1−t) weighting). The CPM encodes tactile and visual observations with frozen DINOv2, fuses them with an N-layer transformer, encodes the noisy action with a 1D CNN+MLP, and outputs a cosine-similarity feasibility score trained contrastively (InfoNCE both directions) on limited expert demos, with noise-augmented actions so it operates on noisy samples. TacUMI is the companion UMI-style handheld collector: Vive tracker + two Lighthouse stations (~$720 base cost, 540 g), rigid fingertips for direct tactile feedback, 640×480 RGB and 200×350 tactile images at 30 Hz plus an onboard-computed 20×35×3 force field.
Results
On five contact-rich tasks (Shoe Lacing, Chip Handover, Cucumber Peeling, Vase Wiping, Lock Opening; Bi-ARX5 dual-arm and Flexiv Rizon4; 20 trials each), TouchGuide lifts DP's average from 16.3% to 36.2% and π0.5's from 35.9% to 58.0% (tactile-image variant), beating RDP (30.3%), PolicyConsensus (24.7%), SafeDiff, and Tactile-Dynamics baselines; on Chip Handover with π0.5 it reaches 60% vs 25% base. Ablations: noise pretraining of the CPM raises the three-task average from 39.17% to 62.50%; removing vision or touch from the CPM drops it to ~43%. TouchGuide stays robust under unseen objects/scenes and visual occlusion (0.950 average on occluded Cucumber Peeling). On Lock Opening, TacUMI-collected data trains better policies (30% with TouchGuide) than SLAM-based UMI (5%) or VR teleop (15%), and a user study rates TacUMI highest (100% valid rate, 9.6/10 satisfaction).
Significance
A clean demonstration that tactile feedback can be added to strong visuomotor policies (including VLAs like π0.5) purely at inference time via action-space guidance — decoupled from base-policy training. The remaining limitation is that the CPM is task-specific. Related wiki threads: Review-LBM-Cotraining · Review-Human-Video-Transfer.
← Back to RSS 2026 survey · RSS-2026-Papers · Home