RSS 2026 TactAlign - Heungwoo/research GitHub Wiki
TactAlign: Human-to-Robot Policy Transfer via Tactile Alignment
Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Manipulation 1 · paper #6 Authors: Youngsun Wi, Jessica Yin, Elvis Xiang, Akash Sharma, Jitendra Malik, Mustafa Mukadam, Nima Fazeli, Tess Hellebrekers arXiv: 2602.13579 · program page
Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Left: unpaired human (tactile-glove) and robot demonstrations feed two stages — tactile self-supervised learning, then cross-sensor tactile alignment via rectified flow that maps glove latents into the robot tactile latent space. Right: the resulting human-to-robot policies shown on real hardware — object-level generalization with <5 min of human demos, generalization beyond demonstrated object instances, dexterous manipulation from human data only, and task-level generalization.
Problem
Human demonstrations collected with wearable tactile gloves are fast and naturally dexterous, but transferring the tactile signals to a robot is hard because sensors and embodiments differ. Existing human-to-robot (H2R) approaches that use touch typically assume identical tactile sensors, need paired data, or tolerate little embodiment gap; the concurrent UniTacHand instead requires strict spatiotemporal human–robot pairing, which is impractical during sliding contact or dynamic object motion.
Method
TactAlign aligns human and robot tactile observations in a shared latent space from unpaired demonstrations of the same task. Stage 1: separate human and robot tactile encoders are pretrained with self-supervised learning (JEPA-style architecture with MSE reconstruction and cross-attention pooling to a fixed-dimensional latent), handling very different resolutions (OSMO glove 1×3 vs. Xela robot skin 30×3). Stage 2: a rectified flow learns a velocity field transporting glove latents to robot latents, trained on noisy pseudo-pairs mined from hand–object interaction similarity (matching poses and pose deltas of transitions under a threshold δ), with no paired datasets, manual labels, or privileged information; sampling uses plain Euler integration. Hardware: OSMO tactile glove on the human side; a Franka Emika Panda with Xela sensing and a RealSense D455 on the robot side.
Results
On H2R co-training across pivoting, insertion, and lid closing, TactAlign reaches 76%, 72%, and 74% success respectively — 100% on seen-by-both objects, 71% on human-only objects, and 65.5% on unseen-by-both objects (averaged over the three tasks). Ablations: removing tactile input costs −59% average success (up to −100% on pivoting); removing alignment (raw tactile features) costs −51% on average and causes near-complete failure on seen objects for pivoting/insertion. On the zero-shot dexterous light-bulb-screwing task trained from human data only, TactAlign achieves 100% success (~61 s to illumination) while both the no-tactile and no-alignment baselines score 0%.
Significance
A rare demonstration that tactile signals — not just vision or kinematics — can be transferred across the human–robot embodiment gap without paired data, using flow matching as a cross-sensor bridge. Connects directly to the wiki threads Review-Tactile-VLA and Review-Dexterous-Manipulation.
← Back to RSS 2026 survey · RSS-2026-Papers · Home