RSS 2026 ViTacFormer - Heungwoo/research GitHub Wiki
ViTacFormer — Cross-Modal Representation for Visuo-Tactile Dexterous Manipulation
Venue: RSS 2026 (Manipulation session) · Authors: Liang Heng, Haoran Geng, Kaifeng Zhang, Pieter Abbeel, Jitendra Malik (Berkeley) · arXiv: 2506.15953 · project Category: Visuo-tactile representation learning for dexterous hands Trend tag: RSS 2026 thread 4 — contact as representation Hardware: dual Realman arms + two SharpaWave anthropomorphic hands (5 digits, 17 DoF each) with high-resolution 320×240 tactile sensors on all 10 fingertips; exoskeleton-glove + VR teleoperation
Compiled from the verified RSS 2026 abstract and the paper's §III–IV / Fig. 2.
Key figure

Figure 2 of the paper — the model is a conditional variational auto-encoder. Left: a transformer encoder maps the action sequence + proprioception to a style variable z (CVAE prior, ACT-style). Right: the ViTacFormer decoder consumes z, joints, multi-camera images, and touch through cross-attention, and — the key design — carries a future-touch-prediction pathway: the predicted tactile signal is fed back as an input alongside real touch while the model auto-regressively generates the action sequence. Contact anticipation is thus built into the representation the policy acts from, not appended as an observation.
Problem
Vision-based dexterous manipulation fails exactly where it matters — fine-grained control under occlusion and contact. Tactile signals are usually appended as extra observations rather than fused into a representation that anticipates contact.
Method
- Cross-attention encoder fusing high-resolution vision and touch into a shared latent.
- Autoregressive tactile-prediction head: the representation is trained to forecast future contact signals, not just encode current ones — making anticipated contact part of the state.
- Easy-to-challenging curriculum progressively refines the visuo-tactile latent space.
- The learned representation drives imitation learning on multi-fingered anthropomorphic hands.
Results (as reported)
- ≈50% higher success rates than prior state-of-the-art across challenging real-world benchmarks.
- First system (per the authors) to autonomously complete long-horizon dexterous tasks of up to 11 sequential stages, sustaining 2.5 minutes of continuous operation with an anthropomorphic hand.
Significance
The flagship of RSS 2026's contact-as-representation thread: where Contact-Grounded Policy grounds contact through prediction-to-control mapping, ViTacFormer grounds it through predictive representation — the tactile analogue of world-model forecasting. The 11-stage/2.5-minute result sets the long-horizon bar for anthropomorphic-hand autonomy. Slots into Review-Tactile-VLA's "predict-touch" family and Review-Dexterous-Manipulation's IL branch; the Berkeley (Abbeel/Malik) lineage connects it to the broader hand-scaling agenda.
← RSS 2026 survey · Home