RSS 2026 ViTacFormer - Heungwoo/research GitHub Wiki

ViTacFormer — Cross-Modal Representation for Visuo-Tactile Dexterous Manipulation

Venue: RSS 2026 (Manipulation session) · Authors: Liang Heng, Haoran Geng, Kaifeng Zhang, Pieter Abbeel, Jitendra Malik (Berkeley) · arXiv: 2506.15953 · project Category: Visuo-tactile representation learning for dexterous hands Trend tag: RSS 2026 thread 4 — contact as representation Hardware: dual Realman arms + two SharpaWave anthropomorphic hands (5 digits, 17 DoF each) with high-resolution 320×240 tactile sensors on all 10 fingertips; exoskeleton-glove + VR teleoperation

Compiled from the verified RSS 2026 abstract and the paper's §III–IV / Fig. 2.

Key figure

ViTacFormer architecture (Figure 2 of arXiv 2506.15953, © the authors)

Figure 2 of the paper — the model is a conditional variational auto-encoder. Left: a transformer encoder maps the action sequence + proprioception to a style variable z (CVAE prior, ACT-style). Right: the ViTacFormer decoder consumes z, joints, multi-camera images, and touch through cross-attention, and — the key design — carries a future-touch-prediction pathway: the predicted tactile signal is fed back as an input alongside real touch while the model auto-regressively generates the action sequence. Contact anticipation is thus built into the representation the policy acts from, not appended as an observation.

Problem

Vision-based dexterous manipulation fails exactly where it matters — fine-grained control under occlusion and contact. Tactile signals are usually appended as extra observations rather than fused into a representation that anticipates contact.

Method

  • Cross-attention encoder fusing high-resolution vision and touch into a shared latent.
  • Autoregressive tactile-prediction head: the representation is trained to forecast future contact signals, not just encode current ones — making anticipated contact part of the state.
  • Easy-to-challenging curriculum progressively refines the visuo-tactile latent space.
  • The learned representation drives imitation learning on multi-fingered anthropomorphic hands.

Results (as reported)

  • ≈50% higher success rates than prior state-of-the-art across challenging real-world benchmarks.
  • First system (per the authors) to autonomously complete long-horizon dexterous tasks of up to 11 sequential stages, sustaining 2.5 minutes of continuous operation with an anthropomorphic hand.

Significance

The flagship of RSS 2026's contact-as-representation thread: where Contact-Grounded Policy grounds contact through prediction-to-control mapping, ViTacFormer grounds it through predictive representation — the tactile analogue of world-model forecasting. The 11-stage/2.5-minute result sets the long-horizon bar for anthropomorphic-hand autonomy. Slots into Review-Tactile-VLA's "predict-touch" family and Review-Dexterous-Manipulation's IL branch; the Berkeley (Abbeel/Malik) lineage connects it to the broader hand-scaling agenda.

← RSS 2026 survey · Home