CVPR 2026 HapticVLA - Heungwoo/research GitHub Wiki
Venue: CVPR 2026 (likely — pending CVF virtual-page confirmation) Category: Tactile / Haptic Manipulation Trend tag: Trend 1 Affiliations: Skoltech
flowchart LR
TRAIN["training time"] --> TVLA["SA-RWFM teacher<br/>SmolVLA + vision + tactile"]
REW["safety-aware tactile rewards<br/>(reward-weighted flow matching)"] --> TVLA
TVLA --> DISTILL["tactile distillation<br/>blended action targets"]
DISTILL --> VLA["vision-only student<br/>no tactile at deploy"]
DEPLOY["deployment"] --> VLA
VLA --> ACT["contact-rich action"]
Tactile sensing is invaluable for contact-rich manipulation, but tactile sensors are expensive, fragile, and not always available in deployment. Existing tactile VLAs (Tactile-VLA, VLA-Touch) require tactile inputs at inference. The question: can the tactile knowledge be learned but not required at deploy time?
Base model: SmolVLA (0.45B), a lightweight flow-matching VLA for high-frequency edge deployment.
Two-stage training:
- Stage 1 — SA-RWFM (Safety-Aware Reward-Weighted Flow Matching). Fine-tune SmolVLA's flow-matching action expert via offline RL, weighting samples by precomputed safety-aware tactile rewards. The per-step reward penalizes over-force, under-force during holding, peak pressure, pressure concentration, force asymmetry, and slip; the episode-level risk aggregates 95th-percentile force exceedance and cumulative slip. Tactile sensing is used only at training time here.
- Stage 2 — Tactile Distillation (TD). Distill the SA-RWFM teacher into a vision-only student that sees only vision + proprioceptive state. The student trains on blended action targets that interpolate ground-truth demonstrations and teacher predictions (ã = (1−α)a^GT + α·â^T, α = 0.5), transferring the tactile-informed behavior without tactile inputs at deploy.
86.7 % mean real-world success across three contact-rich tasks, with no tactile sensor at deploy time. Notably, full HapticVLA (86.7 %) outperforms the SA-RWFM variant that still uses tactile sensing at inference (75 %) — i.e., the distilled vision-only student beats its tactile-equipped counterpart. Ablation (Table I): without TD it drops to 81.7 % (async) / 75 % (sync); synchronous inference with TD is best. Baselines X-VLA (0.9B) and VLA-0 scored 0 % on these tasks.
If the distillation works as claimed, this is a major practical win: tactile-quality manipulation behavior without the tactile-sensor cost or fragility. Closest sibling: the "privileged-information distillation" pattern common in sim-to-real (VIRAL, et al.) — but here the privileged information is tactile, not state.
- arXiv: 2603.15257
- Dexterous Manipulation review (tactile cluster)
- CVPR 2026 survey
← Back to CVPR-2026