Review OmniVTA - Heungwoo/research GitHub Wiki

In-Depth Review โ€” OmniVTA: Visuo-Tactile World Modeling for Contact-Rich Manipulation

Venue: arXiv, Mar 2026 (no conference stated) ยท Authors: Yuhang Zheng, Songen Gu, Weize Li, Yupeng Zheng, Yujie Zang, Shuai Tian, Xiang Li, Ce Hao, Chen Gao, Si Liu, Haoran Li, Yilun Chen, Shuicheng Yan, Wenchao Ding ยท arXiv: 2603.19201 ยท Code: github.com/MrSecant/OmniVTA Category: Sensory-augmented (tactile) VLA โ€” Category J ร— world-model-for-actions (Review-World-Models) Trend tag: Predictive visuo-tactile world model + high-rate reflexive control

Companion reviews: Dexterous Manipulation ยงJ tactile ยท World Models ยท ICRA 2026 Tactile/Force cluster ยท VLA Architectures.

Approach diagram

flowchart LR
  V["RGB observation"] --> WM
  T["GelSight Mini<br/>visuo-tactile image"] --> ENC["Self-supervised tactile encoder<br/>(TactileVAE)"]
  ENC --> WM["Two-stream visuo-tactile<br/>world model<br/>(predict short-horizon<br/>contact evolution)"]
  WM --> POL["Contact-aware fusion policy<br/>(diffusion action chunks)"]
  POL --> RC["60 Hz Reflexive Latent<br/>Tactile Controller"]
  WM -. predicted tactile .-> RC
  T -. observed tactile .-> RC
  RC --> A["corrected actions โ†’ robot"]
Loading

Problem

Vision-only and even visuo-tactile imitation policies treat touch as just another input channel to encode, and they run open-loop within an action chunk โ€” the policy commits a chunk and only re-plans at the next inference step. For contact-rich tasks (insertion, tool use, fragile grasping) that is too slow and too blind: contact state changes in milliseconds, and a chunk planned a moment ago is quickly stale under perturbation or slip. The paper's framing is that reflexive correction โ€” closing the loop on predicted vs. observed contact at high rate โ€” is largely missing from current visuo-tactile policies.

Method

OmniVTA is a world-model approach to touch: it does not merely encode the current tactile image, it predicts how contact will evolve and corrects against that prediction at sensor rate. Four tightly-coupled modules:

  1. Self-supervised tactile encoder (TactileVAE). Learns compact tactile tokens from raw GelSight-Mini visuo-tactile images without action/label supervision โ€” the representation backbone for the rest of the stack.
  2. Two-stream visuo-tactile world model. Jointly ingests the vision stream and the tactile-token stream and predicts short-horizon contact evolution (the near-future tactile signal), giving the policy a forward model of contact dynamics.
  3. Contact-aware fusion policy. A diffusion policy that fuses vision + tactile + the world model's contact prediction to generate action chunks.
  4. 60 Hz Reflexive Latent Tactile Controller. At execution it runs a closed loop at 60 Hz, comparing the world model's predicted tactile against the observed tactile and emitting single-step corrective actions to recover stable contact โ€” the reflex layer that the chunk-level diffusion policy lacks.

Dataset โ€” OmniViTac. A large visuo-tactile-action dataset: 21,000+ trajectories across 86 tasks and 100+ objects, organized into six physics-grounded interaction patterns. Sensor: GelSight Mini (vision-based tactile). Proprioception, vision, and tactile are logged at 60 Hz.

Results

Real-world evaluation (per-task averaged over 5โ€“10 trials, success rate; Table III in the paper):

  • OmniVTA achieves the highest success rate across the diverse contact-rich tasks versus the baselines (notably Diffusion Policy and FoAR).
  • Closed-loop โ‰ซ open-loop: the 60 Hz reflexive controller "significantly improves performance compared with the open-loop variant," and recovers stable contact quickly under perturbation (the closed-loop advantage is the headline ablation).
  • Reported robustness to perturbations and generalization to novel objects/configurations.

(The paper reports per-task percentages in Table III; exact numbers are not reproduced here โ€” confirm against the source. No simulation benchmark, e.g. LIBERO/SIMPLER, is reported โ€” evaluation is real-robot only.)

Significance

OmniVTA is the clearest 2026 instance of touch-as-a-world-model rather than touch-as-an-input. Two contributions stand out: (1) a predictive visuo-tactile world model of short-horizon contact, and (2) a sensor-rate reflexive controller that closes the predicted-vs-observed loop โ€” directly targeting the missing reflex layer in chunk-based diffusion/VLA policies. It also adds a new role to the world-model taxonomy (see Review-World-Models): the world model as a short-horizon predictive reference for reflexive control, distinct from the backbone / RL-environment / data-factory / planner roles catalogued there.


Analysis โ€” OmniVTA in the 2026 tactile-VLA landscape (ICRA 2026 ยท CVPR 2026)

2026's tactile-for-manipulation work clusters into four camps. Placing OmniVTA against the latest ICRA 2026 tactile/force cohort (92 papers) and CVPR 2026 makes its distinctiveness โ€” and its risks โ€” precise.

Camp 1 โ€” Data-collection ergonomics (get touch data cheaply). ICRA 2026's FreeTacMan (2506.01941) โ€” a robot-free wearable visuo-tactile gripper, ~50% higher success than vision-only โ€” and ManipForce (2509.19047, handheld high-frequency F/T + RGB, 83% on six contact tasks) argue the bottleneck is data, not algorithms. OmniVTA's answer is its own large dataset (OmniViTac, 21k+ traj / 86 tasks), but its bet is on modeling, not just collection.

Camp 2 โ€” Representation & fusion (stop vision drowning out touch). ICRA 2026's Multi-Modal Policy Consensus (2509.23468, per-modality diffusion experts + a learned router), TranTac (2509.16550, a single 6-axis IMU in the gripper tip โ€” transient contact dynamics carry the signal, 79% avg), and SaTA (2510.14647, anchor tactile features to the hand frame, +30% on sub-mm USB-C mating), plus the earlier OmniVTLA (2508.08706, semantic-aligned tactile ViT). OmniVTA's TactileVAE + two-stream sits here on representation, but uniquely forecasts the tactile stream rather than only fusing the current one.

Camp 3 โ€” Sensor-free at deployment (distill touch away). This is the camp OmniVTA most sharply contrasts with. ICRA 2026's FD-VLA (2602.02142) distills a force token aligned to a real F/T latent so no force sensor is needed at inference (the distilled token reportedly beats the real reading); CVPR 2026's HapticVLA (Skoltech) trains with safety-aware tactile rewards then distills tactile knowledge into a vision-only VLA (~87% real, no sensor at deploy). OmniVTA makes the opposite bet: it keeps the tactile sensor in a 60 Hz closed loop at run time. The trade is explicit โ€” FD-VLA/HapticVLA win on hardware cost and deployability; OmniVTA wins on reflexive, contact-reactive correction that a sensor-free policy structurally cannot do.

Camp 4 โ€” Predictive world-modeling + reflexive control (OmniVTA's camp). Almost alone in the 2026 cohort, OmniVTA treats contact as something to predict and servo against rather than encode or remove. The closest neighbours are conceptual, not tactile-specific: the broader world-model-for-actions wave at ICLR 2026 (see Review-World-Models) and physically-calibrated tactile simulation like ICRA 2026's DOT-Sim (2604.27367, differentiable MPM optical-tactile sim) โ€” but DOT-Sim builds a simulator, whereas OmniVTA builds an online predictive model used for reflexive control. (Also relevant: ICRA 2026's Symmetry-Aware Vision-Tactile, 2602.13689.)

Net read for researchers.

  • OmniVTA's genuine novelty is the reflex layer: a sensor-rate predicted-vs-observed correction loop on top of chunk-level diffusion โ€” the contact analogue of what Real-Time Chunking does for chunk-boundary smoothness, but driven by tactile prediction error rather than action inpainting.
  • It is not a language-conditioned generalist VLA โ€” there is no VLM backbone or instruction generalization in the OmniVTLA / Tactile-VLA sense; it is a tactile world-model policy. A natural next step is grafting this reflex world-model onto a VLA backbone (the FD-VLA-style token-injection or a Category J sensory channel).
  • The sensor-in-the-loop vs sensor-free axis is now the field's central tactile-VLA fault line: FD-VLA / HapticVLA (remove it) vs OmniVTA (servo on it). Which wins is task-dependent โ€” sensor-free for cost-sensitive fleets, sensor-in-loop for high-precision/perturbation-heavy contact.
  • Caveat: GelSight-Mini-specific, real-robot-only evaluation (no sim benchmark, no cross-sensor generalization shown), and no head-to-head against the VLA-backbone tactile methods above โ€” so its standing relative to FD-VLA / HapticVLA on shared tasks is open.

Limitations

  • Single sensor family (GelSight Mini vision-based tactile) โ€” cross-sensor / force-based generalization not demonstrated (contrast OmniVTLA's dual-path encoder for vision- and force-based sensors).
  • Real-robot-only evaluation; no standardized sim benchmark (LIBERO/SIMPLER) for cross-paper comparison.
  • Sensor required at deployment at 60 Hz โ€” the cost/deployability trade vs the sensor-free distillation camp.
  • Not a generalist / language-conditioned VLA โ€” no instruction-following or open-vocabulary generalization claims.

Links

Related pages

โ† Back to Home

โš ๏ธ **GitHub.com Fallback** โš ๏ธ