ICML 2026 DECO - Heungwoo/research GitHub Wiki

DECO: Decoupled Multimodal Diffusion Transformer for Bimanual Dexterous Manipulation with a Plugin Tactile Adapter โ€” Disentangling vision, proprioception, and touch for contact-rich dual-arm control

Venue: ICML 2026 (Poster) Category: Tactile Traction (2026-06): 0 citations (arXiv)

Problem

Bimanual dexterous manipulation depends on integrating heterogeneous sensory modalities โ€” vision, proprioception, and tactile feedback โ€” but most policies fuse these signals in a coupled, undifferentiated way. DECO argues that different modalities contribute unequally to action generation and therefore should influence the policy through distinct mechanisms rather than being concatenated into a single mixed conditioning vector. A second obstacle is data: there is little large-scale, real-robot bimanual data that pairs dexterous hands with tactile sensing.

Method

DECO is a decoupled multimodal diffusion transformer whose velocity predictor is built on a Multi-Modal Diffusion Transformer (MMDiT) block. The core principle is to decouple modality-specific conditioning from the attention structure so each sensory stream enters through its own pathway:

  • Vision (the dominant modality) is fused via joint self-attention between visual tokens and action tokens, letting images directly guide action generation. Binocular images are encoded with a shared ResNet-34 backbone, flattened into tokens with rotary positional embeddings (RoPE) applied per-view, then concatenated.
  • Proprioception is injected via AdaLN (adaptive layer normalization).
  • Tactile signals enter through cross-attention.

Multimodal Diffusion Transformer block with decoupled conditioning โ€” images via self-attention, proprioception via AdaLN, tactile via cross-attention (Figure 3 from Li et al., 2026)

For tactile, DECO adds a plug-in tactile adapter for parameter-efficient conditioning. A tactile encoder produces region-level features (averaging raw values within each pad โ€” finger tips, pulps, ends, palm โ€” plus a learnable linear projection), and LoRA selectively fine-tunes the attention layers of a frozen, pretrained visionโ€“action backbone. During this second stage only the adapter is optimized, tuning less than 10% of model parameters.

Alongside the model, the authors release DECO-50: 50 hours / over 5M frames of teleoperated real dual-arm bimanual dexterous manipulation data with tactile sensing.

Results

Trained on DECO-50 and evaluated over 2,000+ real-robot rollouts across four tasks (Pick and Place, Material Sorting on a conveyor, Waste Disposal, Assembly), DECO achieves a 72.25% average success rate, a 21% improvement over the baseline, and the best performance on all tasks. The tactile adapter adds +10.25% average success across tasks and +20% on complex contact-rich tasks. On Material Sorting, DECO reaches 101/120 (and 105/120 with tactile) versus 97/120 for ACT and 67/120 for Diffusion Policy. Ablations show that coupled tactile injection (DECO.cs) underperforms the decoupled cross-attention design, particularly on the hardest interference-fit assembly stages.

Significance

DECO contributes both an architectural insight โ€” that disentangling modality-specific conditioning pathways improves controllable multimodal fusion โ€” and a substantial open bimanual tactile dataset. The LoRA-based plug-in adapter is notable for delivering large contact-rich gains while keeping the pretrained vision policy frozen and tuning under 10% of parameters, making tactile a cheap add-on rather than a from-scratch retrain.

Links

โ† Back to ICML-2026