ICML 2026 DECO - Heungwoo/research GitHub Wiki
DECO: Decoupled Multimodal Diffusion Transformer for Bimanual Dexterous Manipulation with a Plugin Tactile Adapter โ Disentangling vision, proprioception, and touch for contact-rich dual-arm control
Venue: ICML 2026 (Poster) Category: Tactile Traction (2026-06): 0 citations (arXiv)
Problem
Bimanual dexterous manipulation depends on integrating heterogeneous sensory modalities โ vision, proprioception, and tactile feedback โ but most policies fuse these signals in a coupled, undifferentiated way. DECO argues that different modalities contribute unequally to action generation and therefore should influence the policy through distinct mechanisms rather than being concatenated into a single mixed conditioning vector. A second obstacle is data: there is little large-scale, real-robot bimanual data that pairs dexterous hands with tactile sensing.
Method
DECO is a decoupled multimodal diffusion transformer whose velocity predictor is built on a Multi-Modal Diffusion Transformer (MMDiT) block. The core principle is to decouple modality-specific conditioning from the attention structure so each sensory stream enters through its own pathway:
- Vision (the dominant modality) is fused via joint self-attention between visual tokens and action tokens, letting images directly guide action generation. Binocular images are encoded with a shared ResNet-34 backbone, flattened into tokens with rotary positional embeddings (RoPE) applied per-view, then concatenated.
- Proprioception is injected via AdaLN (adaptive layer normalization).
- Tactile signals enter through cross-attention.

For tactile, DECO adds a plug-in tactile adapter for parameter-efficient conditioning. A tactile encoder produces region-level features (averaging raw values within each pad โ finger tips, pulps, ends, palm โ plus a learnable linear projection), and LoRA selectively fine-tunes the attention layers of a frozen, pretrained visionโaction backbone. During this second stage only the adapter is optimized, tuning less than 10% of model parameters.
Alongside the model, the authors release DECO-50: 50 hours / over 5M frames of teleoperated real dual-arm bimanual dexterous manipulation data with tactile sensing.
Results
Trained on DECO-50 and evaluated over 2,000+ real-robot rollouts across four tasks (Pick and Place, Material Sorting on a conveyor, Waste Disposal, Assembly), DECO achieves a 72.25% average success rate, a 21% improvement over the baseline, and the best performance on all tasks. The tactile adapter adds +10.25% average success across tasks and +20% on complex contact-rich tasks. On Material Sorting, DECO reaches 101/120 (and 105/120 with tactile) versus 97/120 for ACT and 67/120 for Diffusion Policy. Ablations show that coupled tactile injection (DECO.cs) underperforms the decoupled cross-attention design, particularly on the hardest interference-fit assembly stages.
Significance
DECO contributes both an architectural insight โ that disentangling modality-specific conditioning pathways improves controllable multimodal fusion โ and a substantial open bimanual tactile dataset. The LoRA-based plug-in adapter is notable for delivering large contact-rich gains while keeping the pretrained vision policy frozen and tuning under 10% of parameters, making tactile a cheap add-on rather than a from-scratch retrain.
Links
- arXiv: 2602.05513
- ICML 2026: https://icml.cc/virtual/2026/poster/66358
โ Back to ICML-2026