ICLR 2026 MoE Diffusion Skills - Heungwoo/research GitHub Wiki

SMP — abstracting manipulation skills via MoE diffusion policies

Venue: ICLR 2026 Authors: Ce Hao, Xuanran Zhai, Yaohua Liu, Harold Soh arXiv: 2601.21251 Category: VLA architecture — diffusion policy / mixture-of-experts Trend tag: Skill abstraction / MoE diffusion / efficient multi-task

Approach diagram

flowchart LR
  Obs[Observation] --> Router[Sticky router]
  Router --> Sel[Select small task-relevant<br/>subset of experts per step]
  Basis[Compact orthogonal<br/>skill basis] --> Experts[Skill experts]
  Sel --> Experts
  Experts --> Diff[Diffusion action denoiser]
  Diff --> Act[Action chunk]
  Var[Variational training objective] -.trains.-> Basis
  Var -.trains.-> Router
  Act --> Inf[Adaptive expert activation<br/>fast sampling, no oversized backbone]
Loading

Problem

Diffusion-based policies perform well in robot manipulation, but extending them to multi-task settings is costly: scaling model size and demonstrations to cover many tasks is expensive, and large monolithic backbones make inference slow. The paper seeks reusable, composable skills that can be selectively activated per task.

Method

SMP (Skill Mixture-of-Experts Policy) is a diffusion-based MoE policy that:

  • Learns a compact orthogonal skill basis so experts capture distinct, reusable skills.
  • Uses sticky routing to compose actions from a small, task-relevant subset of experts at each step.
  • Is trained with a variational objective supporting this design.
  • At inference, uses adaptive expert activation for fast sampling without an oversized backbone.

Results

SMP is validated in simulation and on a real dual-arm platform across multi-task learning and transfer-learning tasks, where it reports higher success rates and markedly lower inference cost than large diffusion baselines, and enables rapid adaptation when the task set changes.

Significance

SMP frames skill abstraction as learning an orthogonal expert basis plus sparse, sticky routing — a practical route to scalable, transferable multi-task diffusion policies that avoids paying for a monolithic large backbone at inference. It connects the MoE trend in VLAs with diffusion action heads.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️