RSS 2026 Semantically Structured Mixture of Experts for Compositional - Heungwoo/research GitHub Wiki
Semantically Structured Mixture-of-Experts for Compositional Robotic Manipulation
Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Manipulation 2 · paper #57 Authors: Chengyu Deng, Guanqi Chen, Yizhou Chen, Zejia Liu, Zhiwen Ruan, Guanhua Chen, Jia Pan arXiv: 2605.23477 · program page
Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Figure 2 in two panels: (a) the offline skill-abstraction workflow, where a VLM temporally downsamples a demonstration and segments it into open-vocabulary verb-noun skills ("approach bowl", "pick up bowl", "move towards drawer", "put bowl in drawer"); (b) the skill-conditioned policy framework — a lightweight skill predictor produces an anticipated skill token that is aligned to frozen text embeddings (InterCL) and injected into the MoE router of a 6-layer diffusion transformer, where all action-chunk tokens share one skill-conditioned routing decision over 4 experts (top-2 active) and IntraCL regularizes routing consistency.
Problem
Diffusion policies are precise but hit a scalability trilemma: precision, broad multi-task generalization, and real-time efficiency. MoE diffusion policies (e.g. MoDE, Sparse-DP) restore efficiency via sparse expert activation, but their routing is driven by low-level signals such as diffusion noise level, causing "routing confusion" at skill boundaries — reusable behaviors like grasping get fragmented across experts, hurting transfer and interpretability.
Method
SMoDP (Semantically Structured Mixture-of-Experts Diffusion Policy) grounds expert routing in language-defined skill structure. Offline, Qwen3-VL annotates demonstrations (temporal downsampling λ = 5) with open-vocabulary verb-noun skill segments and their temporal boundaries — no manual labels, no inference-time VLM calls. Online, a lightweight skill predictor (one cross-attention layer + projection) anticipates the current skill embedding from observation and instruction; a Diffusion-Transformer denoiser (6 layers, 4 experts per MoE layer, top-2 activation, action-chunk horizon H = 10) routes all action tokens in a chunk through a single skill-conditioned router token. A dual contrastive alignment ties everything together: InterCL aligns predicted skill embeddings to frozen Sentence-BERT text embeddings with continuous similarity weights, while IntraCL makes semantically similar skills produce similar routing distributions. Training combines diffusion, load-balancing and the two contrastive losses; the 401M-parameter model trains in ~4 h on one RTX 4090 and infers a chunk in ~121 ms with 10 denoising steps.
Results
On LIBERO-10/90 (50 demos/task, 3 seeds), SMoDP reaches 0.95/0.97 average success, beating DP-T, DP-CNN, QUEST, SDP, MoDE, and even OXE-pretrained MoDE (0.94/0.95 shown for the strongest baselines). With only 50% of LIBERO-90 demonstrations, it matches MoDE (Pretrain), which uses ~1.85× more parameters plus OXE pretraining. For few-shot transfer with frozen experts (updating only the 13.7M-parameter skill predictor + router, 3.4% of the model), it beats MoDE+LoRA by an average of 31.5 points on LIBERO-90→LIBERO-10 (e.g. 0.840 vs 0.509 at 10 demos) and wins on LIBERO-GOAL-OOD (0.660 vs 0.635 at 10 demos). On four real bimanual ALOHA tasks (20 demos/task, 20 rollouts), it averages 91.25% vs MoDE's 47.50%, including 20/20 on Handoff Cup where MoDE scores 0/20. Ablations: removing both contrastive losses drops LIBERO-90 from 0.970 to 0.946.
Significance
A concrete demonstration that language-grounded skill semantics is a better routing signal than noise level for MoE diffusion policies, buying data efficiency, interpretability (skill-aligned expert heatmaps), and cheap compositional transfer via frozen-expert recombination. Fits the wiki's efficiency and modularity threads, e.g. Review-VLA-Architecture and Review-Realtime-Execution.
← Back to RSS 2026 survey · RSS-2026-Papers · Home