ICLR 2026 HiMoE VLA - Heungwoo/research GitHub Wiki

HiMoE-VLA โ€” Hierarchical Mixture-of-Experts for Generalist VLA

Venue: ICLR 2026 (arXiv:2512.05693, Dec 2025) Authors: Zhiying Du, Bei Liu, Yaobo Liang, Yichao Shen, Haidong Cao, Xiangyu Zheng, Zhiyuan Feng, Zuxuan Wu, Jiaolong Yang, Yu-Gang Jiang (Fudan University & Microsoft Research Asia) Category: VLA Architecture โ€” Cross-Embodiment Trend tag: Trend 4

Approach diagram

The hierarchy is depth-wise across the action module, not a category classifier. Each MoE layer uses standard top-k token routing (32 experts, top-4); there is no explicit "arm vs hand" router.

flowchart TB
  In[Action-module input tokens] --> AS[Shallow layers: AS-MoE<br/>action-space specialization<br/>joint-angle vs end-effector]
  AS --> HB[Deeper layers: HB-MoE<br/>balances broader heterogeneity<br/>embodiment / sensors]
  HB --> D[Central dense Transformer blocks<br/>shared cross-domain knowledge]
  D --> A[Action]
Loading

Problem

A single dense backbone trained jointly on heterogeneous robots suffers from negative transfer: different action spaces (joint-angle vs. end-effector control), embodiment kinematics, sensor configurations, and control frequencies push the shared weights toward compromises that hurt individual robots. The paper directly demonstrates this โ€” co-training a dense ฯ€0 across mismatched action spaces lowers its score.

Method

Hierarchical (depth-wise) Mixture-of-Experts in the action module: shallow AS-MoE layers specialize per action space (joint-angle vs. end-effector control); deeper HB-MoE layers balance broader heterogeneity (embodiment kinematics, sensor configs); central dense Transformer blocks consolidate the heterogeneous signals into shared representations for cross-domain transfer. Each MoE layer routes tokens with standard top-k gating (32 experts, top-4).

Two regularizers shape specialization:

  • AS-Reg โ€” a contrastive objective treating experts assigned to the same action-space token as positive pairs, sharpening action-space specialization.
  • HB-Reg โ€” aligns empirical routing frequency with expected routing probability, spreading heterogeneous inputs evenly across experts (load balancing).

The VLM backbone is ~4B params.

Results

State-of-the-art on simulation and real robots. LIBERO: 97.8% overall avg (vs UniVLA 95.2%, OpenVLA-OFT 97.1%, ฯ€0 94.2%). CALVIN Dโ†’D: 3.967 avg consecutive tasks (vs ฯ€0 3.758). Real xArm7: 75.0% avg success (ฯ€0 62.5%). Real ALOHA (dual-arm): 63.7% (ฯ€0 54.2%, RDT-1B 47.5%).

Decisive ablation on negative transfer: co-training CALVIN-ABC (EEF) + CALVIN-D (joint), ฯ€0 degrades โˆ’0.259 (3.806โ†’3.547) while HiMoE improves +0.186 (3.826โ†’4.012). Removing all MoE layers drops the model to 3.777 vs 4.012; a single flat MoE (3.813) underperforms the hierarchy, confirming depth-wise separation matters.

Limitations (per authors): features from all VLM layers are fused without selective weighting; model scale (~4B) is modest, constrained by available robotics data. Note: the method demonstrates quick adaptation / fine-tuning to new robots (e.g. 50k steps), not zero per-robot fine-tuning.

Significance

The MoE-based counterpart to X-VLA (soft prompts) and XR-1 (UVMC). All three are bets on where in the architecture embodiment heterogeneity should be handled โ€” HiMoE places it in the feed-forward layers, X-VLA in the input prompt, XR-1 in the tokenizer / shared codebook.

Links

Related pages

โ† Back to ICLR-2026

โš ๏ธ **GitHub.com Fallback** โš ๏ธ