ICLR 2026 HiMoE VLA - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 (arXiv:2512.05693, Dec 2025) Authors: Zhiying Du, Bei Liu, Yaobo Liang, Yichao Shen, Haidong Cao, Xiangyu Zheng, Zhiyuan Feng, Zuxuan Wu, Jiaolong Yang, Yu-Gang Jiang (Fudan University & Microsoft Research Asia) Category: VLA Architecture โ Cross-Embodiment Trend tag: Trend 4
The hierarchy is depth-wise across the action module, not a category classifier. Each MoE layer uses standard top-k token routing (32 experts, top-4); there is no explicit "arm vs hand" router.
flowchart TB
In[Action-module input tokens] --> AS[Shallow layers: AS-MoE<br/>action-space specialization<br/>joint-angle vs end-effector]
AS --> HB[Deeper layers: HB-MoE<br/>balances broader heterogeneity<br/>embodiment / sensors]
HB --> D[Central dense Transformer blocks<br/>shared cross-domain knowledge]
D --> A[Action]
A single dense backbone trained jointly on heterogeneous robots suffers from negative transfer: different action spaces (joint-angle vs. end-effector control), embodiment kinematics, sensor configurations, and control frequencies push the shared weights toward compromises that hurt individual robots. The paper directly demonstrates this โ co-training a dense ฯ0 across mismatched action spaces lowers its score.
Hierarchical (depth-wise) Mixture-of-Experts in the action module: shallow AS-MoE layers specialize per action space (joint-angle vs. end-effector control); deeper HB-MoE layers balance broader heterogeneity (embodiment kinematics, sensor configs); central dense Transformer blocks consolidate the heterogeneous signals into shared representations for cross-domain transfer. Each MoE layer routes tokens with standard top-k gating (32 experts, top-4).
Two regularizers shape specialization:
- AS-Reg โ a contrastive objective treating experts assigned to the same action-space token as positive pairs, sharpening action-space specialization.
- HB-Reg โ aligns empirical routing frequency with expected routing probability, spreading heterogeneous inputs evenly across experts (load balancing).
The VLM backbone is ~4B params.
State-of-the-art on simulation and real robots. LIBERO: 97.8% overall avg (vs UniVLA 95.2%, OpenVLA-OFT 97.1%, ฯ0 94.2%). CALVIN DโD: 3.967 avg consecutive tasks (vs ฯ0 3.758). Real xArm7: 75.0% avg success (ฯ0 62.5%). Real ALOHA (dual-arm): 63.7% (ฯ0 54.2%, RDT-1B 47.5%).
Decisive ablation on negative transfer: co-training CALVIN-ABC (EEF) + CALVIN-D (joint), ฯ0 degrades โ0.259 (3.806โ3.547) while HiMoE improves +0.186 (3.826โ4.012). Removing all MoE layers drops the model to 3.777 vs 4.012; a single flat MoE (3.813) underperforms the hierarchy, confirming depth-wise separation matters.
Limitations (per authors): features from all VLM layers are fused without selective weighting; model scale (~4B) is modest, constrained by available robotics data. Note: the method demonstrates quick adaptation / fine-tuning to new robots (e.g. 50k steps), not zero per-robot fine-tuning.
The MoE-based counterpart to X-VLA (soft prompts) and XR-1 (UVMC). All three are bets on where in the architecture embodiment heterogeneity should be handled โ HiMoE places it in the feed-forward layers, X-VLA in the input prompt, XR-1 in the tokenizer / shared codebook.
- In-depth review: HiMoE-VLA โ in-depth ยท multi-task context: Multi-Task VLA
- arXiv:2512.05693
- OpenReview
- ICLR 2026 listing
โ Back to ICLR-2026