Review HiMoE VLA - Heungwoo/research GitHub Wiki
Paper: "HiMoE-VLA: Hierarchical Mixture-of-Experts for Generalist Vision-Language-Action Policies" ā arXiv 2512.05693 Ā· ICLR 2026 Ā· Fudan University & Microsoft Research Asia (Zhiying Du, Bei Liu, ⦠Yu-Gang Jiang) What it is: the MoE-routing answer to multi-task/cross-embodiment negative transfer ā put a depth-wise mixture-of-experts inside the action module so heterogeneous tasks don't average out. Cluster B in Multi-Task VLA §3. Venue entry: ICLR-2026-HiMoE-VLA Ā· companions X-VLA Ā· XR-1 Ā· Cross-Embodiment.
- Negative transfer is real and measurable. Co-training a dense Ļ0 across mismatched action spaces (joint-angle vs end-effector) lowers its score ā the paper's decisive datapoint that one dense backbone can't serve heterogeneous robots.
- Fix = depth-wise hierarchical MoE in the action module. Shallow AS-MoE layers specialize by action space; deeper HB-MoE layers absorb broader heterogeneity (embodiment kinematics, sensors); central dense blocks consolidate the shared cross-domain knowledge. Standard top-k gating (32 experts, top-4); no hand-built "arm vs hand" router.
- Two regularizers shape the experts: AS-Reg (contrastive ā same-action-space tokens are positive pairs) sharpens action-space specialization; HB-Reg (route-frequency ā expected-probability alignment) load-balances.
- SOTA across sim + real on a ~4B backbone: LIBERO 97.8%, CALVIN DāD 3.967, real xArm7 75.0%, real dual-arm ALOHA 63.7% ā all above Ļ0.
- It localizes the interference. The multi-task failure (Review-Multitask-VLA §1) is dense parameter sharing forcing compromises. HiMoE's answer is structural: give each action-space/embodiment its own feed-forward experts, and only consolidate in the central dense blocks ā so specialization and sharing happen in different layers.
- Depth matters, not just sparsity. A single flat MoE (3.813) underperforms the hierarchy (4.012), and removing MoE entirely drops to 3.777 ā evidence that where you separate (shallow=action-space, deep=embodiment) is the contribution, not merely "add experts."
- Complementary to the other cross-embodiment bets. It's the MoE-in-FFN counterpart to X-VLA (heterogeneity in the input prompt) and XR-1 (heterogeneity in the shared codebook) ā three different loci for the same problem.
flowchart TB
In[Action-module input tokens] --> AS[Shallow ā AS-MoE<br/>action-space specialization<br/>joint-angle vs end-effector Ā· top-4/32]
AS --> HB[Deeper ā HB-MoE<br/>embodiment / sensor heterogeneity Ā· top-4/32]
HB --> D[Central dense Transformer blocks<br/>shared cross-domain knowledge]
D --> A[Action]
ASR[AS-Reg ā contrastive on same-action-space experts] -.-> AS
HBR[HB-Reg ā route-frequency load balance] -.-> HB
- Depth-wise hierarchy (not a category classifier): shallow layers handle the coarsest split (action space); deeper layers handle finer heterogeneity; dense blocks fuse.
- Routing: standard token-level top-k (32 experts, top-4) per MoE layer.
- AS-Reg: contrastive objective pulling together experts routed by same-action-space tokens ā cleaner action-space specialization.
- HB-Reg: aligns empirical routing frequency with expected probability ā even expert utilization (anti-collapse).
- Backbone: ~4B-param VLM.
| Benchmark | HiMoE-VLA | Ļ0 | others |
|---|---|---|---|
| LIBERO (avg) | 97.8% | 94.2% | OpenVLA-OFT 97.1 Ā· UniVLA 95.2 |
| CALVIN DāD (avg consecutive) | 3.967 | 3.758 | ā |
| Real xArm7 (avg success) | 75.0% | 62.5% | ā |
| Real ALOHA dual-arm | 63.7% | 54.2% | RDT-1B 47.5 |
Negative-transfer ablation (co-train CALVIN-ABC EEF + CALVIN-D joint): dense Ļ0 degrades ā0.259 (3.806ā3.547) while HiMoE improves +0.186 (3.826ā4.012). Remove-all-MoE ā 3.777; flat MoE 3.813 < hierarchy 4.012.
Significance. HiMoE-VLA is the cleanest demonstration that hierarchical, depth-wise MoE converts cross-task/embodiment heterogeneity from a liability (negative transfer) into a gain ā the routing-based cure in the multi-task fix landscape.
Limitations (authors' + reviewer's).
- All VLM-layer features are fused without selective weighting ā a coarse consolidation.
- ~4B is modest ā scale is constrained by available robotics data; behavior at larger scale / many more experts is untested.
- Not zero-shot to new robots ā it demonstrates quick adaptation (e.g. 50k steps), not per-robot-fine-tune-free deployment (contrast single-checkpoint deployment).
- Router scaling ā top-k gating at 32 experts is shown; hundreds of tasks/experts unproven.
Design lesson. Separate what conflicts (action-space in shallow experts, embodiment in deep experts) and share what transfers (central dense blocks) ā depth-wise placement beats a flat expert pool.
- Paper: arXiv 2512.05693 Ā· OpenReview Ā· venue entry ICLR-2026-HiMoE-VLA
- Multi-task context: Multi-Task VLA (cluster B) Ā· sibling optimizer approach DyGRO-VLA
- Cross-embodiment kin: X-VLA Ā· XR-1 Ā· Cross-Embodiment Ā· VLA Architectures