ICML 2026 SkillNet - Heungwoo/research GitHub Wiki
Venue: ICML 2026 (Poster) Category: VLA Architecture Affiliations: Senwei Xie, Yuntian Zhang, Zhenzhou Tan, Ruiping Wang, Pengwei Wang, Shanghang Zhang, Xilin Chen
Transfer across diverse task compositions and unseen behaviors remains a major challenge for vision-language-action (VLA) models. Skills are repeatable, atomic components shared across many tasks, and similarities between skills are evidence for transferability across behaviors. Yet existing skill-centric methods have two problems. First, skills are loosely organized, lacking a hierarchy that captures similarities and differences across skills. Second, they lack a mechanism with the capacity to express transferable skill attributes in a structured parametric space.
SkillNet models skill attributes hierarchically and regulates a compositional model structure with those transferable attributes.
- Hierarchical skill modeling. SkillNet exploits motion code and the VerbNet framework to explicitly model the similarities of skills along two axes — mechanical properties and semantic roles — and organizes skills into a hierarchy that captures both their similarities and their differences.
- Compositional structure via MoE. Building on this hierarchy, SkillNet leverages the scalability of a mixture-of-experts (MoE) mechanism. It develops skill embeddings as soft constraints so that similar skills induce similar expert activations, enabling compositional generalization across task compositions and to unseen behaviors.
flowchart TD
A[Skill set] --> B[Motion code]
A --> C[VerbNet framework]
B --> D[Hierarchical skill attributes<br/>mechanical + semantic]
C --> D
D --> E[Transferable skill embeddings<br/>soft constraints]
E --> F[MoE expert activation<br/>similar skills -> similar experts]
F --> G[Compositional generalization<br/>zero-shot / few-shot transfer]
SkillNet is evaluated on zero-shot and few-shot transfer experiments in both simulators and real-world environments.
- Zero-shot transfer: improvement of 16.0%.
- Few-shot transfer: improvement of 23.9%.
- In-domain settings: achieves state-of-the-art performance.
These gains demonstrate that organizing skills into a similarity-aware hierarchy and translating that hierarchy into structured, soft constraints on expert activation yields measurable transfer benefits over loosely organized skill-centric baselines.
SkillNet's contribution is a principled bridge between symbolic/linguistic skill structure (motion codes and VerbNet's verb semantics) and the parametric structure of a modern VLA (MoE experts). By making skill similarity an explicit, transferable signal rather than an emergent property, it offers a route to compositional generalization — combining known skills into novel task compositions and extending to unseen behaviors — while still topping in-domain benchmarks. This positions structured skill hierarchies as a useful inductive bias for scaling VLA models to open-ended manipulation.
- ICML 2026: https://icml.cc/virtual/2026/poster/65559
← Back to ICML-2026