ICML 2026 Decompose and Recompose - Heungwoo/research GitHub Wiki
Decompose and Recompose: Reasoning New Skills from Existing Abilities for Cross-Task Robotic Manipulation — Atomic skill–action pairs for zero-shot cross-task generalization
Venue: ICML 2026 (Poster) Category: Reasoning Affiliations: Xitie Zhang, Aming Wu, Yahong Han

Problem
Cross-task generalization is a core challenge in open-world robotic manipulation, and the key is extracting transferable manipulation knowledge from seen tasks. Recent in-context learning (ICL) approaches feed seen-task demonstrations to a large model to generate actions for unseen tasks without parameter updates. However, existing methods provide only low-level continuous action sequences as context, which fails to capture composable skill knowledge and causes the model to degenerate into superficial trajectory imitation.
Method
Decompose and Recompose is a skill-reasoning framework built on atomic skill–action pairs as intermediate representations that bridge high-level task semantics and low-level control.

- Decompose: seen demonstrations are broken into interpretable skill–action alignments (atomic skill–action pairs) via keyframe detection, producing composable intermediate representations.
- Dual-library demonstration retrieval: a task-adaptive dynamic library is built via visual-semantic retrieval combined with skill sequences from a planning agent, and a coverage-aware static library complements it to fill missing skill patterns. Together they yield skill-comprehensive demonstration sets.
- Recompose: the skill-augmented demonstrations explicitly elicit the LLM's compositional reasoning, so it composes existing skills and infers execution ordering for unseen tasks — zero-shot, with no parameter updates.
For execution, continuous control u = [p, q, g] (end-effector position, orientation quaternion, binary gripper) interfaces with a standard RLBench motion planner; for LLM interaction, translation and rotation are discretized into integer bins and represented as a 7-tuple action token, then decoded back to continuous control.
Results
Evaluated on the AGNOSTOS benchmark (23 unseen tasks split into Level-1 and Level-2 tiers; standard protocol of 25 rollouts × 3 seeds = 75 episodes per task) plus real-world experiments:
- The method achieves the highest overall success rates across both difficulty levels versus diverse VLA baselines (foundation VLA models and other ICL approaches).
- It exceeds 60% success on four distinct tasks (Microwave, Seat, LampOff, USB), whereas individual baselines hit that threshold on at most three.
- Ablations: with no in-context demonstrations the model fails entirely (0% success) — demonstrations are essential. Overall success rises from 19.8% → 26.4% as demonstrations grow from 5 to 20 (Level-1 27.3%→32.5%, Level-2 13.1%→18.5%). For the coverage-aware static library, performance climbs from 24.9% (dynamic-only) to a best 26.4% at 3 static demos, then slightly drops at 4 — indicating an optimal balance between coverage and task relevance.
- Real-world tasks (e.g., stack cups, stack blocks) confirm the framework generates appropriate skill sequences and executes precise actions across diverse scenarios.
Significance
By making atomic skill–action pairs the unit of in-context reasoning, Decompose and Recompose moves beyond trajectory imitation toward genuine compositional skill reuse, enabling training-free zero-shot transfer to novel objects and goals. The dual-library retrieval strategy — pairing task-adaptive dynamic retrieval with coverage-aware static complementation — is a practical recipe for supplying an LLM with the right skill vocabulary, and it sets a new bar on AGNOSTOS cross-task generalization.
Links
- arXiv: 2605.01448
- ICML 2026: https://icml.cc/virtual/2026/poster/63250
← Back to ICML-2026