ICLR 2026 Compose Your Policies - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Category: Training / Test-Time Trend tag: Training (compositional)
flowchart LR
P1[Policy 1<br/>diffusion or flow] --> S1[Score 1]
P2[Policy 2] --> S2[Score 2]
P3[Policy K] --> S3[Score K]
S1 --> CO[Convex combination of scores<br/>Σ wᵢ·scoreᵢ, Σ wᵢ=1]
S2 --> CO
S3 --> CO
CO --> Best[Test-time weight search<br/>per task]
Best --> Out[Composed action<br/>NO retraining]
Given multiple trained diffusion or flow-matching policies for the same task (but differing in input modality — RGB vs. point cloud, architecture — Diffusion/Flow/Mamba Policy, or conditioning — VA vs. VLA), how to combine them without retraining? Ensembling that averages outputs gives suboptimal combinations; what's needed is principled composition at the distribution level.
General Policy Composition (GPC): form a convex combination of the distributional scores of K pre-trained policies,
Consistent gains over any single base policy without additional training: Robomimic & Push-T +2–7.55% average success; RoboTwin +5–7% across six bimanual tasks; four real-world manipulation tasks improved (e.g., Clean Table 14/20 vs. 7–12/20 baselines). Composition helps most when both policies are at least moderately competent (>~30%); gains shrink when one policy is much weaker. Inference overhead is modest (~0.09s→0.13s per action chunk); the full weight search costs ~2.5h (≈1h optimized) vs. days of retraining.
Opens a third axis for improving policies (beyond "train more" and "train differently"): squeeze gains out of policies you already have by composing at inference. Practical for deployment too — keep a library of specialized policies and compose per task.
- π0.6 (flow-based policy candidate)
- Discrete Diffusion VLA (diffusion-based policy candidate)
← Back to ICLR-2026 · Topic: RL (adjacent — test-time composition)