ICLR 2026 Compose Your Policies - Heungwoo/research GitHub Wiki

Compose Your Policies! — Test-Time Distribution-Level Composition

Venue: ICLR 2026 Category: Training / Test-Time Trend tag: Training (compositional)

Approach diagram

flowchart LR
  P1[Policy 1<br/>diffusion or flow] --> S1[Score 1]
  P2[Policy 2] --> S2[Score 2]
  P3[Policy K] --> S3[Score K]
  S1 --> CO[Convex combination of scores<br/>Σ wᵢ·scoreᵢ, Σ wᵢ=1]
  S2 --> CO
  S3 --> CO
  CO --> Best[Test-time weight search<br/>per task]
  Best --> Out[Composed action<br/>NO retraining]
Loading

Problem

Given multiple trained diffusion or flow-matching policies for the same task (but differing in input modality — RGB vs. point cloud, architecture — Diffusion/Flow/Mamba Policy, or conditioning — VA vs. VLA), how to combine them without retraining? Ensembling that averages outputs gives suboptimal combinations; what's needed is principled composition at the distribution level.

Method

General Policy Composition (GPC): form a convex combination of the distributional scores of K pre-trained policies, $\sum_i w_i, s_\theta(\tau_t, t, c_i)$ with $\sum_i w_i = 1$, then perform a test-time weight search (discrete grid, 0.0–1.0 in 0.1 steps; optionally constrain the stronger policy to weight ≥ 0.6). No optimization-based solve and no retraining. Both diffusion and flow-based policies are supported. Theoretical backing: (1) a convex combination of score estimators can attain lower MSE than any individual score (one-step functional improvement); (2) a Grönwall-type bound shows this single-step gain propagates through the full generation trajectory rather than blowing up, yielding systemic improvement. Non-convex "logical" variants (AND/OR superposition) are also explored.

Results

Consistent gains over any single base policy without additional training: Robomimic & Push-T +2–7.55% average success; RoboTwin +5–7% across six bimanual tasks; four real-world manipulation tasks improved (e.g., Clean Table 14/20 vs. 7–12/20 baselines). Composition helps most when both policies are at least moderately competent (>~30%); gains shrink when one policy is much weaker. Inference overhead is modest (~0.09s→0.13s per action chunk); the full weight search costs ~2.5h (≈1h optimized) vs. days of retraining.

Significance

Opens a third axis for improving policies (beyond "train more" and "train differently"): squeeze gains out of policies you already have by composing at inference. Practical for deployment too — keep a library of specialized policies and compose per task.

Links

Related pages

← Back to ICLR-2026 · Topic: RL (adjacent — test-time composition)

⚠️ **GitHub.com Fallback** ⚠️