ICML 2026 From Noise to Control - Heungwoo/research GitHub Wiki
From Noise to Control: Parameterized Diffusion Policies — steering diffusion policies via a geometry-aligned behavior latent
Venue: ICML 2026 (Poster) Category: Diffusion-Flow Policy Affiliations: Renhao Zhang, Haotian Fu, Mingxi Jia, George Konidaris, Yilun Du, Bruno Castro da Silva Traction (2026-06): 0 citations (arXiv)

Problem
Diffusion policies (DP) are a strong default for behavior cloning because they model multimodal, high-dimensional action sequences and produce diverse rollouts via stochastic denoising. But diversity is not the same as controllability. Standard evaluations stress observation-side distribution shift (appearance/camera changes that should not alter the intended behavior), whereas many real deployment failures arise from constraint-induced behavior shifts: a new obstacle, contact constraint, or feasibility condition invalidates previously demonstrated trajectory modes, forcing the robot to select or discover a new mode. In standard DP, steering via the sampling noise is ill-conditioned — small perturbations cause large, unpredictable trajectory changes.
Method
The authors propose Parameterized Diffusion Policy (PDP), a diffusion policy conditioned on a low-dimensional, continuous latent code z embedded in a learned behavior manifold whose geometry is aligned with trajectory similarity.
- Geometry-aligned behavior latent. A trajectory encoder
E_ϕmaps each demonstrationτto a Gaussian posteriorq_ϕ(z|τ); the deterministic meanz̄(τ) = μ_ϕ(τ)is used for conditioning. A symmetric decoderD_ψreconstructsτduring training so the latent preserves behavioral information. - Joint embedding objective
L_embed = L_rec + β_KL·L_KL + β_geo·L_geo— a standard VAE ELBO (reconstruction + KL-to-N(0,I) prior) plus a geometry-alignment term. The geometry term uses (Soft-)DTW over positional/proprioceptive featuresX(τ)so that Euclidean distance in latent space matches the physical trajectory distance, makingza smooth space for search and interpolation. - Latent-conditioned denoising via global modulation. The diffusion denoiser is conditioned on
zthrough global modulation, turning diffusion from a stochastic-diversity mechanism into a precise, optimizable steering tool. - Test-time control. Adaptation is done by optimizing only
zvia gradient-based latent fitting (with an encoder warm-start) — no policy-weight updates. Long-horizon adaptation uses segmented latents.

Results
On multimodal benchmarks in both simulation and real-robot experiments, PDP achieves consistently high success in the original (unshifted) multimodal domains, indicating that conditioning on the behavior latent does not compromise fidelity. Under constraint-induced shifts, PDP maintains high success across all evaluations while baselines (standard DP, BC, BC-GMM, IBC, Diff-ES, ADPro, BESO, VQ-BeT) degrade severely: in the Existing Mode Fitting setting PDP succeeds nearly perfectly, and in the Novel Behavior Discovery (zero-mode-feasible) setting PDP is the only method that consistently succeeds — including over BC, which is allowed to update its weights from the new demonstration. The paper presents these primarily as qualitative/relative comparisons (Tables 1–2 plus trajectory visualizations) rather than a single headline number.
Significance
PDP reframes diffusion-policy adaptation as navigation in a geometry-aligned behavior space rather than weight fine-tuning or brittle noise-space optimization. This gives a principled, low-dimensional handle for both interpolating between known strategies and synthesizing genuinely novel behaviors when constraints change — a direct attack on a failure mode (constraint-induced shift) that standard VLA/diffusion evaluations largely ignore.
Links
- arXiv: 2606.00336
- ICML 2026: https://icml.cc/virtual/2026/poster/62938
← Back to ICML-2026