ICML 2026 From Noise to Control - Heungwoo/research GitHub Wiki

From Noise to Control: Parameterized Diffusion Policies — steering diffusion policies via a geometry-aligned behavior latent

Venue: ICML 2026 (Poster) Category: Diffusion-Flow Policy Affiliations: Renhao Zhang, Haotian Fu, Mingxi Jia, George Konidaris, Yilun Du, Bruno Castro da Silva Traction (2026-06): 0 citations (arXiv)

Observation-side shift vs. constraint-induced behavior shift, and why PDP enables stable behavior steering (Figure 1 from Zhang et al., 2026)

Problem

Diffusion policies (DP) are a strong default for behavior cloning because they model multimodal, high-dimensional action sequences and produce diverse rollouts via stochastic denoising. But diversity is not the same as controllability. Standard evaluations stress observation-side distribution shift (appearance/camera changes that should not alter the intended behavior), whereas many real deployment failures arise from constraint-induced behavior shifts: a new obstacle, contact constraint, or feasibility condition invalidates previously demonstrated trajectory modes, forcing the robot to select or discover a new mode. In standard DP, steering via the sampling noise is ill-conditioned — small perturbations cause large, unpredictable trajectory changes.

Method

The authors propose Parameterized Diffusion Policy (PDP), a diffusion policy conditioned on a low-dimensional, continuous latent code z embedded in a learned behavior manifold whose geometry is aligned with trajectory similarity.

  • Geometry-aligned behavior latent. A trajectory encoder E_ϕ maps each demonstration τ to a Gaussian posterior q_ϕ(z|τ); the deterministic mean z̄(τ) = μ_ϕ(τ) is used for conditioning. A symmetric decoder D_ψ reconstructs τ during training so the latent preserves behavioral information.
  • Joint embedding objective L_embed = L_rec + β_KL·L_KL + β_geo·L_geo — a standard VAE ELBO (reconstruction + KL-to-N(0,I) prior) plus a geometry-alignment term. The geometry term uses (Soft-)DTW over positional/proprioceptive features X(τ) so that Euclidean distance in latent space matches the physical trajectory distance, making z a smooth space for search and interpolation.
  • Latent-conditioned denoising via global modulation. The diffusion denoiser is conditioned on z through global modulation, turning diffusion from a stochastic-diversity mechanism into a precise, optimizable steering tool.
  • Test-time control. Adaptation is done by optimizing only z via gradient-based latent fitting (with an encoder warm-start) — no policy-weight updates. Long-horizon adaptation uses segmented latents.

PDP framework: trajectory encoder embeds demonstrations into a behavior latent z; latent-conditioned diffusion policy (Figure 2 from Zhang et al., 2026)

Results

On multimodal benchmarks in both simulation and real-robot experiments, PDP achieves consistently high success in the original (unshifted) multimodal domains, indicating that conditioning on the behavior latent does not compromise fidelity. Under constraint-induced shifts, PDP maintains high success across all evaluations while baselines (standard DP, BC, BC-GMM, IBC, Diff-ES, ADPro, BESO, VQ-BeT) degrade severely: in the Existing Mode Fitting setting PDP succeeds nearly perfectly, and in the Novel Behavior Discovery (zero-mode-feasible) setting PDP is the only method that consistently succeeds — including over BC, which is allowed to update its weights from the new demonstration. The paper presents these primarily as qualitative/relative comparisons (Tables 1–2 plus trajectory visualizations) rather than a single headline number.

Significance

PDP reframes diffusion-policy adaptation as navigation in a geometry-aligned behavior space rather than weight fine-tuning or brittle noise-space optimization. This gives a principled, low-dimensional handle for both interpolating between known strategies and synthesizing genuinely novel behaviors when constraints change — a direct attack on a failure mode (constraint-induced shift) that standard VLA/diffusion evaluations largely ignore.

Links

← Back to ICML-2026