ICML 2026 PACT - Heungwoo/research GitHub Wiki

PACT: Self-Evolving Physical Safety Alignment for Diffusion Policies in Embodied Manipulation — Projecting trained diffusion policies into constraint-feasible regions, no demos or rewards needed

Venue: ICML 2026 (Poster) Category: Diffusion-Flow Policy (Physical Safety Alignment) Affiliations: Lingxuan Wu, Zijian Zhu, Lizhong Wang, Chengyang Ying, Huayu Chen, Xiao Yang, Fangming Liu, Jun Zhu

Problem

Diffusion policies have achieved remarkable success in robotic manipulation, but they frequently fail to satisfy strict physical constraints required for safe deployment — collision limits, workspace boundaries, and similar hard constraints. The usual fixes (collecting safe demonstrations or designing task rewards) are expensive and task-specific. PACT asks whether an already-trained diffusion policy can be re-aligned into the safe region purely as a post-training step, without any new demonstration data and without task rewards.

Method

PACT is a post-training alignment framework that projects an existing, pretrained diffusion policy into constraint-compliant regions. Its core mechanism is to distill constraint gradients into the diffusion model through a reverse-KL objective with dense supervision across timesteps — so that the safety signal shapes the denoising process at every diffusion step rather than only at the final action.

A central component is a constraint-tightening curriculum: the framework gradually increases how strictly constraints are enforced while explicitly preventing excessive deviation from the original policy. This balance ensures the aligned policy keeps improving on the task instead of collapsing into an over-conservative or degenerate behavior. Because supervision comes from constraint gradients rather than labeled safe trajectories, PACT needs neither demonstrations nor task rewards.

flowchart LR
    A[Pretrained diffusion policy] --> B[Reverse-KL distillation<br/>of constraint gradients]
    B --> C[Dense supervision<br/>across diffusion timesteps]
    C --> D[Constraint-tightening<br/>curriculum]
    D -->|prevents excessive deviation| E[Constraint-feasible<br/>diffusion policy]
Loading

Results

On both simulated and real robotic manipulation tasks, PACT delivers substantial improvements on the safety/performance trade-off: a 31.0% reduction in safety violations on average while improving task success by 30.7%. The simultaneous gains on both axes are the key claim — alignment tightens constraint satisfaction without paying the usual price in task performance, and in fact raises it, which the authors attribute to the curriculum's deviation-control keeping the policy close to a high-performing region.

Significance

PACT addresses a practical bottleneck for deploying diffusion policies: they work well in the average case but violate the hard physical constraints that matter for safety. By framing safety as a demonstration-free, reward-free post-training projection, PACT can be applied to off-the-shelf pretrained policies, making it a lightweight and broadly reusable safety layer. The "self-evolving" curriculum design — gradually tightening constraints while bounding policy drift — is what lets it improve safety and success at the same time rather than trading one for the other.

Links

← Back to ICML-2026

⚠️ **GitHub.com Fallback** ⚠️