RSS 2026 Set Supervised Diffusion Policy - Heungwoo/research GitHub Wiki

Set-Supervised Diffusion Policy: Learning Action-Chunking Diffusion through Corrections

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Imitation learning 1 · paper #80 Authors: Zhaoting Li, Gang Chen, Javier Alonso-Mora, Cosimo Della Santina, Jens Kober arXiv: 2606.01865 · program page

Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

SDP framework (Figure 1 of arXiv 2606.01865, © the authors)

Figure 1: (A) contrastive positive/negative action-chunk pairs are generated from offline demonstrations (by synthesizing negatives) or online human interventions (teacher action = positive, queried unexecuted robot action = negative). (B) Standard DP imitates only the positive chunks; SDP instead constructs a desired action set (black contour) from each pair and trains the diffusion policy to imitate samples drawn from that set.

Problem

Diffusion policies inherit behavior cloning's distributional-shift problem, so deployments lean on human-in-the-loop corrections — but standard BC losses and DAgger-style aggregation only imitate the teacher's positive actions, discarding the negative signal in the robot's undesired actions. This overfits to (possibly noisy) teacher labels and increases reliance on costly expert interventions. CLIC's set-valued supervision exploits both, but uses hard-to-train EBMs and single-step actions, incompatible with action chunking.

Method

SDP extends CLIC's desired action set to action-chunks: each contrastive pair (A−, A+) defines per-timestep balls of radius r·D(a+, a−) around the positive actions, whose product forms a set-valued target. Because diffusion models lack explicit likelihoods, SDP trains by sampling desired action-chunks from the set — a truncated denoising process (from intermediate step K_A) with a reflection operator after every denoising step that reverts out-of-set actions to the positive action — then imitating those samples with the usual diffusion loss. The same recipe supports online interactive imitation learning (interventions logged as paired chunks via a sliding window) and offline learning (negatives synthesized for demonstration data). Ablations show reflection-based sampling clearly beats classifier-guidance variants (0.988/0.946 vs 0.890/0.448 accurate/noisy).

Results

On four simulated tasks (Push-T plus robosuite Square, PickCan, TwoArmLift), interactive SDP averages 0.918 success with accurate feedback vs DP 0.878, DP-DPO 0.691, ADP 0.832, CLIC 0.772, IBC 0.559; with noisy interventions the gap widens (SDP 0.865 vs DP 0.711). SDP also yields better aggregated datasets — policies trained offline on SDP-collected data beat those trained on DP-collected data — and outperforms DP on Robomimic. Real-world: on Insert-T (KUKA iiwa), demonstrations+corrections give SDP 35/40 (hard) and 39/40 (medium) vs DP's 23/40 and 35/40; on FurnitureBench round-table assembly (Franka Panda, 30 demos + 29,194 correction pairs), SDP shows roughly 10% higher success with stage-wise gains concentrated in the contact-rich Tighten Parts stage.

Significance

Makes the negative half of intervention data a first-class training signal for action-chunking diffusion policies, with an RL-like ability to exceed the quality of its labels — directly relevant to the data-quality and human-in-the-loop threads in Review-LBM-Cotraining.

← Back to RSS 2026 survey · RSS-2026-Papers · Home