ICML 2026 Sample from What You See - Heungwoo/research GitHub Wiki

Sample from What You See โ€” BridgePolicy: diffusion-bridge visuomotor control that samples from observations, not noise

Venue: ICML 2026 (Poster) Category: Diffusion-Flow Policy Traction (2026-06): 0 citations (arXiv)

Comparison of Diffusion Policy and BridgePolicy: BridgePolicy's observation modeling lets sampling start from a rich, meaningful prior instead of random noise (Figure 1 from the BridgePolicy authors, 2026)

Problem

Diffusion-model imitation learning captures multi-modal action distributions, but existing visuomotor policies treat observations merely as high-level conditioning inputs to the denoising network rather than embedding them into the stochastic dynamics of the diffusion process itself. Consequently sampling must begin from random Gaussian noise, weakening the coupling between perception and control and often yielding suboptimal performance. The paper asks: what if generation started from the observation rather than from noise?

Method

Overview pipeline: BridgePolicy embeds observations (robot states + point cloud) into the diffusion SDE trajectory via a diffusion-bridge formulation, with a multi-modal fusion module and a semantic aligner; inference samples from the observation and iteratively transforms it into the action (Figure 2 from the BridgePolicy authors, 2026)

BridgePolicy explicitly embeds observations within the stochastic differential equation (SDE) via a diffusion-bridge formulation, constructing an observation-informed trajectory so that sampling starts from a rich, informative prior instead of random noise. The observation consists of robot states and a point cloud.

A key challenge is that classical diffusion bridges connect distributions of matched dimensionality, whereas robotic observations are heterogeneous, multi-modal, and do not naturally align with the action space. BridgePolicy addresses this with two designs:

  • A multi-modal fusion module that unifies visual (point cloud) and state inputs.
  • A semantic aligner that aligns observation and action representations, trained with a contrastive CLIP loss (symmetric L_clip(a, z_obs) + L_clip(z_obs, a)) to bring the observation and action distributions into semantic proximity, making the bridge applicable to heterogeneous robot data.

The overall training objective combines the diffusion-bridge loss with the alignment loss, L = L_DB + ฮฑยทL_align. At inference, fused observations form the latent starting point, which a designed solver iteratively updates via fast sampling to produce actions.

Results

Experiments span 52 tasks across three simulation benchmarks โ€” Adroit, DexArt, and MetaWorld โ€” plus five real-world tasks. Expert demos were collected via scripted policy (MetaWorld) and RL (VRL3 for Adroit, PPO for DexArt), with 10 episodes per Adroit/MetaWorld task and 100 per DexArt task.

Simulation (Table 1, average success rate):

Method MW-Easy MW-Med MW-Hard MW-VeryHard DexArt Adroit Avg
DP 0.79 0.31 0.10 0.26 0.45 0.31 0.37
DP3 0.87 0.61 0.40 0.51 0.57 0.68 0.60
Simple DP3 0.86 0.59 0.38 0.47 0.48 0.68 0.58
FlowPolicy 0.86 0.67 0.59 0.76 0.54 0.70 0.68
BridgePolicy 0.91 0.75 0.58 0.79 0.60 0.81 0.74

Real-world (Table 2, success rate): BridgePolicy averaged 0.90 across Oven-Closing/Opening, Pick-Place, Pour, and Unplug โ€” versus DP3 0.76, Simple DP3 0.66, and FlowPolicy 0.56 โ€” winning or tying every task (e.g., 1.0 on both oven tasks, 0.9 on Unplug).

Significance

BridgePolicy shows that injecting observations into the diffusion SDE itself โ€” rather than as side conditioning โ€” and starting sampling from an observation-conditioned prior consistently outperforms strong 2D and 3D generative baselines (DP, DP3, FlowPolicy) across both simulation and real hardware. The fusion module and CLIP-style aligner make diffusion bridges practical for heterogeneous, dimension-mismatched robot observations, a previously limiting assumption of bridge methods.

Links

โ† Back to ICML-2026