CoRL 2025 DSRL - Heungwoo/research GitHub Wiki

DSRL โ€” Steering Your Diffusion Policy with Latent Space Reinforcement Learning

Venue: CoRL 2025 (Award Finalist) ยท arXiv: 2506.15799 Authors: Andrew Wagenmaker, Mitsuhiko Nakamoto, Yunchu Zhang, Seohong Park, Waleed Yagoub, Anusha Nagabandi, Abhishek Gupta, Sergey Levine (UC Berkeley ยท University of Washington ยท Amazon) Category: RL for VLA / Diffusion Policies Trend tag: RL without log-probs

Approach diagram

flowchart LR
  Z[Initial noise z0] --> FROZEN[Frozen diffusion policy]
  FROZEN --> A[Action chunk]
  A --> R[Environment reward]
  R --> LEARN[Train a policy over z0<br/>RL action space = noise latent]
  LEARN --> Z
Loading

Problem

Diffusion / flow-matching policies don't expose action log-probabilities, so the standard RL toolkit (policy gradient, PPO) can't be applied out of the box. Prior workarounds either finetune the whole model (unstable) or approximate log-probs (noisy).

Method

DSRL (Diffusion Steering via Reinforcement Learning) makes a structural observation: the initial noise fed into a diffusion policy largely determines the sampled trajectory. Therefore you can freeze the diffusion policy and train a small RL actor that chooses the initial noise zโ‚€ given an observation. The noise latent becomes the RL action space โ€” with real log-probs and cheap rollouts. Concretely, DSRL trains a lightweight off-policy actor-critic (SAC) over the latent-noise input, requires only black-box access to the BC policy (it never touches the base weights or gradients), and adds only a small set of trainable parameters (~500K).

Results

Post-hoc improvement of frozen diffusion policies across simulated benchmarks (OpenAI Gym, Robomimic, OGBench), single-task and multi-task (BridgeData V2) diffusion policies, and the pretrained generalist ฯ€โ‚€ policy (steered from its public DROID weights). DSRL reports substantially higher sample efficiency than direct-finetuning RL baselines (e.g., RLPD), and demonstrates real-world autonomous improvement on a physical Franka setup. CoRL 2025 Award Finalist.

Significance

The seed idea for the ICLR 2026 wave of RL-for-flow-matching papers. Every approach on the ICLR 2026 RL topic page can trace an ancestor to DSRL or its contemporaries:

Links

Related pages

โ† Back to CoRL-2025

โš ๏ธ **GitHub.com Fallback** โš ๏ธ