NeurIPS 2025 ReinFlow - Heungwoo/research GitHub Wiki

ReinFlow โ€” Online RL Fine-Tuning for Flow-Matching Policies

Venue: NeurIPS 2025 ยท arXiv: 2505.22094 Category: RL / Flow-Matching Policies

Approach diagram

flowchart LR
  Base[Pretrained flow policy:<br/>Rectified Flow / Shortcut Models] --> Noise[Inject learnable noise<br/>at each flow step]
  Noise --> Markov[Discrete-time Markov process<br/>with EXACT likelihoods]
  Markov --> RL[Online RL fine-tune<br/>PPO-style policy gradient<br/>with exact log-probs]
  RL --> Better[+135% reward locomotion<br/>+40% manipulation<br/>at 1โ€“4 denoising steps]
Loading

Problem

Flow-matching policies are deterministic by construction โ€” the ODE flow path doesn't expose action log-probabilities. That makes the standard RL toolkit (PPO, policy gradient) non-trivial to apply. Prior workarounds either change the architecture (discrete-diffusion VLAs) or use log-prob-free signals (RECAP's outcome conditioning, DSRL's noise-latent). Neither is a drop-in for flow matching.

Method

Inject learnable noise into the deterministic flow path so that each flow integration step becomes a probabilistic Markov transition with exact likelihoods. This converts the flow into a discrete-time Markov process whose per-step log-probs can be computed exactly, so standard policy-gradient RL (a PPO-style objective) can optimize it directly.

Crucially, the noise injection is:

  • Learnable (a state-dependent noise network, not hand-designed)
  • Trainable via RL alongside the policy
  • Drop-in across flow-model variants โ€” the paper fine-tunes Rectified Flow (with various time distributions) and Shortcut Models, at as few as one denoising step.

Scope note: the paper experiments are on Rectified Flow and Shortcut Model policies in simulation. The open-source repo later extended ReinFlow to VLAs (ฯ€0, ฯ€0.5, GR00T-N1.5), but those are not part of the published results.

Results

Benchmarks: OpenAI Gym locomotion (Hopper/Walker2d/Ant/Humanoid), Franka Kitchen (state-based manipulation), and Robomimic (visual manipulation).

  • +135.36% average net reward growth on Rectified Flow locomotion policies vs. baseline, at 82.63% of DPPO's wall-clock time.
  • +40.34% average net success-rate increase on Shortcut Model state/visual manipulation policies, with 23.20% compute-time savings.
  • Effective at very few (4) or even one denoising step.
  • Uses a PPO-style policy-gradient objective; baselines compared against include DPPO, FQL, and a Gaussian policy. (The wiki previously claimed "PPO and GRPO" โ€” GRPO does not appear in the paper.)

Significance

Solves a specific technical blocker: how do you run standard RL on a deterministic flow policy? ReinFlow is the mechanism-complete answer, orthogonal to RECAP's outcome conditioning and the ICLR 2026 residual-RL stack (PLD, RL Tokens, RFS).

Extends DSRL's noise-latent-RL thesis from diffusion policies to flow-matching policies. Authors: Tonghe Zhang (CMU), Chao Yu (Tsinghua), Sichang Su, Yu Wang (Tsinghua).

Links

Related pages

โ† Back to NeurIPS-2025

โš ๏ธ **GitHub.com Fallback** โš ๏ธ