NeurIPS 2025 ReinFlow - Heungwoo/research GitHub Wiki
Venue: NeurIPS 2025 ยท arXiv: 2505.22094 Category: RL / Flow-Matching Policies
flowchart LR
Base[Pretrained flow policy:<br/>Rectified Flow / Shortcut Models] --> Noise[Inject learnable noise<br/>at each flow step]
Noise --> Markov[Discrete-time Markov process<br/>with EXACT likelihoods]
Markov --> RL[Online RL fine-tune<br/>PPO-style policy gradient<br/>with exact log-probs]
RL --> Better[+135% reward locomotion<br/>+40% manipulation<br/>at 1โ4 denoising steps]
Flow-matching policies are deterministic by construction โ the ODE flow path doesn't expose action log-probabilities. That makes the standard RL toolkit (PPO, policy gradient) non-trivial to apply. Prior workarounds either change the architecture (discrete-diffusion VLAs) or use log-prob-free signals (RECAP's outcome conditioning, DSRL's noise-latent). Neither is a drop-in for flow matching.
Inject learnable noise into the deterministic flow path so that each flow integration step becomes a probabilistic Markov transition with exact likelihoods. This converts the flow into a discrete-time Markov process whose per-step log-probs can be computed exactly, so standard policy-gradient RL (a PPO-style objective) can optimize it directly.
Crucially, the noise injection is:
- Learnable (a state-dependent noise network, not hand-designed)
- Trainable via RL alongside the policy
- Drop-in across flow-model variants โ the paper fine-tunes Rectified Flow (with various time distributions) and Shortcut Models, at as few as one denoising step.
Scope note: the paper experiments are on Rectified Flow and Shortcut Model policies in simulation. The open-source repo later extended ReinFlow to VLAs (ฯ0, ฯ0.5, GR00T-N1.5), but those are not part of the published results.
Benchmarks: OpenAI Gym locomotion (Hopper/Walker2d/Ant/Humanoid), Franka Kitchen (state-based manipulation), and Robomimic (visual manipulation).
- +135.36% average net reward growth on Rectified Flow locomotion policies vs. baseline, at 82.63% of DPPO's wall-clock time.
- +40.34% average net success-rate increase on Shortcut Model state/visual manipulation policies, with 23.20% compute-time savings.
- Effective at very few (4) or even one denoising step.
- Uses a PPO-style policy-gradient objective; baselines compared against include DPPO, FQL, and a Gaussian policy. (The wiki previously claimed "PPO and GRPO" โ GRPO does not appear in the paper.)
Solves a specific technical blocker: how do you run standard RL on a deterministic flow policy? ReinFlow is the mechanism-complete answer, orthogonal to RECAP's outcome conditioning and the ICLR 2026 residual-RL stack (PLD, RL Tokens, RFS).
Extends DSRL's noise-latent-RL thesis from diffusion policies to flow-matching policies. Authors: Tonghe Zhang (CMU), Chao Yu (Tsinghua), Sichang Su, Yu Wang (Tsinghua).
- arXiv: https://arxiv.org/abs/2505.22094
- NeurIPS virtual page: https://neurips.cc/virtual/2025/loc/san-diego/poster/119473
- GitHub: https://github.com/ReinFlow/ReinFlow
- DSRL (CoRL 2025 noise-latent-RL ancestor)
- RECAP ยท RL Tokens ยท VLA-RFT ยท SimpleVLA-RL (ICLR 2026 RL-for-VLA siblings)
- RL for VLA
โ Back to NeurIPS-2025