ICLR 2026 Flow Matching PG - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 · OpenReview: eoEmoKoQpJ Category: RL for VLA — Policy parameterization Trend tag: Flow-matching policies / on-policy RL
flowchart LR
Pol[Flow-matching policy v̂_θ<br/>denoising MLP / VLA head] --> Roll[Rollouts: any sampler<br/>10-step Euler / stochastic / det.]
Roll --> Adv[GAE advantage Â_t]
Roll --> MC["Store N_mc<br/>(τ_i, ε_i) pairs per (o_t,a_t)"]
Adv --> Ratio["FPO ratio:<br/>r̂ = exp(L_CFM,θ_old − L_CFM,θ)"]
MC --> Ratio
Ratio --> Clip["PPO-clip surrogate<br/>min(r̂Â, clip(r̂)Â)"]
Clip --> Pol
Standard policy-gradient methods (PPO) parameterise actions as diagonal Gaussians, which struggle with multimodal action distributions and under-conditioned tasks (e.g. humanoid root-only goal conditioning). Flow / diffusion policies capture multimodality well but lack tractable likelihoods, making them awkward to plug into the PPO-clip ratio. Existing on-policy RL methods for diffusion (DDPO, DPPO, Flow-GRPO) treat each denoising step as its own MDP transition, which (i) multiplies horizon length by 10–50, (ii) treats initial noise as part of observation, increasing problem dimensionality, and (iii) restricts to stochastic samplers.
FPO is a drop-in replacement for the PPO likelihood ratio that uses the conditional flow matching (CFM) loss as a likelihood proxy.
Given clean action a_t and noise ε ∼ N(0,I), interpolate at flow timestep τ ∈ [0,1]:
a_t^τ = (1-τ) a_t + τ ε (OT schedule)
Train velocity field v̂_θ to predict the conditional flow u(a_t^τ, τ | a_t) = a_t − ε:
L_CFM,θ = E_{τ, q(a), p_τ(a^τ|a)} [ ‖ v̂_θ(a^τ, τ) − (a − ε) ‖² ]
Replace the PPO log-likelihood ratio with a CFM-loss-based proxy:
r̂_FPO(θ) = exp( L̂_CFM,θ_old(a_t; o_t) − L̂_CFM,θ(a_t; o_t) )
where each L̂ is a Monte-Carlo estimate over N_mc draws of (τ_i, ε_i):
L̂_CFM,θ(a_t; o_t) = (1/N_mc) Σ_i ‖ v̂_θ(a_t^{τ_i}, τ_i; o_t) − (a_t − ε_i) ‖²
The (τ_i, ε_i) pairs are frozen between the θ_old and θ evaluations so that the difference cancels noise.
Plug r̂_FPO into PPO-clip:
max_θ E_{a_t∼π_old} [ min( r̂_FPO Â_t, clip(r̂_FPO, 1−ε_clip, 1+ε_clip) Â_t ) ]
Kingma & Gao (2023) show the diffusion-loss-with-constant-weight equals -ELBO + const; thus
r_FPO(θ) = exp(ELBO_θ - ELBO_θ_old) = (π_θ/π_θ_old) · exp(D_KL_θ_old − D_KL_θ)
The first factor is the true likelihood ratio; the second factor tightens the variational bound. Maximising the FPO ratio simultaneously increases the modelled likelihood of high-advantage actions and tightens the ELBO.
The Monte-Carlo single-sample estimator r̂_FPO^(τ,ε) is an upward biased estimate of r_FPO (Jensen). However the gradient is unbiased (Eq. 19-20):
∇_θ r̂_FPO = − r̂_FPO · ∇_θ ℓ_θ(τ, ε) E[ −∇_θ ℓ_θ ] = ∇_θ ELBO_θ
So even with N_mc = 1, FPO produces directionally unbiased gradients. Empirically, more samples help but a single sample already beats Gaussian PPO.
While not converged:
Collect rollouts using any sampler; compute Â_t with GAE
For each action store N_mc (τ_i, ε_i) pairs and ℓ_θ(τ_i, ε_i)
θ_old ← θ
For each epoch:
For each minibatch (o_t, a_t, {(τ_i,ε_i)}):
Compute ℓ_θ(τ_i, ε_i)
r̂_θ = exp( -(1/N_mc) Σ_i (ℓ_θ - ℓ_θ_old) )
L_FPO = min(r̂Â, clip(r̂, 1±ε)Â)
θ ← Optimizer(θ, ∇L_FPO)
Update value head as in standard PPO
- Optimiser: Adam.
- 60 M total environment steps, batch size 1024, 16 updates per batch.
- 10 sampling steps for FPO and DPPO.
- Learning rate: 3 × 10⁻⁴ for FPO and DPPO.
- Clip ε swept ∈ {0.01, 0.05, 0.1, 0.2, 0.3}; final ε = 0.05 for FPO.
- For DPPO: per-step Gaussian noise σ_t swept ∈ {0.01, 0.05, 0.1}; σ_t = 0.05, ε = 0.2 used.
- N_mc = 8 in main experiments; ablations at 1 and 4.
- ε-CFM (compute CFM loss on noise prediction ε̂) used over u-CFM (CFM on velocity) for scale invariance.
Table 1 reports the average evaluation reward across MuJoCo tasks:
| Method | Avg Reward |
|---|---|
| Gaussian PPO | 667.8 ± 66.0 |
| Gaussian PPO† (default HPs) | 577.2 ± 74.4 |
| DPPO | 652.5 ± 83.7 |
| FPO‡ | 759.3 ± 45.3 |
| FPO, 1 (τ,ε) | 691.6 ± 50.3 |
| FPO, 4 (τ,ε) | 731.2 ± 58.2 |
| FPO, u-MSE (velocity) | 664.6 ± 48.5 |
| FPO, ε_clip = 0.1 | 623.3 ± 76.3 |
| FPO, ε_clip = 0.2 | 526.4 ± 76.8 |
‡ = 8 (τ,ε) pairs, ε-MSE, ε_clip = 0.05.
FPO wins on 8 of 10 Playground tasks (visualised in Figures 2 and 3 against Gaussian PPO and DPPO over BallInCup, FingerSpin, FingerTurnEasy/Hard, FishSwim, PointMass, ReacherEasy/Hard, CartpoleBalance, CheetahRun).
Goal-conditioned MoCap tracking on AMASS, evaluated by success rate (joint distance ≤ 0.5 m), alive duration, and global MPJPE:
| Method | Goal | Success ↑ | Alive ↑ | MPJPE ↓ |
|---|---|---|---|---|
| Gaussian PPO | All joints | 98.7 % | 200.46 | 31.62 |
| FPO | All joints | 96.4 % | 198.00 | 41.98 |
| Gaussian PPO | Root + Hands | 46.5 % | 142.50 | 97.65 |
| FPO | Root + Hands | 70.6 % | 171.32 | 62.91 |
| Gaussian PPO | Root only | 29.8 % | 114.06 | 123.70 |
| FPO | Root only | 54.3 % | 152.90 | 73.55 |
When sufficient conditioning is provided (all joints), Gaussian PPO is on par. As the conditioning becomes sparser (root or root+hands only), Gaussian PPO collapses while FPO retains 54–71 % success — a ~24-point gain. Prior methods solving sparse-goal humanoid control require a teacher → student distillation pipeline; FPO learns end-to-end.
Beyond MoCap tracking, the authors also show FPO trains a humanoid that walks across procedurally generated rough terrain (Figure 4c).
Designed to probe multimodality. FPO learns a bimodal action distribution at saddle-point states (where multiple optima exist), visible in the denoising flow visualisation (Figure 1). Gaussian PPO consistently picks the nearest goal with low diversity.
- N_mc sweep (Table 1). N_mc = 1 → 691.6, N_mc = 4 → 731.2, N_mc = 8 → 759.3. More samples help but FPO with a single sample still beats Gaussian PPO (667.8) and DPPO (652.5).
- ε-CFM vs u-CFM. ε-MSE (predict noise) gives 759.3; u-MSE (predict velocity) drops to 664.6. The authors hypothesise ε is invariant to action scale, which makes a single ε_clip transferable across tasks.
- Clipping ε. ε = 0.05 best (759.3); ε = 0.1 → 623.3; ε = 0.2 → 526.4. Tight clipping critical, similar to but tighter than Gaussian PPO.
- DPPO denoising-MDP comparison (Sec. 3.5). DPPO multiplies horizon by 10–50 (one PPO step per denoising step) and is restricted to stochastic samplers; FPO is sampler-agnostic at both train and test.
- Goal-conditioning sweep on humanoid. Full → root+hands → root reveals widening FPO advantage as conditioning sparsens.
The authors do not include a separate "Limitations" section, but the Discussion surfaces:
- Compute cost. "The training and deployment of flow-based policies is generally more computationally intensive than for corresponding Gaussian policies" (explicit Discussion statement).
- Missing PPO machinery. FPO "lacks established machinery such as KL divergence estimation for adaptive learning rates and entropy regularization" — features standard PPO tooling provides for Gaussian policies.
- No real-robot deployment — all experiments are in simulation (MuJoCo Playground, Isaac Gym, GridWorld). The rough-terrain locomotion result is described as showing "potential for sim-to-real transfer," but no hardware results are reported.
- Image-diffusion fine-tuning unstable. The authors explored applying FPO to fine-tune a pretrained image diffusion model with RL and "found this setting to be unstable in practice," attributing the instability to the broader difficulty of RL on image generation rather than to FPO itself. (The paper does not evaluate or name large diffusion VLAs such as π0 or GR00T-N1.)
- MC-bias trade-off. Single-sample r̂ is upward-biased; gradients are directionally unbiased but variance is higher. The authors use N_mc = 8 as the main-experiment default.
The paper's stated future direction is applying FPO where flow-based policies are already pretrained — e.g. behavior-cloned diffusion policies in robotics — where its PPO-clip compatibility and simplicity may help fine-tuning with task reward.
FPO is the simplest known on-policy RL recipe for flow / diffusion policies that does not require approximating likelihoods or unrolling the sampler as an MDP. Three properties make it especially valuable for the VLA agenda:
- Sampler-agnostic. The same flow actor can be trained with any deterministic or stochastic integrator, and inference can use any number of steps. This decouples training compute from inference latency — important for VLAs that may want fast 1-step inference but stable many-step training.
- PPO-clip drop-in. Existing PPO codebases need only swap the likelihood ratio. The authors implemented FPO on top of Brax PPO with minimal changes.
- Under-conditioning gain. The 24-point success-rate gap on root-only humanoid control is the most direct evidence yet that flow policies' expressivity matters for under-specified tasks — exactly the regime where VLAs often operate (sparse language goal, multimodal answer).
Within the 2026 RL-for-VLA cluster:
- VLA-RFT also targets RL fine-tuning of flow VLAs but inside a learned world model; FPO is model-free.
- SimpleVLA-RL offers a Gaussian PPO recipe for VLA RL fine-tuning; FPO is the flow-native counterpart.
- Guided Flow Policy is the offline-RL counterpart that adds Q-guidance to flow training; FPO complements it with online policy gradients.
- Flow-To-Policy uses flows as planners rather than policies; FPO targets policies directly.
- ReinFlow (NeurIPS 2025) explored a similar territory but with Gaussian-step formulation; FPO replaces the Gaussian-step trick with a CFM-loss ratio.
- Compared to flow-VLA architectures themselves (Unified Diffusion VLA, dVLA, π0 family), FPO provides the missing post-training tool: online RL without giving up the flow generative structure.
The paper is also one of the cleanest theoretical contributions of the cycle — the equivalence between FPO ratio and ELBO-ratio (Eq. 12) provides a justification that previous diffusion-RL recipes lacked.
- OpenReview: https://openreview.net/forum?id=eoEmoKoQpJ
- Guided Flow Policy — offline-RL counterpart
- Flow-To-Policy — flow-as-planner alternative
- VLA-RFT — RL fine-tuning in world model
- SimpleVLA-RL
- ReinFlow
- Unified Diffusion VLA
- dVLA