ICRA 2026 RFT Flow - Heungwoo/research GitHub Wiki

RFT for Flow-Matching VLA — RL Fine-Tuning of Flow Policies

Venue: ICRA 2026 · Authors: Mingyang Lyu, Yinqian Sun, Erliang Lin, Huangrui Li, Ruolin Chen, Feifei Zhao, Yi Zeng · arXiv: 2510.09976 Category: RL for VLA Trend tag: RL fine-tuning of flow-matching action experts

Approach diagram

flowchart LR
  Base[π₀ flow-matching VLA<br/>imitation prior] --> Roll[Online rollouts<br/>sparse rewards]
  Roll --> Ratio[Likelihood-free policy ratio:<br/>per-sample change in<br/>conditional flow-matching loss<br/>Δℓ_cfm = ℓ_old − ℓ_θ]
  Ratio --> PPO[Clipped surrogate<br/>PPO-style objective]
  PPO --> Comp[+ structure-aware credit assignment<br/>+ multi-step latent exploration<br/>+ Q-ensemble value estimation]
  Comp --> Better[Fine-tuned π₀<br/>above imitation prior]
Loading

Problem

VLA models (OpenVLA, Octo, π₀) are bounded by the quality and coverage of their supervised demonstrations. RL via online interaction is the natural way to push past the imitation prior — but conventional policy-gradient RL is computationally infeasible for flow-matching policies. PPO-style updates need an importance-sampling ratio, which requires explicit policy likelihoods; the flow-matching ODE does not expose tractable action log-probabilities, so the policy ratio cannot be computed directly.

Method

The paper proposes Flow Policy Optimization (FPO), which sidesteps explicit likelihoods entirely. Instead of computing log-probs, FPO builds a likelihood-free policy ratio from per-sample changes in the conditional flow-matching (CFM) objective: for each stored sample it tracks the CFM-loss differential between the old and current policy (Δℓ_cfm = ℓ_cfm(·;θ_old) − ℓ_cfm(·;θ)) and uses it as a structure-aligned surrogate for the importance ratio inside a PPO-style clipped objective. On top of this core reformulation, FPO integrates four components:

  • Structure-aware credit assignment to improve gradient efficiency;
  • Clipped surrogate objective to stabilize optimization;
  • Multi-step latent exploration to diversify policy updates;
  • Q-ensemble for robust value estimation.

The base policy fine-tuned is π₀, and learning is stable under sparse rewards.

Results

Evaluated on the LIBERO benchmark and an ALOHA simulation task (Transfer Cube) against supervised, preference-aligned, diffusion-based, autoregressive online-RL, and π₀-FAST baselines, FPO shows consistent gains over the imitation prior and strong alternatives, with stable convergence under sparse reward. Reported LIBERO success rates: Spatial 97.2 · Object 97.3 · Goal 89.4 · Long 65.3 · average 87.2; on ALOHA Transfer Cube the learning curve reaches roughly 65%. Ablations and latent-space analyses validate the individual modules and the stable convergence of the CFM objective during online RL.

Significance

FPO offers a drop-in online-RL recipe for flow-matching VLAs that needs no architectural change — no added noise, no SDE reformulation — by reusing the CFM loss the policy is already trained with as the source of the policy ratio. This makes RFT directly applicable to off-the-shelf flow-matching action experts like π₀.

How it relates to prior flow-RL work

  • vs. ReinFlow: ReinFlow makes the deterministic flow stochastic by injecting learnable noise at each integration step, turning the ODE into a discrete-time Markov chain with exact per-step log-probabilities, then runs standard PPO on those exact log-probs. FPO instead keeps the flow as-is and is likelihood-free: it never computes log-probs, substituting the CFM-loss differential as a ratio proxy. ReinFlow changes the sampler to get an exact likelihood; FPO leaves the sampler alone and approximates the ratio.
  • vs. the ODE→SDE log-prob trick in Qwen-VLA review: Qwen-VLA's stage-IV PPO obtains an analytic flow-matching log-probability via an ODE→SDE conversion, then optimizes that exact likelihood. FPO again differs by avoiding any likelihood estimate: rather than converting the ODE to an SDE to recover a tractable density, it derives its ratio from changes in the CFM training objective.

In short, all three target the same blocker (flow policies lack a usable action likelihood for policy-gradient RL), but ReinFlow and Qwen-VLA recover an exact log-prob (via noise injection / ODE→SDE), whereas FPO is likelihood-free, trading exactness for not having to touch the flow sampler. See RL for VLA for the broader landscape.

Links

Related pages

← Back to ICRA-2026

⚠️ **GitHub.com Fallback** ⚠️