ICRA 2026 RFT Flow - Heungwoo/research GitHub Wiki
Venue: ICRA 2026 · Authors: Mingyang Lyu, Yinqian Sun, Erliang Lin, Huangrui Li, Ruolin Chen, Feifei Zhao, Yi Zeng · arXiv: 2510.09976 Category: RL for VLA Trend tag: RL fine-tuning of flow-matching action experts
flowchart LR
Base[π₀ flow-matching VLA<br/>imitation prior] --> Roll[Online rollouts<br/>sparse rewards]
Roll --> Ratio[Likelihood-free policy ratio:<br/>per-sample change in<br/>conditional flow-matching loss<br/>Δℓ_cfm = ℓ_old − ℓ_θ]
Ratio --> PPO[Clipped surrogate<br/>PPO-style objective]
PPO --> Comp[+ structure-aware credit assignment<br/>+ multi-step latent exploration<br/>+ Q-ensemble value estimation]
Comp --> Better[Fine-tuned π₀<br/>above imitation prior]
VLA models (OpenVLA, Octo, π₀) are bounded by the quality and coverage of their supervised demonstrations. RL via online interaction is the natural way to push past the imitation prior — but conventional policy-gradient RL is computationally infeasible for flow-matching policies. PPO-style updates need an importance-sampling ratio, which requires explicit policy likelihoods; the flow-matching ODE does not expose tractable action log-probabilities, so the policy ratio cannot be computed directly.
The paper proposes Flow Policy Optimization (FPO), which sidesteps explicit likelihoods entirely. Instead of computing log-probs, FPO builds a likelihood-free policy ratio from per-sample changes in the conditional flow-matching (CFM) objective: for each stored sample it tracks the CFM-loss differential between the old and current policy (Δℓ_cfm = ℓ_cfm(·;θ_old) − ℓ_cfm(·;θ)) and uses it as a structure-aligned surrogate for the importance ratio inside a PPO-style clipped objective. On top of this core reformulation, FPO integrates four components:
- Structure-aware credit assignment to improve gradient efficiency;
- Clipped surrogate objective to stabilize optimization;
- Multi-step latent exploration to diversify policy updates;
- Q-ensemble for robust value estimation.
The base policy fine-tuned is π₀, and learning is stable under sparse rewards.
Evaluated on the LIBERO benchmark and an ALOHA simulation task (Transfer Cube) against supervised, preference-aligned, diffusion-based, autoregressive online-RL, and π₀-FAST baselines, FPO shows consistent gains over the imitation prior and strong alternatives, with stable convergence under sparse reward. Reported LIBERO success rates: Spatial 97.2 · Object 97.3 · Goal 89.4 · Long 65.3 · average 87.2; on ALOHA Transfer Cube the learning curve reaches roughly 65%. Ablations and latent-space analyses validate the individual modules and the stable convergence of the CFM objective during online RL.
FPO offers a drop-in online-RL recipe for flow-matching VLAs that needs no architectural change — no added noise, no SDE reformulation — by reusing the CFM loss the policy is already trained with as the source of the policy ratio. This makes RFT directly applicable to off-the-shelf flow-matching action experts like π₀.
- vs. ReinFlow: ReinFlow makes the deterministic flow stochastic by injecting learnable noise at each integration step, turning the ODE into a discrete-time Markov chain with exact per-step log-probabilities, then runs standard PPO on those exact log-probs. FPO instead keeps the flow as-is and is likelihood-free: it never computes log-probs, substituting the CFM-loss differential as a ratio proxy. ReinFlow changes the sampler to get an exact likelihood; FPO leaves the sampler alone and approximates the ratio.
- vs. the ODE→SDE log-prob trick in Qwen-VLA review: Qwen-VLA's stage-IV PPO obtains an analytic flow-matching log-probability via an ODE→SDE conversion, then optimizes that exact likelihood. FPO again differs by avoiding any likelihood estimate: rather than converting the ODE to an SDE to recover a tractable density, it derives its ratio from changes in the CFM training objective.
In short, all three target the same blocker (flow policies lack a usable action likelihood for policy-gradient RL), but ReinFlow and Qwen-VLA recover an exact log-prob (via noise injection / ODE→SDE), whereas FPO is likelihood-free, trading exactness for not having to touch the flow sampler. See RL for VLA for the broader landscape.
- arXiv abstract: arxiv.org/abs/2510.09976
- arXiv HTML: arxiv.org/html/2510.09976v1
- ReinFlow — learnable noise injection for flow RL
- Qwen-VLA review — ODE→SDE flow-matching log-prob for PPO
- RL for VLA — RL-for-VLA topic landing
- ICRA 2026 Survey
← Back to ICRA-2026