ICLR 2026 Guided Flow Policy - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 · OpenReview: EBjy1rmpv0 Category: RL for VLA — Offline RL with flow-matching policies Trend tag: Flow-matching policies / value-weighted BC
flowchart LR
Data[Offline dataset] --> FM[Multi-step flow policy<br/>VaBC π_ω]
FM -- distill --> Actor[Distilled<br/>one-step actor π_θ]
Actor -- weighted BC<br/>direct toward high-value --> FM
FM -. regularize / constrain .-> Actor
Crit[Critic] --> Actor
Standard offline-RL behavior regularization treats every dataset action equally, so the policy clones low-value transitions just as much as high-value ones. Flow-matching policies model the dataset distribution well but, on their own, inherit this unselective imitation.
GFP is a dual-policy BRAC (behavior-regularized actor-critic) framework that couples a multi-step flow-matching policy — termed Value-aware Behavior Cloning (VaBC, the policy
Reports state-of-the-art performance across 144 state- and pixel-based tasks spanning OGBench (105 tasks, 100 state + 5 pixel), D4RL (18, AntMaze/Adroit), and Minari (21, Adroit/Gym-MuJoCo), with substantial gains specifically on suboptimal datasets and challenging tasks. On the OGBench average GFP scores 53.2 vs. FQL 46.7, ReBRAC 43.9, and IQL 20.4; the gap widens on noisy tasks (e.g. cube-double-noisy GFP ≈63 vs. FQL ≈38). Ablations show the guidance temperature
Shows that the flow-matching policy class — increasingly common in VLAs — is a strong fit for offline RL when paired with value-weighted distillation. Complements on-policy Flow Matching Policy Gradients on the offline side, and is a candidate plug-in for offline post-training of flow / diffusion VLAs like Unified Diffusion VLA.
- OpenReview: https://openreview.net/forum?id=EBjy1rmpv0
- arXiv: https://arxiv.org/abs/2512.03973
Authors: Franki Nguimatsia Tiofack, Théotime Le Hellard, Fabian Schramm (Inria / DI-ENS, PSL), Nicolas Perrin-Gilbert (Sorbonne Université), Justin Carpentier.