ICLR 2026 Guided Flow Policy - Heungwoo/research GitHub Wiki

Guided Flow Policy — Offline RL from High-Value Actions

Venue: ICLR 2026 · OpenReview: EBjy1rmpv0 Category: RL for VLA — Offline RL with flow-matching policies Trend tag: Flow-matching policies / value-weighted BC

Approach diagram

flowchart LR
  Data[Offline dataset] --> FM[Multi-step flow policy<br/>VaBC π_ω]
  FM -- distill --> Actor[Distilled<br/>one-step actor π_θ]
  Actor -- weighted BC<br/>direct toward high-value --> FM
  FM -. regularize / constrain .-> Actor
  Crit[Critic] --> Actor
Loading

Problem

Standard offline-RL behavior regularization treats every dataset action equally, so the policy clones low-value transitions just as much as high-value ones. Flow-matching policies model the dataset distribution well but, on their own, inherit this unselective imitation.

Method

GFP is a dual-policy BRAC (behavior-regularized actor-critic) framework that couples a multi-step flow-matching policy — termed Value-aware Behavior Cloning (VaBC, the policy $\pi_\omega$) — with a distilled one-step actor ($\pi_\theta$) through mutual guidance. The actor directs the flow policy via weighted behavior cloning: a soft-max guidance function $g_\eta(s,a)$ compares each dataset action $a$ against a proposal of the actor, so VaBC clones high-value actions rather than indiscriminately imitating every transition. In turn, VaBC acts as a distributional regularizer that constrains the actor to stay aligned with the dataset's best transitions while it maximizes the critic.

Results

Reports state-of-the-art performance across 144 state- and pixel-based tasks spanning OGBench (105 tasks, 100 state + 5 pixel), D4RL (18, AntMaze/Adroit), and Minari (21, Adroit/Gym-MuJoCo), with substantial gains specifically on suboptimal datasets and challenging tasks. On the OGBench average GFP scores 53.2 vs. FQL 46.7, ReBRAC 43.9, and IQL 20.4; the gap widens on noisy tasks (e.g. cube-double-noisy GFP ≈63 vs. FQL ≈38). Ablations show the guidance temperature $\eta$ controls filtering sharpness — moderate values are best, while very low temperatures over-concentrate on narrow action sets and destabilize training.

Significance

Shows that the flow-matching policy class — increasingly common in VLAs — is a strong fit for offline RL when paired with value-weighted distillation. Complements on-policy Flow Matching Policy Gradients on the offline side, and is a candidate plug-in for offline post-training of flow / diffusion VLAs like Unified Diffusion VLA.

Links

Authors: Franki Nguimatsia Tiofack, Théotime Le Hellard, Fabian Schramm (Inria / DI-ENS, PSL), Nicolas Perrin-Gilbert (Sorbonne Université), Justin Carpentier.

Related pages

← Back to ICLR-2026 · Topic: RL

⚠️ **GitHub.com Fallback** ⚠️