RSS 2026 Latent Policy Steering through One Step - Heungwoo/research GitHub Wiki
Latent Policy Steering through One-Step Flow Policies
Venue: RSS 2026 (Sydney, Jul 13–17) · Session: RL · paper #152 Authors: Hokyun Im, Andrey Kolobov, Jianlong Fu, Youngwoon Lee arXiv: 2603.05296 · program page
Summary compiled from the arXiv paper (v2); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Top: QC-FQL constrains a one-step actor with an explicit regularizer, creating the return-vs-constraint trade-off. Middle: DSRL steers a flow-matching policy's latents but needs a latent-space critic Q(s,z) obtained via lossy distillation. Bottom: LPS backpropagates action-space critic gradients ∇aQ(s,a) directly through a differentiable one-step MeanFlow base policy into the latent actor — no proxy critic, no regularization weight.
Problem
Offline RL for robots hinges on a brittle trade-off between return maximization (which pushes policies off dataset support) and behavior regularization (whose weight α is highly sensitive to reward scale, dataset diversity, and capacity — infeasible to sweep on real robots). Latent steering (DSRL) sidesteps α but, in the fully offline setting, must distill an approximate latent-space critic via noise aliasing, which is lossy and can fail to converge.
Method
Latent Policy Steering (LPS) (Yonsei, Microsoft Research) keeps a behavior-cloned MeanFlow one-step generative policy fixed as a behavior-constrained prior, and trains a latent-space actor by backpropagating original-action-space Q-gradients through the differentiable one-step policy — eliminating proxy latent critics entirely. Value learning uses Q-Chunking (shared across all compared methods; chunk length h=5). Simulation uses 4-layer MLP base policies; the real-robot setup uses a 114M-parameter DiT base policy on the DROID platform with 50 teleoperated demos per task, semi-sparse reward (−1 per step, 0 on success), γ=0.99, 10K gradient steps.
Results
On OGBench (cube-single/double, scene, puzzle-3x3/4x4 plus a pixel-based visual task), LPS consistently outperforms one-step-distillation baselines (QC-FQL, QC-MFQL), CFGRL, and DSRL, which shows high variance and fails on cube-double; LPS is stable across α values from 0.01 to 300 where QC-FQL peaks sharply. On four real DROID tasks (pick-and-place carrots, eggplant-to-bin, refill tape, plug-in-bulb; 20 trials each), LPS attains the highest success on every task and the best average, while DSRL scores 0% on plug-in-bulb — worse than its own BC base. LPS also fine-tunes online (insert-pen: surpasses offline baseline and DSRL within 5K environment steps from 20 demos), trains faster than DSRL, and inherits MeanFlow's one-step inference latency.
Significance
A clean fix for offline latent steering: make the generative prior one-step and differentiable, and the action-space critic can drive latent optimization end-to-end — tuning-free structural regularization with real-robot evidence. Slots into the wiki's RL thread on offline-to-real policy improvement and the one-step-generation efficiency theme of Review-Realtime-Execution.
← Back to RSS 2026 survey · RSS-2026-Papers · Home