CoRL 2025 Streaming Flow Policy - Heungwoo/research GitHub Wiki
Venue: CoRL 2025 (Oral) · arXiv: 2505.21851 Authors: Sunshine Jiang, Xiaolin Fang, Nicholas Roy, Tomás Lozano-Pérez, Leslie Pack Kaelbling, Siddharth Ancha (MIT) Also: Best Paper Nominee (3/21), ICRA 2025 Beyond Pick-and-Place Workshop Category: Flow / Diffusion Policies Trend tag: Trajectory streaming
flowchart LR
H[Observation history h] --> V[History-conditioned<br/>velocity field vθ a,t|h]
V -- integrate from a≈a_prev --> T["Flow time t = execution time"]
T --> A1[action a_t1]
T --> A2[action a_t2]
T --> A3[action a_t3]
A1 & A2 & A3 --> EXEC[Stream to robot on-the-fly]
Standard diffusion/flow policies sample an entire action chunk from noise in the trajectory space 𝒜^T before any action can execute, paying full denoising latency up front and producing discontinuities at chunk boundaries.
Streaming Flow Policy treats the action trajectory itself as the flow trajectory: it learns a velocity field directly in action space 𝒜 and integrates it so that the flow-integration time coincides with execution time. Integration starts from a narrow Gaussian around the previous action, and each integration step emits the next action, which can be streamed to the robot on-the-fly during sampling (no "trajectory of trajectories", no chunk windowing).
-
History-conditioned velocity field
vθ(a, t | h): inputs are the current actiona, normalized flow timet ∈ [0,1], and observation historyh. -
Stabilizing conditional flow: each demonstration ξ defines
v(a,t) = ξ̇(t) − k·(a − ξ(t)), a feedback term with gainkthat contracts a thin Gaussian "tube" (varianceσ₀²·e^(−2kt)) around the demo, reducing distribution shift. -
Multimodality is preserved by training the marginal velocity over a mixture of these per-demonstration tubes
p*(a|t,h) = ∫ pξ(a|t)·p𝒟(ξ|h) dξ, matching the per-timestep training distribution without sampling whole trajectories.
CoRL 2025 Oral. On Push-T (state), SFP reaches 95.1 / 96.0 (avg/max) vs Diffusion Policy 92.9 / 94.4 and flow-matching policy 80.6 / 82.6. On RoboMimic (state): Lift 100/100, Can 98.4/100 (DP 94.8/98.0), Square 78.0/84.0 (DP 77.2/84.0). Per-action latency ≈ 3.5 ms vs DP-100-DDPM 40.2 ms and 10-step DDIM 4.4 ms; the streaming property lets sampling overlap with execution, avoiding chunk-boundary stalls and jerk.
Part of the CoRL 2025 trajectory-streaming trend alongside SAIL and DemoSpeedup. By collapsing flow time into execution time, it removes the chunk-boundary problem for flow-matching policies entirely rather than smoothing over it. Related to ICLR 2026's runtime-efficiency work (FASTER, OmniSAT, HyperVLA).
- arXiv: https://arxiv.org/abs/2505.21851
- Project page: https://streaming-flow-policy.github.io
← Back to CoRL-2025