CoRL 2025 Streaming Flow Policy - Heungwoo/research GitHub Wiki

Streaming Flow Policy

Venue: CoRL 2025 (Oral) · arXiv: 2505.21851 Authors: Sunshine Jiang, Xiaolin Fang, Nicholas Roy, Tomás Lozano-Pérez, Leslie Pack Kaelbling, Siddharth Ancha (MIT) Also: Best Paper Nominee (3/21), ICRA 2025 Beyond Pick-and-Place Workshop Category: Flow / Diffusion Policies Trend tag: Trajectory streaming

Approach diagram

flowchart LR
  H[Observation history h] --> V[History-conditioned<br/>velocity field vθ a,t|h]
  V -- integrate from a≈a_prev --> T["Flow time t = execution time"]
  T --> A1[action a_t1]
  T --> A2[action a_t2]
  T --> A3[action a_t3]
  A1 & A2 & A3 --> EXEC[Stream to robot on-the-fly]
Loading

Problem

Standard diffusion/flow policies sample an entire action chunk from noise in the trajectory space 𝒜^T before any action can execute, paying full denoising latency up front and producing discontinuities at chunk boundaries.

Method

Streaming Flow Policy treats the action trajectory itself as the flow trajectory: it learns a velocity field directly in action space 𝒜 and integrates it so that the flow-integration time coincides with execution time. Integration starts from a narrow Gaussian around the previous action, and each integration step emits the next action, which can be streamed to the robot on-the-fly during sampling (no "trajectory of trajectories", no chunk windowing).

  • History-conditioned velocity field vθ(a, t | h): inputs are the current action a, normalized flow time t ∈ [0,1], and observation history h.
  • Stabilizing conditional flow: each demonstration ξ defines v(a,t) = ξ̇(t) − k·(a − ξ(t)), a feedback term with gain k that contracts a thin Gaussian "tube" (variance σ₀²·e^(−2kt)) around the demo, reducing distribution shift.
  • Multimodality is preserved by training the marginal velocity over a mixture of these per-demonstration tubes p*(a|t,h) = ∫ pξ(a|t)·p𝒟(ξ|h) dξ, matching the per-timestep training distribution without sampling whole trajectories.

Results

CoRL 2025 Oral. On Push-T (state), SFP reaches 95.1 / 96.0 (avg/max) vs Diffusion Policy 92.9 / 94.4 and flow-matching policy 80.6 / 82.6. On RoboMimic (state): Lift 100/100, Can 98.4/100 (DP 94.8/98.0), Square 78.0/84.0 (DP 77.2/84.0). Per-action latency ≈ 3.5 ms vs DP-100-DDPM 40.2 ms and 10-step DDIM 4.4 ms; the streaming property lets sampling overlap with execution, avoiding chunk-boundary stalls and jerk.

Significance

Part of the CoRL 2025 trajectory-streaming trend alongside SAIL and DemoSpeedup. By collapsing flow time into execution time, it removes the chunk-boundary problem for flow-matching policies entirely rather than smoothing over it. Related to ICLR 2026's runtime-efficiency work (FASTER, OmniSAT, HyperVLA).

Links

Related pages

← Back to CoRL-2025

⚠️ **GitHub.com Fallback** ⚠️