ICML 2026 OMP - Heungwoo/research GitHub Wiki
OMP: One-step MeanFlow Policy — directional alignment for high-precision single-step manipulation
Venue: ICML 2026 (Poster) Category: Diffusion-Flow Policy Affiliations: Han Fang, Yize Huang, Yuheng Zhao, Paul Weng, Xiao Li, Yutong Ban Traction (2026-06): 4 citations (arXiv)

Problem
Generative robot policies face a hard trade-off: diffusion models (DP3, NFE=10) achieve high success but suffer ~132 ms latency on Adroit, while flow-based methods (FlowPolicy, AdaFlow) reach single-step inference at the cost of architectural complexity and over-constrained training that hurts generalization. MeanFlow and its first robotics adaptation MP1 enable NFE=1 inference (6.8 ms, 19x faster than DP3) but have two flaws: the model predicts instantaneous velocity yet is trained on interval-averaged velocities (a mismatch that degrades trajectory accuracy), and the required Jacobian-Vector Product (JVP) operator consumes heavy GPU memory.
Method
OMP improves MeanFlow-based policies along two axes. (1) Directional Alignment. The authors analyze the geometry of the MSE loss in high-dimensional velocity regression and show that the directional gradient term scales as 2·ρ·ρ*·sin α — i.e., the gradient that corrects direction vanishes as the target velocity magnitude ρ*→0. This is exactly the regime of high-precision tasks (peg insert, thread-in-hole), explaining why standard flow models stall on fine manipulation. OMP adds a lightweight Cosine Loss that directly aligns the direction of the predicted interval-averaged velocity with the true mean velocity, restoring directional gradient signal independent of magnitude. (2) DDE for the JVP operator. OMP replaces the exact JVP with a Differential Derivation Equation (DDE) that approximates the Jacobian-Vector Product via finite differences, sharply reducing GPU memory for complex tasks at the cost of small approximation error. A Dispersive Loss term is retained for latent-feature discrimination.

Results
Evaluated on 3 Adroit tasks and 34 Meta-World tasks (10 expert demos each, RTX 4090, three seeds). OMP reaches 82.3% ± 1.7% average success, vs. MP1's 78.9% ± 2.1% — a 3.4% gain over MP1 and 10.7% over FlowPolicy. Because MP1 is already near-optimal on the 21 Meta-World "easy" tasks (88.2%), the headline average understates OMP; on harder splits OMP improves Meta-World Medium by 9.4%, Hard by 4.4%, and Very-Hard by 10.6%. Training curves show OMP converges faster and more stably than baselines. Ablations confirm removing the directional alignment (Cosine Loss) causes a significant drop across all Adroit and Meta-World tasks; swapping JVP for DDE trades a small accuracy drop for reduced GPU memory.
Significance
OMP gives a clean theoretical account of why single-step flow policies fail on fine manipulation — the vanishing directional gradient — and fixes it with a near-free cosine term, while the DDE trick makes MeanFlow-style NFE=1 policies practical on memory-limited hardware. It pushes the real-time, single-step generative-policy frontier that matters for high-frequency robot control.
Links
- arXiv: 2512.19347
- ICML 2026: https://icml.cc/virtual/2026/poster/66693
← Back to ICML-2026