ICML 2026 FocalPolicy - Heungwoo/research GitHub Wiki
FocalPolicy: Frequency-Optimized Chunking and Locally Anchored Flow Matching for Coherent Visuomotor Policy — synergizing proximal precision with distal coherence
Venue: ICML 2026 (Poster) Category: Diffusion-Flow Policy Traction (2026-06): 0 citations (arXiv)

Problem
Visuomotor policies learn manipulation skills from expert demonstrations, but generating smooth, coherent trajectories remains hard because it requires balancing proximal precision (getting the immediate actions right) with distal foresight (staying coherent over the longer horizon). Existing chunk-based policies typically optimize the action distribution within a single chunk while neglecting inter-chunk coherence. The resulting inter-chunk discontinuities accumulate as compounding errors and impede the learning of coherent long-horizon behavior.
Method
FocalPolicy is a foresight-aware visuomotor policy that combines Frequency-Optimized Chunking with Locally Anchored flow matching. It has two core components:
- Foresight Composite Objective (FCO): supervises time-domain alignment on the proximal (near-term) actions while regularizing frequency-domain structure across multiple future action chunks. This couples short-term precision with cross-chunk coherence rather than refining each chunk in isolation.
- Locally Anchored Sampling (LAS): a time-sampling scheme for consistency flow matching. Standard uniform sampling of the flow timepoints can attenuate target-signal propagation for early time τ. LAS instead biases the anchor time r toward the terminal region (τ → 1), strengthening target-signal propagation and improving training efficiency. The authors' optimal configuration is μ_r = 4.0, σ_r = 1.6 (biased toward 1.0).
The components are modular: the paper shows FCO and LAS can be transplanted into other baselines (DP3 w. FCO, FlowPolicy w. LAS).

Results
On a simulation suite spanning Adroit and MetaWorld (53 tasks; averages over 53 tasks), FocalPolicy reaches an average success of 83.6% with NFE = 1 (a single function evaluation), versus DP3 at 79.9% (NFE = 10), FlowPolicy at 79.4% (NFE = 1), FreqPolicy at 75.1%, SDM at 74.4% and DP at 55.9%. Per-category gains are large on harder splits (e.g., MetaWorld Medium 81.9%, Very Hard 85.1%).
Ablations on 24 tasks isolate the contributions: full model 77.9% average, w/o LAS drops to 64.3%, and w/o FCO & LAS drops further — confirming both modules matter. Learning curves show FocalPolicy converges faster and higher than FlowPolicy. The method also reduces compounding error: 3D end-effector trajectories track expert paths more closely (lower Euclidean error on Bin Picking and Push Wall), with lower Action Total Variation (ATV) indicating smoother, more temporally coherent action sequences. Real-world experiments on a UR-10e arm with an AG-95 gripper and a RealSense D435i camera evaluate six staged tasks (Water Pouring, Pot Loading, Cup Matching, Drawer Loading, Tower Stacking, Object Sorting), reporting per-stage success scores against FlowPolicy, DP3 and FreqPolicy. Compute is modest: training in ~1.5 h using ~8 GB on a single RTX 4090.
Significance
FocalPolicy reframes visuomotor learning around inter-chunk coherence, showing that a frequency-domain consistency objective plus a terminal-biased sampling scheme yields smoother, more accurate long-horizon manipulation at single-step inference (NFE = 1). Because FCO and LAS generalize to other flow/diffusion policies, the contribution is a reusable recipe for reducing compounding error in chunk-based action generation.
Links
- arXiv: 2605.15944
- Project: https://focalpolicy.github.io/
- ICML 2026: https://icml.cc/virtual/2026/poster/66520
← Back to ICML-2026