ICML 2026 FocalPolicy - Heungwoo/research GitHub Wiki

FocalPolicy: Frequency-Optimized Chunking and Locally Anchored Flow Matching for Coherent Visuomotor Policy — synergizing proximal precision with distal coherence

Venue: ICML 2026 (Poster) Category: Diffusion-Flow Policy Traction (2026-06): 0 citations (arXiv)

Comparison with chunk-based baselines: previous approaches (a) prioritize intra-chunk refinement but overlook inter-chunk discontinuities, while FocalPolicy (b) uses a Foresight Composite Objective to synergize proximal precision with distal coherence across chunks (Figure 1 from He et al., 2026)

Problem

Visuomotor policies learn manipulation skills from expert demonstrations, but generating smooth, coherent trajectories remains hard because it requires balancing proximal precision (getting the immediate actions right) with distal foresight (staying coherent over the longer horizon). Existing chunk-based policies typically optimize the action distribution within a single chunk while neglecting inter-chunk coherence. The resulting inter-chunk discontinuities accumulate as compounding errors and impede the learning of coherent long-horizon behavior.

Method

FocalPolicy is a foresight-aware visuomotor policy that combines Frequency-Optimized Chunking with Locally Anchored flow matching. It has two core components:

  • Foresight Composite Objective (FCO): supervises time-domain alignment on the proximal (near-term) actions while regularizing frequency-domain structure across multiple future action chunks. This couples short-term precision with cross-chunk coherence rather than refining each chunk in isolation.
  • Locally Anchored Sampling (LAS): a time-sampling scheme for consistency flow matching. Standard uniform sampling of the flow timepoints can attenuate target-signal propagation for early time τ. LAS instead biases the anchor time r toward the terminal region (τ → 1), strengthening target-signal propagation and improving training efficiency. The authors' optimal configuration is μ_r = 4.0, σ_r = 1.6 (biased toward 1.0).

The components are modular: the paper shows FCO and LAS can be transplanted into other baselines (DP3 w. FCO, FlowPolicy w. LAS).

Pipeline of FocalPolicy: Locally Anchored Sampling improves consistency-flow-matching training efficiency, and the Foresight Composite Objective synergizes proximal time-domain precision with distal frequency-domain coherence (Figure 2 from He et al., 2026)

Results

On a simulation suite spanning Adroit and MetaWorld (53 tasks; averages over 53 tasks), FocalPolicy reaches an average success of 83.6% with NFE = 1 (a single function evaluation), versus DP3 at 79.9% (NFE = 10), FlowPolicy at 79.4% (NFE = 1), FreqPolicy at 75.1%, SDM at 74.4% and DP at 55.9%. Per-category gains are large on harder splits (e.g., MetaWorld Medium 81.9%, Very Hard 85.1%).

Ablations on 24 tasks isolate the contributions: full model 77.9% average, w/o LAS drops to 64.3%, and w/o FCO & LAS drops further — confirming both modules matter. Learning curves show FocalPolicy converges faster and higher than FlowPolicy. The method also reduces compounding error: 3D end-effector trajectories track expert paths more closely (lower Euclidean error on Bin Picking and Push Wall), with lower Action Total Variation (ATV) indicating smoother, more temporally coherent action sequences. Real-world experiments on a UR-10e arm with an AG-95 gripper and a RealSense D435i camera evaluate six staged tasks (Water Pouring, Pot Loading, Cup Matching, Drawer Loading, Tower Stacking, Object Sorting), reporting per-stage success scores against FlowPolicy, DP3 and FreqPolicy. Compute is modest: training in ~1.5 h using ~8 GB on a single RTX 4090.

Significance

FocalPolicy reframes visuomotor learning around inter-chunk coherence, showing that a frequency-domain consistency objective plus a terminal-biased sampling scheme yields smoother, more accurate long-horizon manipulation at single-step inference (NFE = 1). Because FCO and LAS generalize to other flow/diffusion policies, the contribution is a reusable recipe for reducing compounding error in chunk-based action generation.

Links

← Back to ICML-2026