ICLR 2026 Primary Fine Decoupling - Heungwoo/research GitHub Wiki

PF-DAG — primary-fine decoupling for action generation in imitation

Venue: ICLR 2026 Authors: Xiaohan Lei, Min Wang, Wengang Zhou, Xingyu Lu, Houqiang Li arXiv: 2602.21684 Category: Policy learning — generative action heads for imitation Trend tag: Multimodal action generation / MeanFlow / mode decoupling

Approach diagram

flowchart LR
  Demo[Demonstration action chunks] --> Comp[Stage 1: compress chunks<br/>into small set of discrete modes]
  Obs[Observation] --> Sel[Lightweight policy:<br/>select consistent coarse mode]
  Comp --> Sel
  Sel --> Mode[Coarse mode<br/>avoids mode bouncing]
  Mode --> MF[Stage 2: mode-conditioned<br/>MeanFlow policy]
  Obs --> MF
  MF --> Act[High-fidelity continuous actions]
Loading

Problem

Robotic manipulation action sequences are multi-modal. Existing approaches either discretize actions into tokens (losing fine-grained variation) or generate continuous actions in a single stage (producing unstable mode transitions, i.e. "mode bouncing"). The paper aims to keep coarse-mode consistency while preserving fine continuous detail.

Method

PF-DAG (Primary-Fine Decoupling for Action Generation) is a two-stage framework that decouples coarse action consistency from fine-grained variation:

  1. Compress action chunks into a small set of discrete modes, so a lightweight policy can select consistent coarse modes and avoid mode bouncing.
  2. Learn a mode-conditioned MeanFlow policy to generate high-fidelity continuous actions.

Theoretically, the authors prove the two-stage design achieves a strictly lower MSE bound than single-stage generative policies.

Results

PF-DAG outperforms state-of-the-art baselines across 56 tasks from Adroit, DexArt, and MetaWorld, and further generalizes to real-world tactile dexterous manipulation tasks.

Significance

PF-DAG offers a principled decomposition of action generation into a discrete coarse-mode selector and a continuous fine generator, backed by an MSE-bound argument — addressing the instability of single-stage continuous policies while retaining fine detail that pure tokenization discards. The MeanFlow fine stage ties it to the broader flow-matching action-head trend.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️