ICLR 2026 Primary Fine Decoupling - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Authors: Xiaohan Lei, Min Wang, Wengang Zhou, Xingyu Lu, Houqiang Li arXiv: 2602.21684 Category: Policy learning — generative action heads for imitation Trend tag: Multimodal action generation / MeanFlow / mode decoupling
flowchart LR
Demo[Demonstration action chunks] --> Comp[Stage 1: compress chunks<br/>into small set of discrete modes]
Obs[Observation] --> Sel[Lightweight policy:<br/>select consistent coarse mode]
Comp --> Sel
Sel --> Mode[Coarse mode<br/>avoids mode bouncing]
Mode --> MF[Stage 2: mode-conditioned<br/>MeanFlow policy]
Obs --> MF
MF --> Act[High-fidelity continuous actions]
Robotic manipulation action sequences are multi-modal. Existing approaches either discretize actions into tokens (losing fine-grained variation) or generate continuous actions in a single stage (producing unstable mode transitions, i.e. "mode bouncing"). The paper aims to keep coarse-mode consistency while preserving fine continuous detail.
PF-DAG (Primary-Fine Decoupling for Action Generation) is a two-stage framework that decouples coarse action consistency from fine-grained variation:
- Compress action chunks into a small set of discrete modes, so a lightweight policy can select consistent coarse modes and avoid mode bouncing.
- Learn a mode-conditioned MeanFlow policy to generate high-fidelity continuous actions.
Theoretically, the authors prove the two-stage design achieves a strictly lower MSE bound than single-stage generative policies.
PF-DAG outperforms state-of-the-art baselines across 56 tasks from Adroit, DexArt, and MetaWorld, and further generalizes to real-world tactile dexterous manipulation tasks.
PF-DAG offers a principled decomposition of action generation into a discrete coarse-mode selector and a continuous fine generator, backed by an MSE-bound argument — addressing the instability of single-stage continuous policies while retaining fine detail that pure tokenization discards. The MeanFlow fine stage ties it to the broader flow-matching action-head trend.
- arXiv: https://arxiv.org/abs/2602.21684
- OpenReview: https://openreview.net/forum?id=wySMuWHmt4
← Back to ICLR-2026