ICRA 2026 Topic Diffusion Policy - Heungwoo/research GitHub Wiki

ICRA 2026 — Diffusion & Flow-Matching Policies (Topic Analysis)

Scope: 46 ICRA 2026 papers whose core contribution is a diffusion or flow-matching generative policy for manipulation, navigation, or motion generation (sourced from the official ICRA 2026 program; session codes shown in the table). Sibling reviews: VLA Architectures review · RL for VLA ← Back to ICRA-2026-VLA-Manipulation-Survey

Denoising diffusion policies (Chi et al., 2023) remain the dominant recipe for learning continuous, multi-modal action distributions from demonstrations, and ICRA 2026 confirms it: of the imitation-learning track, this 46-paper cluster is the single largest methodological family. The defining tension across the cluster is expressivity vs. speed — diffusion captures multi-modal behavior beautifully but pays an iterative-sampling tax at inference. ICRA 2026 answers this in two ways: (1) a clear rise of flow-matching / rectified-flow policies that trade many denoising steps for one or a few ODE integration steps (DynaFlow, SeFA-Policy, the Motion Manifold Flow line, JFTO), and (2) a wave of structured / conditioned diffusion that injects geometry, object motion, modality priors, or guidance so the generative head has to do less work. Below, eight analytical sub-trends organize the cluster, followed by deep-dives on the papers with confirmed arXiv records and concrete numbers.

Sub-trends

1. Faster diffusion: distillation, flow alignment & efficient training

The largest practical theme is cutting the iterative-sampling cost. SeFA-Policy (Selective Flow Alignment) starts from a rectified-flow policy and adds a consistency-correction step that re-aligns one-step-generated actions with the current observation, recovering accuracy lost in distillation while keeping one-step inference — the authors report cutting inference latency by over 98% vs. multi-step baselines. Mini Diffuser attacks the training cost instead: by exploiting the asymmetry between action-diffusion and image-diffusion it uses two-level mini-batches (many noised action samples per vision-language condition), reaching 95% of SOTA multi-task performance at ~5% of the training time and ~7% of the memory. DynaFlow is a flow-matching model that bakes a differentiable simulator into the flow so trajectories are physically consistent by construction. Together these mark the maturation of "diffusion is great but too slow" into concrete distillation, alignment, and batching recipes.

2. Flow-matching & rectified-flow policies

Beyond pure speed, flow-matching is increasingly the modeling substrate of choice. SeFA-Policy and DynaFlow above are flow-native; Joint Flow Trajectory Optimization (JFTO) uses flow to generate feasible robot motion from video demonstrations while respecting joint-feasibility constraints; Motion Manifold Flow Primitives and the earlier DA-MMP (Dynamics-Aware Motion Manifold Primitives) learn flows on a learned motion manifold so task-conditioned trajectory generation respects complex task–motion dependencies. Conditional Flow-VAE applies the flow family to safety-critical AV scenario generation. This sub-trend is the ICRA-2026 echo of the flow-matching VLA action experts (FPO, π0-style) now permeating classical IL.

3. Structured, equivariant & geometry-aware diffusion

A recurring idea is to stop making the denoiser relearn spatial structure. Hybrid Diffusion Policies with Projective Geometric Algebra (hPGA-DP) embeds geometric inductive biases (translations/rotations as PGA primitives) into the encoder/decoder around a standard U-Net/Transformer denoiser, reporting substantially faster convergence and higher success. Disentangled Point Diffusion diffuses in a disentangled point representation for precise object placement, generalizing across novel object geometries. S²-Diffusion couples a promptable semantic module with a spatial representation to generalize instance-level skills to category-level. These papers share the thesis that the right representation does most of the work the diffusion head would otherwise have to learn.

4. Conditioning & guidance: steering the generative head

Many papers keep the diffusion backbone but add conditioning or test-time guidance. PPGuide steers diffusion policies with a learned performance-predictive guidance signal to suppress compounding action errors. GRITS is a spillage-aware guided diffusion policy for food scooping. DISCO uses VLM-derived 3D keyframes to guide diffusion through constrained inpainting for open-vocabulary language-conditioned manipulation. Motion before Action (MBA) cascades two diffusion processes — first diffuse the object's future motion, then condition robot-action diffusion on it — as a plug-and-play head. Factorizing Diffusion Policies explicitly re-weights observation modalities (proprioception/vision/tactile) by their influence. Camera-conditioning (Plücker-ray extrinsics) and NavDP's privileged-information critic are guidance by another name. This is the cluster's largest conceptual bucket: keep diffusion, add a smarter condition.

5. 3D, point-cloud & trajectory-as-condition diffusion

Spatial grounding shows up via point clouds and 2D/3D trajectory intermediates. Disentangled Point Diffusion is the clearest point-cloud entry. Diffusion Trajectory-Guided Policy (DTP) first diffuses a 2D guidance trajectory with a generative VLM, then refines the imitation policy with it (+25% on CALVIN). DRAW2ACT turns depth-encoded trajectories into robot demonstration videos via video diffusion. Seeing Motion, Generating Action and Motion before Action both route through explicit motion intermediates before action generation. The common pattern: diffuse a spatially meaningful intermediate (points, 2D path, object motion) and condition the action head on it.

6. Cross-embodiment & learning-from-human-video diffusion

Data scarcity pushes several papers to learn diffusion policies from heterogeneous embodiments and human video. X-Diffusion exploits the forward diffusion process: a classifier finds the minimum indistinguishability step and human actions are mixed in only at noise levels where they are indistinguishable from robot actions, giving coarse high-level guidance without dynamically infeasible supervision. Latent Action Diffusion learns a contrastive latent action space unifying anthropomorphic hands, a human hand, and a parallel gripper, co-training a single diffusion policy across them (+25.3% success). Masquerade edits in-the-wild egocentric human video (inpaint arms, overlay a rendered robot) and fine-tunes a diffusion-policy head on ~50 robot demos per task (5–6× over baselines). Inference-Stage Adaptation-Projection adapts a trained diffusion policy to unseen manipulators at inference. This sub-trend treats embodiment gap as a noise-level / latent-space alignment problem.

7. Robustness, safety & long-horizon recovery

As diffusion policies move toward deployment, several papers add guarantees and recovery. MoE-DP inserts a Mixture-of-Experts layer between encoder and denoiser, yielding interpretable skill decomposition and a 36% relative robustness gain under disturbances on long-horizon tasks. Path-Consistent Safety Filtering wraps a diffusion policy in a safety filter that preserves path consistency; SafeFlowMPC fuses learned policies with predictive safe trajectory planning; Diffusion Stabilizer Policy targets surgical-robot stability. Human-in-the-loop entries — HITL-D (diffusion-assisted shared control) and "Uncertainty Comes for Free" (diffusion's own denoising variance as an intervention signal) — turn the generative model's uncertainty into a deployment-safety lever. Compose by Focus and FUNCanon target compositional / long-horizon execution via skill primitives.

8. Diffusion beyond tabletop manipulation: navigation, multi-agent, dexterous & driving

The recipe travels. NavDP is a sim-to-real navigation diffusion policy with a privileged-information critic (363 km / 1244 scenes of sim data; zero-shot to quadruped, wheeled, humanoid). Ventura adapts image-diffusion models for task-conditioned navigation. MIMIC-D does decentralized multi-agent diffusion (CTDE) for implicit coordination (95% on a real bimanual basket-lift from 16 demos), and ADM-DP fuses vision-tactile-graph modalities for multi-agent manipulation. Physics-Informed Diffusion Mamba Transformer brings diffusion planning to real-world driving. Application-specific diffusion policies span surgery (Diffusion Stabilizer), assistive dressing (Diffusion Policy for Robot-Assisted Dressing), food (GRITS, SCOOP'D), throwing (DA-MMP), and humanoid skills (MAKP kicking, Unified Fall-Safety). This breadth is the clearest signal that diffusion/flow is now the default continuous-action IL backbone, not a niche.

Standout deep-dives

Seven papers with confirmed arXiv records and concrete reported numbers. (Numbers below are as stated by the authors; this cluster's RL-fine-tuning cousin is Flow Policy Optimization (FPO), and the backbone taxonomy is in the VLA Architectures review.)

  • SeFA-Policy — Selective Flow Alignment (arXiv 2511.08583). A rectified-flow visuomotor policy with a consistency-correction step that uses expert demos to re-align one-step-generated actions with the current observation while preserving multimodality. Reported to surpass SOTA diffusion/flow policies while reducing inference latency by over 98%. The cleanest answer in the cluster to "one-step flow loses accuracy."

  • Mini Diffuser — two-level mini-batch training (arXiv 2505.09430). Exploits the asymmetry between action-diffusion and image-diffusion by pairing many noised action samples with each vision-language condition. On RLBench it reaches 95% of SOTA multi-task diffusion-policy performance using ~5% of the training time and ~7% of the memory — an order-of-magnitude training-cost cut, the standout efficiency result here.

  • X-Diffusion — cross-embodiment human demos (arXiv 2511.04671). Trains a classifier to detect the minimum indistinguishability step in the forward diffusion process; human actions are injected into training only at noise levels where the classifier cannot tell human from robot, so feasible motions give low-level supervision and mismatched ones give only coarse high-level guidance. A principled noise-level treatment of the embodiment gap.

  • Latent Action Diffusion (arXiv 2506.14608). Learns a contrastive latent action space unifying anthropomorphic hands, a human hand, and a parallel-jaw gripper, then co-trains a single diffusion policy across them for multi-robot control with up to +25.3% manipulation success from cross-embodiment transfer.

  • NavDP — sim-to-real navigation diffusion (arXiv 2505.08712). A transformer that jointly diffuses candidate trajectories and learns a critic from simulator privileged information to select among them. Trained purely in sim (~2,500 trajectories/GPU/day; 363.2 km across 1244 scenes), it transfers zero-shot to quadruped, wheeled, and humanoid robots — the cluster's strongest evidence that diffusion policies scale beyond tabletop.

  • MoE-DP — MoE-enhanced diffusion policy (arXiv 2511.05007). Inserts a Mixture-of-Experts layer between the visual encoder and the denoiser; experts specialize to semantic primitives (approach, grasp), giving an interpretable skill decomposition and a 36% average relative robustness improvement under disturbances on six long-horizon tasks.

  • Diffusion Trajectory-Guided Policy (DTP) (arXiv 2502.10040). Two stages: a generative VLM diffuses a 2D guidance trajectory, then the imitation policy is refined with it. Outperforms SOTA by 25% on CALVIN from scratch (no external pretraining), exemplifying the "diffuse a spatial intermediate, condition the action head" pattern.

Complete paper list (46)

Code Title arXiv
ThBT2.9 Factorizing Diffusion Policies for Observation Modality Prioritization 2509.16830
ThI1I.11 Motion before Action: Diffusing Object Motion As Manipulation Condition 2411.09658
ThI1I.127 Compose by Focus: Scene Graph-Based Atomic Skills 2509.16053
ThI1I.136 MIMIC-D: Multi-Modal Imitation for Multi-Agent Coordination with Decentralized Diffusion Policies 2509.14159
ThI1I.257 Conditional Flow-VAE for Safety-Critical Traffic Scenario Generation 2605.04366
ThI1I.282 PPGuide: Steering Diffusion Policies with Performance Predictive Guidance 2603.10980
ThI1I.285 Prepare before You Act: Learning from Humans to Rearrange Initial States 2509.18043
ThI1I.367 S²-Diffusion: Generalizing from Instance-Level to Category-Level Skills in Robot Manipulation 2502.09389
ThI1I.379 Mini Diffuser: Fast Multi-Task Diffusion Policy Training Using Two-Level Mini-Batches 2505.09430
ThI1I.58 RoboMatch: A Unified Mobile-Manipulation Teleoperation Platform with Auto-Matching Network Architecture for Long-Horizon Tasks 2509.08522
ThI1I.63 DynaFlow: Dynamics-Embedded Flow Matching for Physically Consistent Motion Generation from State-Only Demonstrations 2509.19804
ThI1I.88 Physically-Based Lighting Generation for Robotic Manipulation 2508.01442
ThI2I.122 HITL-D: Human in the Loop Diffusion Assisted Shared Control 2605.21460
ThI2I.161 DA-MMP: Learning Coordinated and Accurate Throwing with Dynamics-Aware Motion Manifold Primitives 2509.23721
ThI2I.262 WorldPlanner: Monte Carlo Tree Search and MPC with Action-Conditioned Visual World Models 2511.03077
ThI2I.281 X-Diffusion: Training Diffusion Policies on Cross-Embodiment Human Demonstrations 2511.04671
ThI2I.311 Diffusion Stabilizer Policy for Automated Surgical Robot Manipulations 2503.01252
ThI2I.71 SafeFlowMPC: Predictive and Safe Trajectory Planning for Robot Manipulators with Learning-Based Policies 2602.12794
ThI2LB.10 Diffusion Policy for Robot-Assisted Dressing with Moving Human Arms —
TuAT1.1 GRITS: A Spillage-Aware Guided Diffusion Policy for Robot Food Scooping Tasks 2510.00573
TuAT1.4 Do You Know Where Your Camera Is? View-Invariant Policy Learning with Camera Conditioning 2510.02268
TuBT1.2 Uncertainty Comes for Free: Human-In-The-Loop Policies with Diffusion Models 2503.01876
TuI1I.104 Scaling Single Human Demonstrations for Imitation Learning Using Generative Foundational Models 2602.12734
TuI1I.170 Joint Flow Trajectory Optimization for Feasible Robot Motion Generation from Video Demonstrations 2509.20703
TuI1I.191 From Demonstrations to Safe Deployment: Path-Consistent Safety Filtering for Diffusion Policies 2511.06385
TuI1I.215 FUNCanon: Learning Pose-Aware Action Primitives Via Functional Object Canonicalization for Generalizable Robotic Manipulation 2509.19102
TuI1I.259 Physics-Informed Diffusion Mamba Transformer for Real-World Driving 2602.00808
TuI1I.299 Ventura: Adapting Image Diffusion Models for Unified Task Conditioned Navigation 2510.01388
TuI1I.63 Latent Action Diffusion for Cross-Embodiment Manipulation 2506.14608
TuI1I.78 Seeing Motion, Generating Action: Explicit Motion-Aware Policy for Robotic Action Generation —
TuI2I.158 MAKP: Multi-Mode Accurate Kicking Policy for Humanoid Robots —
TuI2I.160 NavDP: Learning Sim-To-Real Navigation Diffusion Policy with Privileged Information Guidance 2505.08712
TuI2I.45 Inference-Stage Adaptation-Projection Strategy Adapts Diffusion Policy to Cross-Manipulators Scenarios 2509.11621
TuI2I.63 SCOOP'D: Learning Mixed-Liquid-Solid Scooping Via Sim2Real Generative Policy 2510.11566
WeBT2.3 Hybrid Diffusion Policies with Projective Geometric Algebra for Efficient Robot Manipulation Learning 2507.05695
WeI1I.114 Unified Humanoid Fall-Safety Policy from a Few Demonstrations 2511.07407
WeI1I.153 MoE-DP: An MoE-Enhanced Diffusion Policy for Robust Long-Horizon Robotic Manipulation with Skill Decomposition and Failure Recovery 2511.05007
WeI1I.192 DRAW2ACT: Turning Depth-Encoded Trajectories into Robotic Demonstration Videos 2512.14217
WeI1I.211 ADM-DP: Adaptive Dynamic Modality Diffusion Policy through Vision-Tactile-Graph Fusion for Multi-Agent Manipulation 2602.21622
WeI1I.222 Masquerade: Learning from In-The-Wild Human Videos Using Data-Editing 2508.09976
WeI1I.23 DISCO: Language-Guided Manipulation with Diffusion Policies and Constrained Inpainting 2406.09767
WeI1I.27 Motion Manifold Flow Primitives for Task-Conditioned Trajectory Generation under Complex Task-Motion Dependencies 2407.19681
WeI2I.132 Disentangled Point Diffusion for Precise Object Placement —
WeI2I.299 SeFA-Policy: Fast and Accurate Visuomotor Policy Learning with Selective Flow Alignment 2511.08583
WeI2I.311 Physically-Grounded Data Generation Via Video Diffusion Models —
WeI2I.340 Diffusion Trajectory-Guided Policy for Long-Horizon Robot Manipulation 2502.10040

Related

← Back to ICRA-2026-VLA-Manipulation-Survey