ICRA 2026 Topic Diffusion Policy - Heungwoo/research GitHub Wiki
ICRA 2026 — Diffusion & Flow-Matching Policies (Topic Analysis)
Scope: 46 ICRA 2026 papers whose core contribution is a diffusion or flow-matching generative policy for manipulation, navigation, or motion generation (sourced from the official ICRA 2026 program; session codes shown in the table). Sibling reviews: VLA Architectures review · RL for VLA ← Back to ICRA-2026-VLA-Manipulation-Survey
Denoising diffusion policies (Chi et al., 2023) remain the dominant recipe for learning continuous, multi-modal action distributions from demonstrations, and ICRA 2026 confirms it: of the imitation-learning track, this 46-paper cluster is the single largest methodological family. The defining tension across the cluster is expressivity vs. speed — diffusion captures multi-modal behavior beautifully but pays an iterative-sampling tax at inference. ICRA 2026 answers this in two ways: (1) a clear rise of flow-matching / rectified-flow policies that trade many denoising steps for one or a few ODE integration steps (DynaFlow, SeFA-Policy, the Motion Manifold Flow line, JFTO), and (2) a wave of structured / conditioned diffusion that injects geometry, object motion, modality priors, or guidance so the generative head has to do less work. Below, eight analytical sub-trends organize the cluster, followed by deep-dives on the papers with confirmed arXiv records and concrete numbers.
Sub-trends
1. Faster diffusion: distillation, flow alignment & efficient training
The largest practical theme is cutting the iterative-sampling cost. SeFA-Policy (Selective Flow Alignment) starts from a rectified-flow policy and adds a consistency-correction step that re-aligns one-step-generated actions with the current observation, recovering accuracy lost in distillation while keeping one-step inference — the authors report cutting inference latency by over 98% vs. multi-step baselines. Mini Diffuser attacks the training cost instead: by exploiting the asymmetry between action-diffusion and image-diffusion it uses two-level mini-batches (many noised action samples per vision-language condition), reaching 95% of SOTA multi-task performance at ~5% of the training time and ~7% of the memory. DynaFlow is a flow-matching model that bakes a differentiable simulator into the flow so trajectories are physically consistent by construction. Together these mark the maturation of "diffusion is great but too slow" into concrete distillation, alignment, and batching recipes.
2. Flow-matching & rectified-flow policies
Beyond pure speed, flow-matching is increasingly the modeling substrate of choice. SeFA-Policy and DynaFlow above are flow-native; Joint Flow Trajectory Optimization (JFTO) uses flow to generate feasible robot motion from video demonstrations while respecting joint-feasibility constraints; Motion Manifold Flow Primitives and the earlier DA-MMP (Dynamics-Aware Motion Manifold Primitives) learn flows on a learned motion manifold so task-conditioned trajectory generation respects complex task–motion dependencies. Conditional Flow-VAE applies the flow family to safety-critical AV scenario generation. This sub-trend is the ICRA-2026 echo of the flow-matching VLA action experts (FPO, π0-style) now permeating classical IL.
3. Structured, equivariant & geometry-aware diffusion
A recurring idea is to stop making the denoiser relearn spatial structure. Hybrid Diffusion Policies with Projective Geometric Algebra (hPGA-DP) embeds geometric inductive biases (translations/rotations as PGA primitives) into the encoder/decoder around a standard U-Net/Transformer denoiser, reporting substantially faster convergence and higher success. Disentangled Point Diffusion diffuses in a disentangled point representation for precise object placement, generalizing across novel object geometries. S²-Diffusion couples a promptable semantic module with a spatial representation to generalize instance-level skills to category-level. These papers share the thesis that the right representation does most of the work the diffusion head would otherwise have to learn.
4. Conditioning & guidance: steering the generative head
Many papers keep the diffusion backbone but add conditioning or test-time guidance. PPGuide steers diffusion policies with a learned performance-predictive guidance signal to suppress compounding action errors. GRITS is a spillage-aware guided diffusion policy for food scooping. DISCO uses VLM-derived 3D keyframes to guide diffusion through constrained inpainting for open-vocabulary language-conditioned manipulation. Motion before Action (MBA) cascades two diffusion processes — first diffuse the object's future motion, then condition robot-action diffusion on it — as a plug-and-play head. Factorizing Diffusion Policies explicitly re-weights observation modalities (proprioception/vision/tactile) by their influence. Camera-conditioning (Plücker-ray extrinsics) and NavDP's privileged-information critic are guidance by another name. This is the cluster's largest conceptual bucket: keep diffusion, add a smarter condition.
5. 3D, point-cloud & trajectory-as-condition diffusion
Spatial grounding shows up via point clouds and 2D/3D trajectory intermediates. Disentangled Point Diffusion is the clearest point-cloud entry. Diffusion Trajectory-Guided Policy (DTP) first diffuses a 2D guidance trajectory with a generative VLM, then refines the imitation policy with it (+25% on CALVIN). DRAW2ACT turns depth-encoded trajectories into robot demonstration videos via video diffusion. Seeing Motion, Generating Action and Motion before Action both route through explicit motion intermediates before action generation. The common pattern: diffuse a spatially meaningful intermediate (points, 2D path, object motion) and condition the action head on it.
6. Cross-embodiment & learning-from-human-video diffusion
Data scarcity pushes several papers to learn diffusion policies from heterogeneous embodiments and human video. X-Diffusion exploits the forward diffusion process: a classifier finds the minimum indistinguishability step and human actions are mixed in only at noise levels where they are indistinguishable from robot actions, giving coarse high-level guidance without dynamically infeasible supervision. Latent Action Diffusion learns a contrastive latent action space unifying anthropomorphic hands, a human hand, and a parallel gripper, co-training a single diffusion policy across them (+25.3% success). Masquerade edits in-the-wild egocentric human video (inpaint arms, overlay a rendered robot) and fine-tunes a diffusion-policy head on ~50 robot demos per task (5–6× over baselines). Inference-Stage Adaptation-Projection adapts a trained diffusion policy to unseen manipulators at inference. This sub-trend treats embodiment gap as a noise-level / latent-space alignment problem.
7. Robustness, safety & long-horizon recovery
As diffusion policies move toward deployment, several papers add guarantees and recovery. MoE-DP inserts a Mixture-of-Experts layer between encoder and denoiser, yielding interpretable skill decomposition and a 36% relative robustness gain under disturbances on long-horizon tasks. Path-Consistent Safety Filtering wraps a diffusion policy in a safety filter that preserves path consistency; SafeFlowMPC fuses learned policies with predictive safe trajectory planning; Diffusion Stabilizer Policy targets surgical-robot stability. Human-in-the-loop entries — HITL-D (diffusion-assisted shared control) and "Uncertainty Comes for Free" (diffusion's own denoising variance as an intervention signal) — turn the generative model's uncertainty into a deployment-safety lever. Compose by Focus and FUNCanon target compositional / long-horizon execution via skill primitives.
8. Diffusion beyond tabletop manipulation: navigation, multi-agent, dexterous & driving
The recipe travels. NavDP is a sim-to-real navigation diffusion policy with a privileged-information critic (363 km / 1244 scenes of sim data; zero-shot to quadruped, wheeled, humanoid). Ventura adapts image-diffusion models for task-conditioned navigation. MIMIC-D does decentralized multi-agent diffusion (CTDE) for implicit coordination (95% on a real bimanual basket-lift from 16 demos), and ADM-DP fuses vision-tactile-graph modalities for multi-agent manipulation. Physics-Informed Diffusion Mamba Transformer brings diffusion planning to real-world driving. Application-specific diffusion policies span surgery (Diffusion Stabilizer), assistive dressing (Diffusion Policy for Robot-Assisted Dressing), food (GRITS, SCOOP'D), throwing (DA-MMP), and humanoid skills (MAKP kicking, Unified Fall-Safety). This breadth is the clearest signal that diffusion/flow is now the default continuous-action IL backbone, not a niche.
Standout deep-dives
Seven papers with confirmed arXiv records and concrete reported numbers. (Numbers below are as stated by the authors; this cluster's RL-fine-tuning cousin is Flow Policy Optimization (FPO), and the backbone taxonomy is in the VLA Architectures review.)
-
SeFA-Policy — Selective Flow Alignment (arXiv 2511.08583). A rectified-flow visuomotor policy with a consistency-correction step that uses expert demos to re-align one-step-generated actions with the current observation while preserving multimodality. Reported to surpass SOTA diffusion/flow policies while reducing inference latency by over 98%. The cleanest answer in the cluster to "one-step flow loses accuracy."
-
Mini Diffuser — two-level mini-batch training (arXiv 2505.09430). Exploits the asymmetry between action-diffusion and image-diffusion by pairing many noised action samples with each vision-language condition. On RLBench it reaches 95% of SOTA multi-task diffusion-policy performance using ~5% of the training time and ~7% of the memory — an order-of-magnitude training-cost cut, the standout efficiency result here.
-
X-Diffusion — cross-embodiment human demos (arXiv 2511.04671). Trains a classifier to detect the minimum indistinguishability step in the forward diffusion process; human actions are injected into training only at noise levels where the classifier cannot tell human from robot, so feasible motions give low-level supervision and mismatched ones give only coarse high-level guidance. A principled noise-level treatment of the embodiment gap.
-
Latent Action Diffusion (arXiv 2506.14608). Learns a contrastive latent action space unifying anthropomorphic hands, a human hand, and a parallel-jaw gripper, then co-trains a single diffusion policy across them for multi-robot control with up to +25.3% manipulation success from cross-embodiment transfer.
-
NavDP — sim-to-real navigation diffusion (arXiv 2505.08712). A transformer that jointly diffuses candidate trajectories and learns a critic from simulator privileged information to select among them. Trained purely in sim (~2,500 trajectories/GPU/day; 363.2 km across 1244 scenes), it transfers zero-shot to quadruped, wheeled, and humanoid robots — the cluster's strongest evidence that diffusion policies scale beyond tabletop.
-
MoE-DP — MoE-enhanced diffusion policy (arXiv 2511.05007). Inserts a Mixture-of-Experts layer between the visual encoder and the denoiser; experts specialize to semantic primitives (approach, grasp), giving an interpretable skill decomposition and a 36% average relative robustness improvement under disturbances on six long-horizon tasks.
-
Diffusion Trajectory-Guided Policy (DTP) (arXiv 2502.10040). Two stages: a generative VLM diffuses a 2D guidance trajectory, then the imitation policy is refined with it. Outperforms SOTA by 25% on CALVIN from scratch (no external pretraining), exemplifying the "diffuse a spatial intermediate, condition the action head" pattern.
Complete paper list (46)
| Code | Title | arXiv |
|---|---|---|
| ThBT2.9 | Factorizing Diffusion Policies for Observation Modality Prioritization | 2509.16830 |
| ThI1I.11 | Motion before Action: Diffusing Object Motion As Manipulation Condition | 2411.09658 |
| ThI1I.127 | Compose by Focus: Scene Graph-Based Atomic Skills | 2509.16053 |
| ThI1I.136 | MIMIC-D: Multi-Modal Imitation for Multi-Agent Coordination with Decentralized Diffusion Policies | 2509.14159 |
| ThI1I.257 | Conditional Flow-VAE for Safety-Critical Traffic Scenario Generation | 2605.04366 |
| ThI1I.282 | PPGuide: Steering Diffusion Policies with Performance Predictive Guidance | 2603.10980 |
| ThI1I.285 | Prepare before You Act: Learning from Humans to Rearrange Initial States | 2509.18043 |
| ThI1I.367 | S²-Diffusion: Generalizing from Instance-Level to Category-Level Skills in Robot Manipulation | 2502.09389 |
| ThI1I.379 | Mini Diffuser: Fast Multi-Task Diffusion Policy Training Using Two-Level Mini-Batches | 2505.09430 |
| ThI1I.58 | RoboMatch: A Unified Mobile-Manipulation Teleoperation Platform with Auto-Matching Network Architecture for Long-Horizon Tasks | 2509.08522 |
| ThI1I.63 | DynaFlow: Dynamics-Embedded Flow Matching for Physically Consistent Motion Generation from State-Only Demonstrations | 2509.19804 |
| ThI1I.88 | Physically-Based Lighting Generation for Robotic Manipulation | 2508.01442 |
| ThI2I.122 | HITL-D: Human in the Loop Diffusion Assisted Shared Control | 2605.21460 |
| ThI2I.161 | DA-MMP: Learning Coordinated and Accurate Throwing with Dynamics-Aware Motion Manifold Primitives | 2509.23721 |
| ThI2I.262 | WorldPlanner: Monte Carlo Tree Search and MPC with Action-Conditioned Visual World Models | 2511.03077 |
| ThI2I.281 | X-Diffusion: Training Diffusion Policies on Cross-Embodiment Human Demonstrations | 2511.04671 |
| ThI2I.311 | Diffusion Stabilizer Policy for Automated Surgical Robot Manipulations | 2503.01252 |
| ThI2I.71 | SafeFlowMPC: Predictive and Safe Trajectory Planning for Robot Manipulators with Learning-Based Policies | 2602.12794 |
| ThI2LB.10 | Diffusion Policy for Robot-Assisted Dressing with Moving Human Arms | — |
| TuAT1.1 | GRITS: A Spillage-Aware Guided Diffusion Policy for Robot Food Scooping Tasks | 2510.00573 |
| TuAT1.4 | Do You Know Where Your Camera Is? View-Invariant Policy Learning with Camera Conditioning | 2510.02268 |
| TuBT1.2 | Uncertainty Comes for Free: Human-In-The-Loop Policies with Diffusion Models | 2503.01876 |
| TuI1I.104 | Scaling Single Human Demonstrations for Imitation Learning Using Generative Foundational Models | 2602.12734 |
| TuI1I.170 | Joint Flow Trajectory Optimization for Feasible Robot Motion Generation from Video Demonstrations | 2509.20703 |
| TuI1I.191 | From Demonstrations to Safe Deployment: Path-Consistent Safety Filtering for Diffusion Policies | 2511.06385 |
| TuI1I.215 | FUNCanon: Learning Pose-Aware Action Primitives Via Functional Object Canonicalization for Generalizable Robotic Manipulation | 2509.19102 |
| TuI1I.259 | Physics-Informed Diffusion Mamba Transformer for Real-World Driving | 2602.00808 |
| TuI1I.299 | Ventura: Adapting Image Diffusion Models for Unified Task Conditioned Navigation | 2510.01388 |
| TuI1I.63 | Latent Action Diffusion for Cross-Embodiment Manipulation | 2506.14608 |
| TuI1I.78 | Seeing Motion, Generating Action: Explicit Motion-Aware Policy for Robotic Action Generation | — |
| TuI2I.158 | MAKP: Multi-Mode Accurate Kicking Policy for Humanoid Robots | — |
| TuI2I.160 | NavDP: Learning Sim-To-Real Navigation Diffusion Policy with Privileged Information Guidance | 2505.08712 |
| TuI2I.45 | Inference-Stage Adaptation-Projection Strategy Adapts Diffusion Policy to Cross-Manipulators Scenarios | 2509.11621 |
| TuI2I.63 | SCOOP'D: Learning Mixed-Liquid-Solid Scooping Via Sim2Real Generative Policy | 2510.11566 |
| WeBT2.3 | Hybrid Diffusion Policies with Projective Geometric Algebra for Efficient Robot Manipulation Learning | 2507.05695 |
| WeI1I.114 | Unified Humanoid Fall-Safety Policy from a Few Demonstrations | 2511.07407 |
| WeI1I.153 | MoE-DP: An MoE-Enhanced Diffusion Policy for Robust Long-Horizon Robotic Manipulation with Skill Decomposition and Failure Recovery | 2511.05007 |
| WeI1I.192 | DRAW2ACT: Turning Depth-Encoded Trajectories into Robotic Demonstration Videos | 2512.14217 |
| WeI1I.211 | ADM-DP: Adaptive Dynamic Modality Diffusion Policy through Vision-Tactile-Graph Fusion for Multi-Agent Manipulation | 2602.21622 |
| WeI1I.222 | Masquerade: Learning from In-The-Wild Human Videos Using Data-Editing | 2508.09976 |
| WeI1I.23 | DISCO: Language-Guided Manipulation with Diffusion Policies and Constrained Inpainting | 2406.09767 |
| WeI1I.27 | Motion Manifold Flow Primitives for Task-Conditioned Trajectory Generation under Complex Task-Motion Dependencies | 2407.19681 |
| WeI2I.132 | Disentangled Point Diffusion for Precise Object Placement | — |
| WeI2I.299 | SeFA-Policy: Fast and Accurate Visuomotor Policy Learning with Selective Flow Alignment | 2511.08583 |
| WeI2I.311 | Physically-Grounded Data Generation Via Video Diffusion Models | — |
| WeI2I.340 | Diffusion Trajectory-Guided Policy for Long-Horizon Robot Manipulation | 2502.10040 |
Related
← Back to ICRA-2026-VLA-Manipulation-Survey