ICRA 2026 Topic Imitation - Heungwoo/research GitHub Wiki

ICRA 2026 β€” Imitation Learning & Learning from Demonstration (Topic Analysis)

Venue: IEEE ICRA 2026 Β· Vienna, Austria Β· June 1–5, 2026 Compiled against the official PaperCept program. This page covers the 116 papers whose primary framing is imitation learning (IL) / learning-from-demonstration (LfD) for manipulation and related embodied control. Diffusion- and flow-policy papers have their own page and are excluded here unless IL is the dominant theme; VLA-architecture papers are tracked in the ICRA 2026 Survey.

Overview

Imitation learning is the dominant manipulation-learning paradigm at ICRA 2026. The program's topic-keyword counts put Imitation Learning at 197 and Learning from Demonstration at 133 β€” together the largest learning cluster, ahead of Reinforcement Learning's 254 only because RL spans locomotion and driving as well. This page groups the 116 papers where IL/LfD is the organizing principle rather than a downstream consumer of a new policy class.

What is new in 2026, relative to the diffusion-policy gold rush of 2023–2024, is that the field has visibly shifted from "fit a policy to demos" toward the demonstrations themselves: how to collect them cheaply (robot-free, VR, smartphone, human-video), how to curate and scale them (influence functions, self-supervised filtering, scaling laws), and how to get more out of fewer (augmentation, retrieval, non-expert/sub-optimal data, cross-embodiment transfer). A second clear thread is the resurgence of classical structured LfD β€” dynamical systems, movement primitives, Riemannian and equivariant formulations β€” now fused with neural function approximation and safety certificates. A third is IL meets foundation models: VLMs supplying saliency/where-to-attend, LLMs auto-augmenting demos, and action-reasoning models that plan in space before acting.

Sub-trends

1. Data collection & teleoperation systems

The single most active cluster. The trend is toward robot-free and portable capture so demonstration collection scales beyond a teleoperated arm. ActiveUMI (WeI2I.276, arXiv 2510.01607) pairs a portable VR kit with sensorized controllers that mirror the robot end-effectors and records the operator's deliberate head motions, learning the visual-attention↔manipulation link for bimanual tasks (70% avg vs 42% fixed-camera baseline). TWIST2 (WeBT3.1, arXiv 2511.02832) is a mocap-free, portable whole-body humanoid collection rig (PICO4U VR + a ~$250 2-DoF robot neck) that gathers 100 demos in 15 min. EgoMI (TuI2I.259, arXiv 2511.00153) captures synchronized end-effector + active-head trajectories from egocentric human demos and retargets to semi-humanoids. COBALT (ThI1I.238) democratizes capture via cloud-based smartphone teleoperation; RAMPA (TuI2I.3) does programming-by-demonstration through AR/XR; AVR (TuI1I.239) adds active viewpoint + focal-length optimization to teleop. Supporting infrastructure includes Robot Control Stack (TuI1I.250, a lean ecosystem for learning at scale) and Hierarchical Grid-Based Sensor Pose Extraction (TuI2LB.10) for demonstration-dataset generation.

2. Demonstration scaling, curation & augmentation

A data-centric counter-movement to "just collect more." SCIZOR (ThBT2.3, arXiv 2505.22626) self-supervisedly removes sub-optimal and redundant state-action pairs (task-progress predictor + dedup), giving +15.4% avg with less data. Quality Over Quantity (ThI2I.153) curates demos via influence functions; Beyond the Majority (WeI1I.286) tackles the long-tail distribution of training demos. On the synthesis side, LLM Trainer (TuI1I.336) turns as few as one human demo into a large dataset via LLM-driven augmentation, SoftMimicGen (WeI2I.305) generates data for deformable-object manipulation, and MOVE (ThI1I.192) is a motion-based collection paradigm for spatial generalization. EDAIL (WeI2I.115) uses exploration-driven augmentation inside adversarial IL. Scaling-law work is explicit: The Curse of Precision (ThI2I.162) studies data-vs-precision for high-precision assembly, and Data Scaling Laws for IL-Based End-to-End Driving (ThI2I.259) extends the relationship to autonomous driving.

3. Learning from human video & play (actionless data)

Action-free video is the cheapest demonstration source, and several papers learn from it directly. AMPLIFY (TuI2I.286, arXiv 2506.14198) encodes visual dynamics into discrete motion tokens from keypoint trajectories, training a forward model on action-free video and a small inverse model on labeled data (first zero-action-data LIBERO generalization). ViSA-Flow (ThI2I.98, arXiv 2505.01288) defines semantic action flow as a video-invariant intermediate and pre-trains on human-object-interaction video. MimicDroid (WeBT3.3, arXiv 2509.09769) does in-context learning for humanoids purely from unlabeled human play videos. EasyMimic (WeI2I.193) is a low-cost framework from human videos, HAND (WeI2I.182) retrieves behavior from human hand-path demos, and MotionTrans (WeI1I.96, arXiv 2509.17759) cotrains human VR data with robot data to transfer motion-level skills (9 zero-shot tasks, +40% finetuning).

4. Classical structured LfD: dynamical systems, movement primitives, manifolds

A strong, distinctly-ICRA revival. Dynamical-system LfD: ODE-driven diffeomorphic mappings (TuI1I.322), constraint-aware DS from human demos (TuI1I.333), Safe and Stable NN Dynamical Systems (SΒ²-NNDS, TuI1I.410), and behavior-controllable stable dynamics on Riemannian configuration manifolds (ThI2I.418). Movement primitives & DMPs: DMPs + control barrier functions for constrained trajectories (WeI1I.30), SafeDMPs with formal safety for HRI (ThI1I.273), DMP + affordance templates for contact tasks (TuI1LB.2), differentiable Motion Manifold Primitives under kinodynamic constraints (ThI1I.31), and the synthesis paper From Movement Primitives to Distance Fields to Dynamical Systems (WeI2I.27). Riemannian / geometric methods: LfD over Riemannian manifolds via Neural ODEs (TuI2LB.24) and Riemannian Time Warping for sequence alignment in curved spaces (WeI1I.358). Also task-parameterized motion with time-sensitive constraints (TuI1I.363), an alignment-based approach to learning motions (TuI2I.358), and a nonparametric IL method (ThI2I.170).

5. Generalization, robustness & data quality at test time

Making IL survive distribution shift. AutoFocus-IL (ThI1I.323) uses VLM saliency maps to steer attention to task-relevant features without extra annotation; PEEK (WeI1I.279, arXiv 2509.18282) offloads where/what reasoning to a VLM via point-based overlays (up to 41Γ— zero-shot gain for a sim-only 3D policy). The Temporal Trap (WeI1I.61) diagnoses temporal entanglement in pre-trained visual representations; Unsupervised Domain Adaptation (WeI1I.218) handles visual perturbations; Task Robustness via Re-Labelling (WeI1I.199) fixes instruction-following. Learning from imperfect supervisors: Using Non-Expert Data to Robustify IL via offline RL (ThI2I.234), Beyond the Teacher (WeI2I.291) and Better Than Diverse Demonstrators (WeBT2.7) both exploit mixed-skill / sub-optimal / heterogeneous demos. Failure Identification in IL (WeI2I.172) filters failures statistically and semantically; BOSS (ThI1I.17) is a benchmark for observation-space shift in long-horizon tasks.

6. Cross-embodiment & transfer learning

Cross-Embodiment Transfer via Behavior-Aligned Representations (ThI2I.283) and Cross-Embodiment Imitation with a unified latent space across single-arm/dual-arm/legged humanoids (ThI2I.404) tackle the embodiment gap head-on. One-Shot Cross-Geometry Skill Transfer (ThI1I.122) generalizes a skill across object shapes via part decomposition; Efficient Learning of Object Placement (TuI2I.367) does intra-category transfer. Spatio-Temporal Motion Retargeting (WeI2I.21) transfers agile motions to quadrupeds, and D-CLING (TuI2I.145) depth-conditions and prior-preserves a navigation foundation model during in-domain finetuning.

7. Long-horizon, hierarchical, lifelong & skill-discovery IL

Decomposition for scale. Playbook (TuI2I.38) does scalable discrete skill discovery from unstructured datasets; lifelong IL appears in T2S (tokenized skill scaling, TuI1I.72) and SPREAD (subspace-representation distillation, TuI2I.102). Hierarchy + imagination: From Dream to Action (TuI1I.146) couples a 3D world model with hierarchical policy learning. Long-horizon teaching: Sequentially Teaching Sequential Tasks (ST)Β² (WeI1I.424). Memory/history modeling: MTIL (TuI2I.348) encodes full history with Mamba, and History-Aware Visuomotor Policy (TuI2I.168) uses point tracking to break the Markov assumption.

8. Action representation, tokenization & reasoning

MolmoAct (TuI1I.327, arXiv 2508.07917) introduces Action Reasoning Models β€” depth-aware perception tokens β†’ editable trajectory traces β†’ action tokens (70.5% SimplerEnv zero-shot). Fast ECoT (ThI2I.60, arXiv 2506.07639) caches/reuses embodied chain-of-thought to cut latency. MLA (TuI2I.195) is a multisensory language-action model that also forecasts; Cross-Modal Instructions (WeI2I.294) and TeNet (WeI1LB.19, text-to-network compact policy synthesis) explore alternative conditioning. CLAR (TuI1I.150) learns 3D representations by fusing masked reconstruction with contrastive alignment. ABPolicy (ThI2I.158) uses asynchronous B-spline action representation for smooth real-time control; Viper (TuI2I.139) targets verifiable, efficient IL policies.

9. IL βˆͺ RL: residual, adversarial, IRL & self-improvement

The boundary with RL is heavily trafficked. Demonstration-augmented RL: Rainbow-DemoRL (ThI1I.299), GRAPE (ThI1I.94, arXiv 2411.19309; trajectory-level preference alignment, +51.8%/+58.2%). Adversarial IL: GAIL for robot swarms (ThI1I.214), EDAIL (WeI2I.115). Inverse RL / reward learning: Masked IRL (ThI2I.233, LLM-guided reward disambiguation), ILCL (ThI2I.353, inverse logic-constraint learning), neural CBFs from demos via inverse constraint learning (ThI2I.123), MEDIRL traversability cost maps (TuI1LB.3). Self-improvement & residual: SOE (TuI1I.134, on-manifold exploration), Robust Online Residual Refinement via Koopman dynamics (WeI1I.134), Unlocking SAC for IL (TuI2I.270), and Quasimetric Decision Transformers (TuI2I.328).

10. Domains: locomotion, driving, navigation, medical, HRI

IL/LfD spreads well beyond tabletop manipulation. Legged/locomotion: KiRAS keyframe self-imitation (ThI2I.142), quadruped walking from seconds of demos (TuI2I.118), Behavior Foundation Model for Humanoids (TuI1I.120), adaptive motion priors (WeI1I.235), RPG humanoid fighting (TuI2I.280), bicycle flip stunts (ThI2I.240). Driving: CAPS priority sampling (ThI1I.202), STAGE personalized style (ThI2I.410), CoPlanner (TuI1I.75), Mimir (TuI2I.388), location-specific behavior priors (TuBT3.5), and the planning study Perfect Prediction or Plenty of Proposals? (TuI2I.72). Navigation: The One RING (ThAT1.4, arXiv 2412.14401; embodiment-agnostic, trained over 1M randomized embodiments), Narrate2Nav (ThI2I.202), ECAHD (ThI2I.280), social navigation from +/- demos (WeI2I.133), sidewalk autopilot (ThI1I.326), synthetic-vs-real navigation data (WeI1I.74), NavGSim (TuI2I.332). Medical: UltraHiT carotid ultrasonography (ThI2I.188), 3D-ultrasound servoing (WeAT2.2), needle insertion by example (ThI2I.160), surgical sim-to-real visuomotor (WeI1I.231). HRI / human factors: Training Humans to Teach Robots (TuI1I.108), passivity-based trajectory/force arbitration (TuI1I.98), soft-rigid feeding system (ThI2I.348).

11. Evaluation, benchmarks & infrastructure

A Taxonomy for Evaluating Generalist Robot Manipulation Policies (ThI1I.435) proposes a framework for quantifying generalization; BOSS (ThI1I.17) benchmarks observation-space shift. Real-Is-Sim (ThI2I.56) bridges sim-to-real with a dynamic digital twin inside behavior-cloning pipelines. Multi-agent IL: R2BC (WeI1I.71, multi-agent from single-agent demos) and behavior models for multi-agent driving simulation (WeI2I.438). Imitation-BT (WeI2I.439) distills RL agents into behavior trees for interpretability.

Standout deep-dives

All arXiv IDs and headline numbers below were confirmed via web search against the papers' own abstracts/pages.

  • The One RING (ThAT1.4, arXiv 2412.14401, UW / AI2) β€” an embodiment-agnostic indoor navigation generalist trained entirely in simulation over 1M randomized embodiments (camera params, collider sizes, center of rotation). Achieves 72.1% sim and 78.9% real-world avg object-goal success across multiple unseen platforms (Stretch RE-1, LoCoBot, Go1). A clean demonstration that massive embodiment randomization can replace per-robot retraining.

  • SCIZOR (ThBT2.3, arXiv 2505.22626, UT-Austin RPL) β€” self-supervised data curation that filters sub-optimal (low task-progress) and redundant (deduplicated state-action) samples. Yields +15.4% avg across OXE, RoboMimic, and real Sirius-Fleet while using less data β€” the strongest case this cycle that demonstration quality beats quantity.

  • MolmoAct (TuI1I.327, arXiv 2508.07917, AI2) β€” Action Reasoning Models: depth-aware perception tokens β†’ editable visual trajectory traces β†’ action tokens. 70.5% zero-shot on SimplerEnv visual matching (beats Ο€β‚€ and GR00T N1.5), 86.6% avg LIBERO, +10%/+22.7% real single-arm/bimanual over Ο€β‚€-FAST. Fully open (weights, data, reasoning dataset).

  • AMPLIFY (TuI2I.286, arXiv 2506.14198) β€” separates visual-motion prediction from action inference using discrete motion tokens from keypoints; trains the forward model on action-free video, the inverse model on few labeled examples. Up to 3.7Γ— better dynamics MSE, 1.2–2.2Γ— downstream gains in low-data, and the first LIBERO generalization from zero in-distribution action data.

  • ViSA-Flow (ThI2I.98, arXiv 2505.01288, U-Michigan / KTH) β€” semantic action flow as a visual-difference-invariant intermediate, pretrained on large-scale human-object-interaction video then adapted with few robot demos. SOTA on CALVIN and real tasks, especially in low-data regimes.

  • GRAPE (ThI1I.94, arXiv 2411.19309, UNC et al.) β€” trajectory-level preference alignment for VLAs that models reward from both successful and failed rollouts, with spatiotemporal constraints from a VLM. +51.8% in-domain / +58.2% unseen success, plus βˆ’37.4% collisions when aligned for safety.

  • TWIST2 (WeBT3.1, arXiv 2511.02832, Stanford) β€” mocap-free, portable whole-body humanoid teleop/collection (PICO4U VR + ~$250 2-DoF robot neck for egocentric vision), 100 demos in 15 min at near-100% success, with a hierarchical visuomotor policy. Fully open-sourced β€” addresses humanoid robotics' missing data-collection infrastructure.

  • ActiveUMI (WeI2I.276, arXiv 2510.01607) β€” robot-free in-the-wild bimanual data via a portable VR kit whose sensorized controllers mirror the robot's end-effectors, crucially capturing the operator's active head/gaze motion. 70% avg task success vs 42% (fixed head cam) / 26% (wrist-only), showing active perception is learnable from human demos.

Complete paper list (116)

Papers are listed under their sub-theme. The "Code/links" column marks an open-source release or project page where a search surfaced one (βœ“). The arXiv column lists IDs confirmed via web search; blank means none was located (not that none exists).

Data collection & teleoperation systems

Code Title Paper ID arXiv
βœ“ ActiveUMI: Robotic Manipulation with Active Perception from Robot-Free Human Demonstrations WeI2I.276 2510.01607
βœ“ TWIST2: Scalable, Portable, and Holistic Humanoid Data Collection System WeBT3.1 2511.02832
βœ“ EgoMI: Learning Active Vision and Whole-Body Manipulation from Egocentric Human Demonstrations TuI2I.259 2511.00153
COBALT: Crowdsourcing Robot Learning Via Cloud-Based Teleoperation with Smartphones ThI1I.238
RAMPA: Robotic Augmented Reality for Machine Programming by DemonstrAtion TuI2I.3 2410.13412
AVR: Active Vision-Driven Precise Robot Manipulation with Viewpoint and Focal Length Optimization TuI1I.239 2503.01439
Robot Control Stack: A Lean Ecosystem for Robot Learning at Scale TuI1I.250 2509.14932
Hierarchical Grid-Based Sensor Pose Extraction for Demonstration Dataset Generation TuI2LB.10

Demonstration scaling, curation & augmentation

Code Title Paper ID arXiv
βœ“ SCIZOR: A Self-Supervised Approach to Data Curation for Large-Scale Imitation Learning ThBT2.3 2505.22626
Quality Over Quantity: Demonstration Curation Via Influence Functions for Data-Centric Robot Learning ThI2I.153 2603.09056
Beyond the Majority: Long-Tail Imitation Learning for Robotic Manipulation WeI1I.286 2602.06512
LLM Trainer: Automated Robotic Data Generating Via Demonstration Augmentation Using LLMs TuI1I.336 2509.20070
SoftMimicGen: A Data Generation System for Scalable Robot Learning in Deformable Object Manipulation WeI2I.305 2603.25725
MOVE: A Simple Motion-Based Data Collection Paradigm for Spatial Generalization in Robotic Manipulation ThI1I.192 2512.04813
The Curse of Precision: A Data Scaling Law for High-Precision Robotic Manipulation ThI2I.162
Data Scaling Laws for Imitation Learning-Based End-To-End Autonomous Driving ThI2I.259 2412.02689

Learning from human video & play (actionless data)

Code Title Paper ID arXiv
βœ“ AMPLIFY: Actionless Motion Priors for Robot Learning from Videos TuI2I.286 2506.14198
βœ“ ViSA-Flow: Accelerating Robot Skill Learning Via Large-Scale Video Semantic Action Flow ThI2I.98 2505.01288
βœ“ MimicDroid: In-Context Learning for Humanoid Robot Manipulation from Human Play Videos WeBT3.3 2509.09769
βœ“ MotionTrans: Human VR Data Enable Motion-Level Learning for Robotic Manipulation Policies WeI1I.96 2509.17759
EasyMimic: A Low-Cost Framework for Robot Imitation Learning from Human Videos WeI2I.193 2602.11464
HAND Me the Data: Fast Robot Adaptation Via Hand Path Retrieval WeI2I.182 2505.20455

Classical structured LfD: dynamical systems, movement primitives, manifolds

Code Title Paper ID arXiv
Learning Dynamical System-Based Robot Motions from Demonstrations Via ODE-Driven Diffeomorphic Mappings TuI1I.322
Learning Constraint-Aware Dynamical Systems from Human Demonstrations for Constrained Manipulation Tasks TuI1I.333
Safe and Stable Neural Network Dynamical Systems for Robot Motion Planning TuI1I.410 2511.20593
Behavior-Controllable Stable Dynamics Models on Riemannian Configuration Manifolds ThI2I.418
Dynamic Movement Primitives with Control Barrier Functions for Constrained Trajectory Planning WeI1I.30
SafeDMPs: Integrating Formal Safety with DMPs for Adaptive HRI ThI1I.273 2603.29708
Learning Contact Tasks Skills Based on DMP and Affordance Templates TuI1LB.2
Differentiable Motion Manifold Primitives for Reactive Motion Generation under Kinodynamic Constraints ThI1I.31 2410.12193
From Movement Primitives to Distance Fields to Dynamical Systems WeI2I.27 2504.09705
Learning from Demonstrations Over Riemannian Manifolds Using Neural ODEs TuI2LB.24
Riemannian Time Warping: Multiple Sequence Alignment in Curved Spaces WeI1I.358 2506.01635
Task-Parameterized Motion Learning with Time-Sensitive Constraints TuI1I.363 2312.03506
An Alignment-Based Approach to Learning Motions from Demonstrations TuI2I.358 2511.14988
A Computationally Efficient Nonparametric Approach for Robot Imitation Learning ThI2I.170

Generalization, robustness & data quality at test time

Code Title Paper ID arXiv
βœ“ PEEK: Guiding and Minimal Image Representations for Zero-Shot Generalization of Robot Manipulation Policies WeI1I.279 2509.18282
AutoFocus-IL: VLM-Based Saliency Maps for Data-Efficient Visual Imitation Learning without Extra Human Annotations ThI1I.323 2511.18617
The Temporal Trap: Entanglement in Pre-Trained Visual Representations for Visuomotor Policy Learning WeI1I.61 2502.03270
Unsupervised Domain Adaptation for Robust Imitation Learning under Visual Perturbations WeI1I.218
Task Robustness Via Re-Labelling Vision-Action Robot Data WeI1I.199
Using Non-Expert Data to Robustify Imitation Learning Via Offline Reinforcement Learning ThI2I.234 2510.19495
Beyond the Teacher: Leveraging Mixed-Skill Demonstrations for Robust Imitation Learning WeI2I.291
Better Than Diverse Demonstrators: Reward Decomposition from Suboptimal and Heterogeneous Demonstrations WeBT2.7
Failure Identification in Imitation Learning Via Statistical and Semantic Filtering WeI2I.172 2604.13788
BOSS: Benchmark for Observation Space Shift in Long-Horizon Task ThI1I.17 2502.15679

Cross-embodiment & transfer learning

Code Title Paper ID arXiv
Cross-Embodiment Transfer Via Behavior-Aligned Representations ThI2I.283
Cross-Embodiment Imitation: Learning a Unified Latent Space for Multi-Robot Control ThI2I.404 2601.15419
One-Shot Cross-Geometry Skill Transfer through Part Decomposition ThI1I.122 2604.15455
Efficient Learning of Object Placement with Intra-Category Transfer TuI2I.367 2411.03408
Spatio-Temporal Motion Retargeting for Quadruped Robots WeI2I.21 2404.11557
D-CLING: Prior-Preserving Depth-Conditioned Fine-Tuning for Navigation Foundation Models TuI2I.145

Long-horizon, hierarchical, lifelong & skill-discovery IL

Code Title Paper ID arXiv
Playbook: Scalable Discrete Skill Discovery from Unstructured Datasets for Long-Horizon Decision-Making Problems TuI2I.38
T2S: Tokenized Skill Scaling for Lifelong Imitation Learning TuI1I.72 2508.01167
SPREAD: Subspace Representation Distillation for Lifelong Imitation Learning TuI2I.102 2603.08763
From Dream to Action: Hierarchical Policy Learning with 3D World Imagination for Robotic Manipulation TuI1I.146
Sequentially Teaching Sequential Tasks (ST)Β²: Teaching Robots Long-Horizon Manipulation Skills WeI1I.424 2510.21046
MTIL: Encoding Full History with Mamba for Temporal Imitation Learning TuI2I.348 2505.12410
History-Aware Visuomotor Policy Learning Via Point Tracking TuI2I.168 2509.17141

Action representation, tokenization & reasoning

Code Title Paper ID arXiv
βœ“ MolmoAct: Action Reasoning Models That Can Reason in Space TuI1I.327 2508.07917
βœ“ Fast ECoT: Efficient Embodied Chain-Of-Thought Via Thoughts Reuse ThI2I.60 2506.07639
MLA: A Multisensory Language–Action Model for Multimodal Understanding and Forecasting in Robotic Manipulation TuI2I.195 2509.26642
Cross-Modal Instructions for Robot Motion Generation WeI2I.294 2509.21107
TeNet: Text-To-Network for Compact Policy Synthesis WeI1LB.19 2601.15912
CLAR: Learning 3D Representations for Robotic Manipulation by Fusing Masked Reconstruction with Multi-Level Contrastive Alignment TuI1I.150
ABPolicy: Asynchronous B-Spline Flow Policy for Real-Time and Smooth Robotic Manipulation ThI2I.158 2602.23901
Viper: Verifiable Imitation Learning Policy for Efficient Robotic Manipulation TuI2I.139

IL βˆͺ RL: residual, adversarial, IRL & self-improvement

Code Title Paper ID arXiv
βœ“ GRAPE: Generalizing Robot Policy Via Preference Alignment ThI1I.94 2411.19309
Rainbow-DemoRL: Combining Improvements in Demonstration-Augmented Reinforcement Learning ThI1I.299 2603.27400
Generative Adversarial Imitation Learning for Robot Swarms: Learning from Human Demonstrations and Trained Policies ThI1I.214 2603.02783
EDAIL: Adversarial Imitation Learning Via Exploration-Driven Data Augmentation WeI2I.115
Masked IRL: LLM-Guided Reward Disambiguation from Demonstrations and Language ThI2I.233 2511.14565
ILCL: Inverse Logic-Constraint Learning from Temporally Constrained Demonstrations ThI2I.353 2507.11000
Learning Neural Control Barrier Functions from Expert Demonstrations Using Inverse Constraint Learning ThI2I.123 2510.21560
Learning Traversability Cost Maps with Decomposed Uncertainties Via Continuous-State MEDIRL TuI1LB.3
SOE: Sample-Efficient Robot Policy Self-Improvement Via On-Manifold Exploration TuI1I.134 2509.19292
Robust Online Residual Refinement Via Koopman-Guided Dynamics Modeling WeI1I.134 2509.12562
Unlocking the Potential of Soft Actor-Critic for Imitation Learning TuI2I.270 2509.24539
Quasimetric Decision Transformers: Enhancing Goal-Conditioned Reinforcement Learning with Structured Distance Guidance TuI2I.328
Data-Efficient Hierarchical Goal-Conditioned Reinforcement Learning Via Normalizing Flows ThI1I.203 2602.11142

Domain: locomotion & humanoid motion

Code Title Paper ID arXiv
KiRAS: Keyframe Guided Self-Imitation for Robust and Adaptive Skill Learning in Quadruped Robots ThI2I.142 2603.15179
Learning Quadruped Walking from Seconds of Demonstration TuI2I.118 2603.06961
Behavior Foundation Model for Humanoid Robots TuI1I.120 2509.13780
Adaptive Motion Priors with Constrained Optimization WeI1I.235
RPG: Robust Policy Gating for Smooth Multi-Skill Transitions in Humanoid Fighting TuI2I.280 2604.21355
Flip Stunts on Bicycle Robots Using Iterative Motion Imitation ThI2I.240 2603.27944
Teaching to Individual Needs: Bidirectional Teacher-Student Learning for Wheeled-Legged Locomotion ThI1I.201

Domain: autonomous driving

Code Title Paper ID arXiv
CAPS: Context-Aware Priority Sampling for Enhanced Imitation Learning in Autonomous Driving ThI1I.202 2503.01650
STAGE: STyle-Controllable Action GEneration for Personalized Autonomous Driving ThI2I.410
CoPlanner: An Interactive Motion Planner with Contingency-Aware Diffusion for Autonomous Driving TuI1I.75 2509.17080
Mimir: Hierarchical Goal-Driven Diffusion with Uncertainty Propagation for End-To-End Autonomous Driving TuI2I.388 2512.07130
Learning Location-Specific Latent Behavior Priors for Occupancy Prediction in Automated Driving TuBT3.5
Perfect Prediction or Plenty of Proposals? What Matters Most in Planning for Autonomous Driving TuI2I.72 2510.15505
Toward Efficient and Robust Behavior Models for Multi-Agent Driving Simulation WeI2I.438 2512.05812

Domain: navigation

Code Title Paper ID arXiv
βœ“ The One RING: A Robotic Indoor Navigation Generalist ThAT1.4 2412.14401
Narrate2Nav: Real-Time Visual Navigation with Implicit Language Reasoning in Human-Centric Environments ThI2I.202 2506.14233
ECAHD: Efficient Collision-Aware Hierarchical Diffusion Navigation ThI2I.280
Learning Social Navigation from Positive and Negative Demonstrations and Rule-Based Specifications WeI2I.133 2510.12215
Learning Sidewalk Autopilot from Multi-Scale Imitation with Corrective Behavior Expansion ThI1I.326 2603.22527
Synthetic vs. Real Training Data for Visual Navigation WeI1I.74 2509.11791
NavGSim: High-Fidelity Gaussian Splatting Simulator for Large-Scale Navigation TuI2I.332 2603.15186

Domain: medical & surgical

Code Title Paper ID arXiv
UltraHiT: A Hierarchical Transformer Architecture for Generalizable Internal Carotid Artery Robotic Ultrasonography ThI2I.188 2509.13832
Breaking the Latency Barrier: Synergistic Perception and Control for High-Frequency 3D Ultrasound Servoing WeAT2.2 2511.00983
Real-Time Robotic Needle Insertion in Deformable and Moving Structure Using Learning-By-Example Method ThI2I.160
Overcoming Imperfect Kinematics in Surgical Robotics through Sim-To-Real Visuomotor Learning WeI1I.231

Domain: HRI, human factors, contact & force

Code Title Paper ID arXiv
Training Humans to Teach Robots: Large and Lasting Skill Gains TuI1I.108
A Passivity-Based Framework for Dynamic Arbitration between Trajectory and Force Tracking Using Human Demonstration TuI1I.98
A Soft-Rigid Hybrid Robot-Assisted Feeding System with a Tendon-Driven Continuum Robot ThI2I.348
Spline-FRIDA: Towards Diverse, Humanlike Robot Painting Styles with a Sample-Efficient, Differentiable Brush Stroke Model WeI1I.380 2412.00597

Evaluation, benchmarks, multi-agent & infrastructure

Code Title Paper ID arXiv
A Taxonomy for Evaluating Generalist Robot Manipulation Policies ThI1I.435 2503.01238
Real-Is-Sim: Bridging the Sim-To-Real Gap with a Dynamic Digital Twin ThI2I.56 2504.03597
R2BC: Multi-Agent Imitation Learning from Single-Agent Demonstrations WeI1I.71 2510.18085
Imitation-BT: Automating Behavior Tree Generation by Echoing Reinforcement Learning Agents WeI2I.439

Other / cross-cutting IL methods

Code Title Paper ID arXiv
Flow-Enabled Generalization to Human Demonstrations in Few-Shot Imitation Learning TuI2I.53 2602.10594
Real-Time Generation of Near-Minimum-Energy Trajectories Via Constraint-Informed Residual Learning WeI2I.383 2501.09450
Learning Multiple Initial Solutions to Optimization Problems WeI1I.45 2411.02158

Related

← Back to ICRA-2026-VLA-Manipulation-Survey