ICRA 2026 Topic Imitation - Heungwoo/research GitHub Wiki
ICRA 2026 β Imitation Learning & Learning from Demonstration (Topic Analysis)
Venue: IEEE ICRA 2026 Β· Vienna, Austria Β· June 1β5, 2026 Compiled against the official PaperCept program. This page covers the 116 papers whose primary framing is imitation learning (IL) / learning-from-demonstration (LfD) for manipulation and related embodied control. Diffusion- and flow-policy papers have their own page and are excluded here unless IL is the dominant theme; VLA-architecture papers are tracked in the ICRA 2026 Survey.
Overview
Imitation learning is the dominant manipulation-learning paradigm at ICRA 2026. The program's topic-keyword counts put Imitation Learning at 197 and Learning from Demonstration at 133 β together the largest learning cluster, ahead of Reinforcement Learning's 254 only because RL spans locomotion and driving as well. This page groups the 116 papers where IL/LfD is the organizing principle rather than a downstream consumer of a new policy class.
What is new in 2026, relative to the diffusion-policy gold rush of 2023β2024, is that the field has visibly shifted from "fit a policy to demos" toward the demonstrations themselves: how to collect them cheaply (robot-free, VR, smartphone, human-video), how to curate and scale them (influence functions, self-supervised filtering, scaling laws), and how to get more out of fewer (augmentation, retrieval, non-expert/sub-optimal data, cross-embodiment transfer). A second clear thread is the resurgence of classical structured LfD β dynamical systems, movement primitives, Riemannian and equivariant formulations β now fused with neural function approximation and safety certificates. A third is IL meets foundation models: VLMs supplying saliency/where-to-attend, LLMs auto-augmenting demos, and action-reasoning models that plan in space before acting.
Sub-trends
1. Data collection & teleoperation systems
The single most active cluster. The trend is toward robot-free and portable capture so demonstration collection scales beyond a teleoperated arm. ActiveUMI (WeI2I.276, arXiv 2510.01607) pairs a portable VR kit with sensorized controllers that mirror the robot end-effectors and records the operator's deliberate head motions, learning the visual-attentionβmanipulation link for bimanual tasks (70% avg vs 42% fixed-camera baseline). TWIST2 (WeBT3.1, arXiv 2511.02832) is a mocap-free, portable whole-body humanoid collection rig (PICO4U VR + a ~$250 2-DoF robot neck) that gathers 100 demos in 15 min. EgoMI (TuI2I.259, arXiv 2511.00153) captures synchronized end-effector + active-head trajectories from egocentric human demos and retargets to semi-humanoids. COBALT (ThI1I.238) democratizes capture via cloud-based smartphone teleoperation; RAMPA (TuI2I.3) does programming-by-demonstration through AR/XR; AVR (TuI1I.239) adds active viewpoint + focal-length optimization to teleop. Supporting infrastructure includes Robot Control Stack (TuI1I.250, a lean ecosystem for learning at scale) and Hierarchical Grid-Based Sensor Pose Extraction (TuI2LB.10) for demonstration-dataset generation.
2. Demonstration scaling, curation & augmentation
A data-centric counter-movement to "just collect more." SCIZOR (ThBT2.3, arXiv 2505.22626) self-supervisedly removes sub-optimal and redundant state-action pairs (task-progress predictor + dedup), giving +15.4% avg with less data. Quality Over Quantity (ThI2I.153) curates demos via influence functions; Beyond the Majority (WeI1I.286) tackles the long-tail distribution of training demos. On the synthesis side, LLM Trainer (TuI1I.336) turns as few as one human demo into a large dataset via LLM-driven augmentation, SoftMimicGen (WeI2I.305) generates data for deformable-object manipulation, and MOVE (ThI1I.192) is a motion-based collection paradigm for spatial generalization. EDAIL (WeI2I.115) uses exploration-driven augmentation inside adversarial IL. Scaling-law work is explicit: The Curse of Precision (ThI2I.162) studies data-vs-precision for high-precision assembly, and Data Scaling Laws for IL-Based End-to-End Driving (ThI2I.259) extends the relationship to autonomous driving.
3. Learning from human video & play (actionless data)
Action-free video is the cheapest demonstration source, and several papers learn from it directly. AMPLIFY (TuI2I.286, arXiv 2506.14198) encodes visual dynamics into discrete motion tokens from keypoint trajectories, training a forward model on action-free video and a small inverse model on labeled data (first zero-action-data LIBERO generalization). ViSA-Flow (ThI2I.98, arXiv 2505.01288) defines semantic action flow as a video-invariant intermediate and pre-trains on human-object-interaction video. MimicDroid (WeBT3.3, arXiv 2509.09769) does in-context learning for humanoids purely from unlabeled human play videos. EasyMimic (WeI2I.193) is a low-cost framework from human videos, HAND (WeI2I.182) retrieves behavior from human hand-path demos, and MotionTrans (WeI1I.96, arXiv 2509.17759) cotrains human VR data with robot data to transfer motion-level skills (9 zero-shot tasks, +40% finetuning).
4. Classical structured LfD: dynamical systems, movement primitives, manifolds
A strong, distinctly-ICRA revival. Dynamical-system LfD: ODE-driven diffeomorphic mappings (TuI1I.322), constraint-aware DS from human demos (TuI1I.333), Safe and Stable NN Dynamical Systems (SΒ²-NNDS, TuI1I.410), and behavior-controllable stable dynamics on Riemannian configuration manifolds (ThI2I.418). Movement primitives & DMPs: DMPs + control barrier functions for constrained trajectories (WeI1I.30), SafeDMPs with formal safety for HRI (ThI1I.273), DMP + affordance templates for contact tasks (TuI1LB.2), differentiable Motion Manifold Primitives under kinodynamic constraints (ThI1I.31), and the synthesis paper From Movement Primitives to Distance Fields to Dynamical Systems (WeI2I.27). Riemannian / geometric methods: LfD over Riemannian manifolds via Neural ODEs (TuI2LB.24) and Riemannian Time Warping for sequence alignment in curved spaces (WeI1I.358). Also task-parameterized motion with time-sensitive constraints (TuI1I.363), an alignment-based approach to learning motions (TuI2I.358), and a nonparametric IL method (ThI2I.170).
5. Generalization, robustness & data quality at test time
Making IL survive distribution shift. AutoFocus-IL (ThI1I.323) uses VLM saliency maps to steer attention to task-relevant features without extra annotation; PEEK (WeI1I.279, arXiv 2509.18282) offloads where/what reasoning to a VLM via point-based overlays (up to 41Γ zero-shot gain for a sim-only 3D policy). The Temporal Trap (WeI1I.61) diagnoses temporal entanglement in pre-trained visual representations; Unsupervised Domain Adaptation (WeI1I.218) handles visual perturbations; Task Robustness via Re-Labelling (WeI1I.199) fixes instruction-following. Learning from imperfect supervisors: Using Non-Expert Data to Robustify IL via offline RL (ThI2I.234), Beyond the Teacher (WeI2I.291) and Better Than Diverse Demonstrators (WeBT2.7) both exploit mixed-skill / sub-optimal / heterogeneous demos. Failure Identification in IL (WeI2I.172) filters failures statistically and semantically; BOSS (ThI1I.17) is a benchmark for observation-space shift in long-horizon tasks.
6. Cross-embodiment & transfer learning
Cross-Embodiment Transfer via Behavior-Aligned Representations (ThI2I.283) and Cross-Embodiment Imitation with a unified latent space across single-arm/dual-arm/legged humanoids (ThI2I.404) tackle the embodiment gap head-on. One-Shot Cross-Geometry Skill Transfer (ThI1I.122) generalizes a skill across object shapes via part decomposition; Efficient Learning of Object Placement (TuI2I.367) does intra-category transfer. Spatio-Temporal Motion Retargeting (WeI2I.21) transfers agile motions to quadrupeds, and D-CLING (TuI2I.145) depth-conditions and prior-preserves a navigation foundation model during in-domain finetuning.
7. Long-horizon, hierarchical, lifelong & skill-discovery IL
Decomposition for scale. Playbook (TuI2I.38) does scalable discrete skill discovery from unstructured datasets; lifelong IL appears in T2S (tokenized skill scaling, TuI1I.72) and SPREAD (subspace-representation distillation, TuI2I.102). Hierarchy + imagination: From Dream to Action (TuI1I.146) couples a 3D world model with hierarchical policy learning. Long-horizon teaching: Sequentially Teaching Sequential Tasks (ST)Β² (WeI1I.424). Memory/history modeling: MTIL (TuI2I.348) encodes full history with Mamba, and History-Aware Visuomotor Policy (TuI2I.168) uses point tracking to break the Markov assumption.
8. Action representation, tokenization & reasoning
MolmoAct (TuI1I.327, arXiv 2508.07917) introduces Action Reasoning Models β depth-aware perception tokens β editable trajectory traces β action tokens (70.5% SimplerEnv zero-shot). Fast ECoT (ThI2I.60, arXiv 2506.07639) caches/reuses embodied chain-of-thought to cut latency. MLA (TuI2I.195) is a multisensory language-action model that also forecasts; Cross-Modal Instructions (WeI2I.294) and TeNet (WeI1LB.19, text-to-network compact policy synthesis) explore alternative conditioning. CLAR (TuI1I.150) learns 3D representations by fusing masked reconstruction with contrastive alignment. ABPolicy (ThI2I.158) uses asynchronous B-spline action representation for smooth real-time control; Viper (TuI2I.139) targets verifiable, efficient IL policies.
9. IL βͺ RL: residual, adversarial, IRL & self-improvement
The boundary with RL is heavily trafficked. Demonstration-augmented RL: Rainbow-DemoRL (ThI1I.299), GRAPE (ThI1I.94, arXiv 2411.19309; trajectory-level preference alignment, +51.8%/+58.2%). Adversarial IL: GAIL for robot swarms (ThI1I.214), EDAIL (WeI2I.115). Inverse RL / reward learning: Masked IRL (ThI2I.233, LLM-guided reward disambiguation), ILCL (ThI2I.353, inverse logic-constraint learning), neural CBFs from demos via inverse constraint learning (ThI2I.123), MEDIRL traversability cost maps (TuI1LB.3). Self-improvement & residual: SOE (TuI1I.134, on-manifold exploration), Robust Online Residual Refinement via Koopman dynamics (WeI1I.134), Unlocking SAC for IL (TuI2I.270), and Quasimetric Decision Transformers (TuI2I.328).
10. Domains: locomotion, driving, navigation, medical, HRI
IL/LfD spreads well beyond tabletop manipulation. Legged/locomotion: KiRAS keyframe self-imitation (ThI2I.142), quadruped walking from seconds of demos (TuI2I.118), Behavior Foundation Model for Humanoids (TuI1I.120), adaptive motion priors (WeI1I.235), RPG humanoid fighting (TuI2I.280), bicycle flip stunts (ThI2I.240). Driving: CAPS priority sampling (ThI1I.202), STAGE personalized style (ThI2I.410), CoPlanner (TuI1I.75), Mimir (TuI2I.388), location-specific behavior priors (TuBT3.5), and the planning study Perfect Prediction or Plenty of Proposals? (TuI2I.72). Navigation: The One RING (ThAT1.4, arXiv 2412.14401; embodiment-agnostic, trained over 1M randomized embodiments), Narrate2Nav (ThI2I.202), ECAHD (ThI2I.280), social navigation from +/- demos (WeI2I.133), sidewalk autopilot (ThI1I.326), synthetic-vs-real navigation data (WeI1I.74), NavGSim (TuI2I.332). Medical: UltraHiT carotid ultrasonography (ThI2I.188), 3D-ultrasound servoing (WeAT2.2), needle insertion by example (ThI2I.160), surgical sim-to-real visuomotor (WeI1I.231). HRI / human factors: Training Humans to Teach Robots (TuI1I.108), passivity-based trajectory/force arbitration (TuI1I.98), soft-rigid feeding system (ThI2I.348).
11. Evaluation, benchmarks & infrastructure
A Taxonomy for Evaluating Generalist Robot Manipulation Policies (ThI1I.435) proposes a framework for quantifying generalization; BOSS (ThI1I.17) benchmarks observation-space shift. Real-Is-Sim (ThI2I.56) bridges sim-to-real with a dynamic digital twin inside behavior-cloning pipelines. Multi-agent IL: R2BC (WeI1I.71, multi-agent from single-agent demos) and behavior models for multi-agent driving simulation (WeI2I.438). Imitation-BT (WeI2I.439) distills RL agents into behavior trees for interpretability.
Standout deep-dives
All arXiv IDs and headline numbers below were confirmed via web search against the papers' own abstracts/pages.
-
The One RING (
ThAT1.4, arXiv 2412.14401, UW / AI2) β an embodiment-agnostic indoor navigation generalist trained entirely in simulation over 1M randomized embodiments (camera params, collider sizes, center of rotation). Achieves 72.1% sim and 78.9% real-world avg object-goal success across multiple unseen platforms (Stretch RE-1, LoCoBot, Go1). A clean demonstration that massive embodiment randomization can replace per-robot retraining. -
SCIZOR (
ThBT2.3, arXiv 2505.22626, UT-Austin RPL) β self-supervised data curation that filters sub-optimal (low task-progress) and redundant (deduplicated state-action) samples. Yields +15.4% avg across OXE, RoboMimic, and real Sirius-Fleet while using less data β the strongest case this cycle that demonstration quality beats quantity. -
MolmoAct (
TuI1I.327, arXiv 2508.07917, AI2) β Action Reasoning Models: depth-aware perception tokens β editable visual trajectory traces β action tokens. 70.5% zero-shot on SimplerEnv visual matching (beats Οβ and GR00T N1.5), 86.6% avg LIBERO, +10%/+22.7% real single-arm/bimanual over Οβ-FAST. Fully open (weights, data, reasoning dataset). -
AMPLIFY (
TuI2I.286, arXiv 2506.14198) β separates visual-motion prediction from action inference using discrete motion tokens from keypoints; trains the forward model on action-free video, the inverse model on few labeled examples. Up to 3.7Γ better dynamics MSE, 1.2β2.2Γ downstream gains in low-data, and the first LIBERO generalization from zero in-distribution action data. -
ViSA-Flow (
ThI2I.98, arXiv 2505.01288, U-Michigan / KTH) β semantic action flow as a visual-difference-invariant intermediate, pretrained on large-scale human-object-interaction video then adapted with few robot demos. SOTA on CALVIN and real tasks, especially in low-data regimes. -
GRAPE (
ThI1I.94, arXiv 2411.19309, UNC et al.) β trajectory-level preference alignment for VLAs that models reward from both successful and failed rollouts, with spatiotemporal constraints from a VLM. +51.8% in-domain / +58.2% unseen success, plus β37.4% collisions when aligned for safety. -
TWIST2 (
WeBT3.1, arXiv 2511.02832, Stanford) β mocap-free, portable whole-body humanoid teleop/collection (PICO4U VR + ~$250 2-DoF robot neck for egocentric vision), 100 demos in 15 min at near-100% success, with a hierarchical visuomotor policy. Fully open-sourced β addresses humanoid robotics' missing data-collection infrastructure. -
ActiveUMI (
WeI2I.276, arXiv 2510.01607) β robot-free in-the-wild bimanual data via a portable VR kit whose sensorized controllers mirror the robot's end-effectors, crucially capturing the operator's active head/gaze motion. 70% avg task success vs 42% (fixed head cam) / 26% (wrist-only), showing active perception is learnable from human demos.
Complete paper list (116)
Papers are listed under their sub-theme. The "Code/links" column marks an open-source release or project page where a search surfaced one (β). The arXiv column lists IDs confirmed via web search; blank means none was located (not that none exists).
Data collection & teleoperation systems
| Code | Title | Paper ID | arXiv |
|---|---|---|---|
| β | ActiveUMI: Robotic Manipulation with Active Perception from Robot-Free Human Demonstrations | WeI2I.276 | 2510.01607 |
| β | TWIST2: Scalable, Portable, and Holistic Humanoid Data Collection System | WeBT3.1 | 2511.02832 |
| β | EgoMI: Learning Active Vision and Whole-Body Manipulation from Egocentric Human Demonstrations | TuI2I.259 | 2511.00153 |
| COBALT: Crowdsourcing Robot Learning Via Cloud-Based Teleoperation with Smartphones | ThI1I.238 | ||
| RAMPA: Robotic Augmented Reality for Machine Programming by DemonstrAtion | TuI2I.3 | 2410.13412 | |
| AVR: Active Vision-Driven Precise Robot Manipulation with Viewpoint and Focal Length Optimization | TuI1I.239 | 2503.01439 | |
| Robot Control Stack: A Lean Ecosystem for Robot Learning at Scale | TuI1I.250 | 2509.14932 | |
| Hierarchical Grid-Based Sensor Pose Extraction for Demonstration Dataset Generation | TuI2LB.10 |
Demonstration scaling, curation & augmentation
| Code | Title | Paper ID | arXiv |
|---|---|---|---|
| β | SCIZOR: A Self-Supervised Approach to Data Curation for Large-Scale Imitation Learning | ThBT2.3 | 2505.22626 |
| Quality Over Quantity: Demonstration Curation Via Influence Functions for Data-Centric Robot Learning | ThI2I.153 | 2603.09056 | |
| Beyond the Majority: Long-Tail Imitation Learning for Robotic Manipulation | WeI1I.286 | 2602.06512 | |
| LLM Trainer: Automated Robotic Data Generating Via Demonstration Augmentation Using LLMs | TuI1I.336 | 2509.20070 | |
| SoftMimicGen: A Data Generation System for Scalable Robot Learning in Deformable Object Manipulation | WeI2I.305 | 2603.25725 | |
| MOVE: A Simple Motion-Based Data Collection Paradigm for Spatial Generalization in Robotic Manipulation | ThI1I.192 | 2512.04813 | |
| The Curse of Precision: A Data Scaling Law for High-Precision Robotic Manipulation | ThI2I.162 | ||
| Data Scaling Laws for Imitation Learning-Based End-To-End Autonomous Driving | ThI2I.259 | 2412.02689 |
Learning from human video & play (actionless data)
| Code | Title | Paper ID | arXiv |
|---|---|---|---|
| β | AMPLIFY: Actionless Motion Priors for Robot Learning from Videos | TuI2I.286 | 2506.14198 |
| β | ViSA-Flow: Accelerating Robot Skill Learning Via Large-Scale Video Semantic Action Flow | ThI2I.98 | 2505.01288 |
| β | MimicDroid: In-Context Learning for Humanoid Robot Manipulation from Human Play Videos | WeBT3.3 | 2509.09769 |
| β | MotionTrans: Human VR Data Enable Motion-Level Learning for Robotic Manipulation Policies | WeI1I.96 | 2509.17759 |
| EasyMimic: A Low-Cost Framework for Robot Imitation Learning from Human Videos | WeI2I.193 | 2602.11464 | |
| HAND Me the Data: Fast Robot Adaptation Via Hand Path Retrieval | WeI2I.182 | 2505.20455 |
Classical structured LfD: dynamical systems, movement primitives, manifolds
| Code | Title | Paper ID | arXiv |
|---|---|---|---|
| Learning Dynamical System-Based Robot Motions from Demonstrations Via ODE-Driven Diffeomorphic Mappings | TuI1I.322 | ||
| Learning Constraint-Aware Dynamical Systems from Human Demonstrations for Constrained Manipulation Tasks | TuI1I.333 | ||
| Safe and Stable Neural Network Dynamical Systems for Robot Motion Planning | TuI1I.410 | 2511.20593 | |
| Behavior-Controllable Stable Dynamics Models on Riemannian Configuration Manifolds | ThI2I.418 | ||
| Dynamic Movement Primitives with Control Barrier Functions for Constrained Trajectory Planning | WeI1I.30 | ||
| SafeDMPs: Integrating Formal Safety with DMPs for Adaptive HRI | ThI1I.273 | 2603.29708 | |
| Learning Contact Tasks Skills Based on DMP and Affordance Templates | TuI1LB.2 | ||
| Differentiable Motion Manifold Primitives for Reactive Motion Generation under Kinodynamic Constraints | ThI1I.31 | 2410.12193 | |
| From Movement Primitives to Distance Fields to Dynamical Systems | WeI2I.27 | 2504.09705 | |
| Learning from Demonstrations Over Riemannian Manifolds Using Neural ODEs | TuI2LB.24 | ||
| Riemannian Time Warping: Multiple Sequence Alignment in Curved Spaces | WeI1I.358 | 2506.01635 | |
| Task-Parameterized Motion Learning with Time-Sensitive Constraints | TuI1I.363 | 2312.03506 | |
| An Alignment-Based Approach to Learning Motions from Demonstrations | TuI2I.358 | 2511.14988 | |
| A Computationally Efficient Nonparametric Approach for Robot Imitation Learning | ThI2I.170 |
Generalization, robustness & data quality at test time
| Code | Title | Paper ID | arXiv |
|---|---|---|---|
| β | PEEK: Guiding and Minimal Image Representations for Zero-Shot Generalization of Robot Manipulation Policies | WeI1I.279 | 2509.18282 |
| AutoFocus-IL: VLM-Based Saliency Maps for Data-Efficient Visual Imitation Learning without Extra Human Annotations | ThI1I.323 | 2511.18617 | |
| The Temporal Trap: Entanglement in Pre-Trained Visual Representations for Visuomotor Policy Learning | WeI1I.61 | 2502.03270 | |
| Unsupervised Domain Adaptation for Robust Imitation Learning under Visual Perturbations | WeI1I.218 | ||
| Task Robustness Via Re-Labelling Vision-Action Robot Data | WeI1I.199 | ||
| Using Non-Expert Data to Robustify Imitation Learning Via Offline Reinforcement Learning | ThI2I.234 | 2510.19495 | |
| Beyond the Teacher: Leveraging Mixed-Skill Demonstrations for Robust Imitation Learning | WeI2I.291 | ||
| Better Than Diverse Demonstrators: Reward Decomposition from Suboptimal and Heterogeneous Demonstrations | WeBT2.7 | ||
| Failure Identification in Imitation Learning Via Statistical and Semantic Filtering | WeI2I.172 | 2604.13788 | |
| BOSS: Benchmark for Observation Space Shift in Long-Horizon Task | ThI1I.17 | 2502.15679 |
Cross-embodiment & transfer learning
| Code | Title | Paper ID | arXiv |
|---|---|---|---|
| Cross-Embodiment Transfer Via Behavior-Aligned Representations | ThI2I.283 | ||
| Cross-Embodiment Imitation: Learning a Unified Latent Space for Multi-Robot Control | ThI2I.404 | 2601.15419 | |
| One-Shot Cross-Geometry Skill Transfer through Part Decomposition | ThI1I.122 | 2604.15455 | |
| Efficient Learning of Object Placement with Intra-Category Transfer | TuI2I.367 | 2411.03408 | |
| Spatio-Temporal Motion Retargeting for Quadruped Robots | WeI2I.21 | 2404.11557 | |
| D-CLING: Prior-Preserving Depth-Conditioned Fine-Tuning for Navigation Foundation Models | TuI2I.145 |
Long-horizon, hierarchical, lifelong & skill-discovery IL
| Code | Title | Paper ID | arXiv |
|---|---|---|---|
| Playbook: Scalable Discrete Skill Discovery from Unstructured Datasets for Long-Horizon Decision-Making Problems | TuI2I.38 | ||
| T2S: Tokenized Skill Scaling for Lifelong Imitation Learning | TuI1I.72 | 2508.01167 | |
| SPREAD: Subspace Representation Distillation for Lifelong Imitation Learning | TuI2I.102 | 2603.08763 | |
| From Dream to Action: Hierarchical Policy Learning with 3D World Imagination for Robotic Manipulation | TuI1I.146 | ||
| Sequentially Teaching Sequential Tasks (ST)Β²: Teaching Robots Long-Horizon Manipulation Skills | WeI1I.424 | 2510.21046 | |
| MTIL: Encoding Full History with Mamba for Temporal Imitation Learning | TuI2I.348 | 2505.12410 | |
| History-Aware Visuomotor Policy Learning Via Point Tracking | TuI2I.168 | 2509.17141 |
Action representation, tokenization & reasoning
| Code | Title | Paper ID | arXiv |
|---|---|---|---|
| β | MolmoAct: Action Reasoning Models That Can Reason in Space | TuI1I.327 | 2508.07917 |
| β | Fast ECoT: Efficient Embodied Chain-Of-Thought Via Thoughts Reuse | ThI2I.60 | 2506.07639 |
| MLA: A Multisensory LanguageβAction Model for Multimodal Understanding and Forecasting in Robotic Manipulation | TuI2I.195 | 2509.26642 | |
| Cross-Modal Instructions for Robot Motion Generation | WeI2I.294 | 2509.21107 | |
| TeNet: Text-To-Network for Compact Policy Synthesis | WeI1LB.19 | 2601.15912 | |
| CLAR: Learning 3D Representations for Robotic Manipulation by Fusing Masked Reconstruction with Multi-Level Contrastive Alignment | TuI1I.150 | ||
| ABPolicy: Asynchronous B-Spline Flow Policy for Real-Time and Smooth Robotic Manipulation | ThI2I.158 | 2602.23901 | |
| Viper: Verifiable Imitation Learning Policy for Efficient Robotic Manipulation | TuI2I.139 |
IL βͺ RL: residual, adversarial, IRL & self-improvement
| Code | Title | Paper ID | arXiv |
|---|---|---|---|
| β | GRAPE: Generalizing Robot Policy Via Preference Alignment | ThI1I.94 | 2411.19309 |
| Rainbow-DemoRL: Combining Improvements in Demonstration-Augmented Reinforcement Learning | ThI1I.299 | 2603.27400 | |
| Generative Adversarial Imitation Learning for Robot Swarms: Learning from Human Demonstrations and Trained Policies | ThI1I.214 | 2603.02783 | |
| EDAIL: Adversarial Imitation Learning Via Exploration-Driven Data Augmentation | WeI2I.115 | ||
| Masked IRL: LLM-Guided Reward Disambiguation from Demonstrations and Language | ThI2I.233 | 2511.14565 | |
| ILCL: Inverse Logic-Constraint Learning from Temporally Constrained Demonstrations | ThI2I.353 | 2507.11000 | |
| Learning Neural Control Barrier Functions from Expert Demonstrations Using Inverse Constraint Learning | ThI2I.123 | 2510.21560 | |
| Learning Traversability Cost Maps with Decomposed Uncertainties Via Continuous-State MEDIRL | TuI1LB.3 | ||
| SOE: Sample-Efficient Robot Policy Self-Improvement Via On-Manifold Exploration | TuI1I.134 | 2509.19292 | |
| Robust Online Residual Refinement Via Koopman-Guided Dynamics Modeling | WeI1I.134 | 2509.12562 | |
| Unlocking the Potential of Soft Actor-Critic for Imitation Learning | TuI2I.270 | 2509.24539 | |
| Quasimetric Decision Transformers: Enhancing Goal-Conditioned Reinforcement Learning with Structured Distance Guidance | TuI2I.328 | ||
| Data-Efficient Hierarchical Goal-Conditioned Reinforcement Learning Via Normalizing Flows | ThI1I.203 | 2602.11142 |
Domain: locomotion & humanoid motion
| Code | Title | Paper ID | arXiv |
|---|---|---|---|
| KiRAS: Keyframe Guided Self-Imitation for Robust and Adaptive Skill Learning in Quadruped Robots | ThI2I.142 | 2603.15179 | |
| Learning Quadruped Walking from Seconds of Demonstration | TuI2I.118 | 2603.06961 | |
| Behavior Foundation Model for Humanoid Robots | TuI1I.120 | 2509.13780 | |
| Adaptive Motion Priors with Constrained Optimization | WeI1I.235 | ||
| RPG: Robust Policy Gating for Smooth Multi-Skill Transitions in Humanoid Fighting | TuI2I.280 | 2604.21355 | |
| Flip Stunts on Bicycle Robots Using Iterative Motion Imitation | ThI2I.240 | 2603.27944 | |
| Teaching to Individual Needs: Bidirectional Teacher-Student Learning for Wheeled-Legged Locomotion | ThI1I.201 |
Domain: autonomous driving
| Code | Title | Paper ID | arXiv |
|---|---|---|---|
| CAPS: Context-Aware Priority Sampling for Enhanced Imitation Learning in Autonomous Driving | ThI1I.202 | 2503.01650 | |
| STAGE: STyle-Controllable Action GEneration for Personalized Autonomous Driving | ThI2I.410 | ||
| CoPlanner: An Interactive Motion Planner with Contingency-Aware Diffusion for Autonomous Driving | TuI1I.75 | 2509.17080 | |
| Mimir: Hierarchical Goal-Driven Diffusion with Uncertainty Propagation for End-To-End Autonomous Driving | TuI2I.388 | 2512.07130 | |
| Learning Location-Specific Latent Behavior Priors for Occupancy Prediction in Automated Driving | TuBT3.5 | ||
| Perfect Prediction or Plenty of Proposals? What Matters Most in Planning for Autonomous Driving | TuI2I.72 | 2510.15505 | |
| Toward Efficient and Robust Behavior Models for Multi-Agent Driving Simulation | WeI2I.438 | 2512.05812 |
Domain: navigation
| Code | Title | Paper ID | arXiv |
|---|---|---|---|
| β | The One RING: A Robotic Indoor Navigation Generalist | ThAT1.4 | 2412.14401 |
| Narrate2Nav: Real-Time Visual Navigation with Implicit Language Reasoning in Human-Centric Environments | ThI2I.202 | 2506.14233 | |
| ECAHD: Efficient Collision-Aware Hierarchical Diffusion Navigation | ThI2I.280 | ||
| Learning Social Navigation from Positive and Negative Demonstrations and Rule-Based Specifications | WeI2I.133 | 2510.12215 | |
| Learning Sidewalk Autopilot from Multi-Scale Imitation with Corrective Behavior Expansion | ThI1I.326 | 2603.22527 | |
| Synthetic vs. Real Training Data for Visual Navigation | WeI1I.74 | 2509.11791 | |
| NavGSim: High-Fidelity Gaussian Splatting Simulator for Large-Scale Navigation | TuI2I.332 | 2603.15186 |
Domain: medical & surgical
| Code | Title | Paper ID | arXiv |
|---|---|---|---|
| UltraHiT: A Hierarchical Transformer Architecture for Generalizable Internal Carotid Artery Robotic Ultrasonography | ThI2I.188 | 2509.13832 | |
| Breaking the Latency Barrier: Synergistic Perception and Control for High-Frequency 3D Ultrasound Servoing | WeAT2.2 | 2511.00983 | |
| Real-Time Robotic Needle Insertion in Deformable and Moving Structure Using Learning-By-Example Method | ThI2I.160 | ||
| Overcoming Imperfect Kinematics in Surgical Robotics through Sim-To-Real Visuomotor Learning | WeI1I.231 |
Domain: HRI, human factors, contact & force
| Code | Title | Paper ID | arXiv |
|---|---|---|---|
| Training Humans to Teach Robots: Large and Lasting Skill Gains | TuI1I.108 | ||
| A Passivity-Based Framework for Dynamic Arbitration between Trajectory and Force Tracking Using Human Demonstration | TuI1I.98 | ||
| A Soft-Rigid Hybrid Robot-Assisted Feeding System with a Tendon-Driven Continuum Robot | ThI2I.348 | ||
| Spline-FRIDA: Towards Diverse, Humanlike Robot Painting Styles with a Sample-Efficient, Differentiable Brush Stroke Model | WeI1I.380 | 2412.00597 |
Evaluation, benchmarks, multi-agent & infrastructure
| Code | Title | Paper ID | arXiv |
|---|---|---|---|
| A Taxonomy for Evaluating Generalist Robot Manipulation Policies | ThI1I.435 | 2503.01238 | |
| Real-Is-Sim: Bridging the Sim-To-Real Gap with a Dynamic Digital Twin | ThI2I.56 | 2504.03597 | |
| R2BC: Multi-Agent Imitation Learning from Single-Agent Demonstrations | WeI1I.71 | 2510.18085 | |
| Imitation-BT: Automating Behavior Tree Generation by Echoing Reinforcement Learning Agents | WeI2I.439 |
Other / cross-cutting IL methods
| Code | Title | Paper ID | arXiv |
|---|---|---|---|
| Flow-Enabled Generalization to Human Demonstrations in Few-Shot Imitation Learning | TuI2I.53 | 2602.10594 | |
| Real-Time Generation of Near-Minimum-Energy Trajectories Via Constraint-Informed Residual Learning | WeI2I.383 | 2501.09450 | |
| Learning Multiple Initial Solutions to Optimization Problems | WeI1I.45 | 2411.02158 |
Related
β Back to ICRA-2026-VLA-Manipulation-Survey