ICRA 2026 Topic Planning - Heungwoo/research GitHub Wiki
ICRA 2026 — Manipulation Planning & TAMP (Topic Analysis)
Venue: IEEE ICRA 2026 · Vienna, Austria · June 1–5, 2026 Compiled against the official ICRA 2026 PaperCept program. Paper IDs (e.g.
ThI1I.72) are program session/slot codes. arXiv IDs were confirmed individually where listed; entries without an arXiv link were not located and are left blank rather than guessed.
This page covers the 50 papers grouped under Manipulation Planning & Task-and-Motion Planning in the ICRA 2026 program. The cluster spans the full planning stack: classical sampling/optimization-based motion planning (RRT-Connect variants, tensor/GPU-batched planners, screw-theoretic IK), learning-augmented planning (diffusion/consistency/flow planners, RL-tuned classical planners), LLM/VLM-driven task planning (failure detection, contact-acceptability reasoning, assembly-manual parsing), and long-horizon symbolic + skill planning (TAMP solvers, skill libraries, symbol/skill co-invention). Where the VLA survey tracks the policy frontier, this cluster tracks the deliberative frontier — how a robot decides what to do and in what order, and how to certify the resulting motions are feasible, optimal, or safe.
A recurring theme: foundation models are no longer only the policy. They increasingly sit above a classical planner — as a sequence proposer, a constraint/cost writer, a failure detector, or a demonstration generator — while the geometric/optimization machinery underneath retains the feasibility and optimality guarantees that learned policies still lack.
Sub-trends
1. LLM/VLM-driven task planning
The largest semantic shift is using VLMs as the deliberative layer rather than the action layer. Robust Task Planning via Failure Detection Using Scene Graph from Multi-View Images (ThAT1.6) argues that LLM/VLM failure detectors over-assume full scene understanding and grounds detection in an explicit multi-view scene graph. IMPACT (WeI1I.64, arXiv 2503.10110) uses a VLM to infer which surfaces tolerate contact, emitting an anisotropic cost map for a contact-aware A* — relaxing the collision-free assumption that makes classical planning brittle in clutter. Manual2Skill++ (WeI2I.159, arXiv 2510.16344) parses assembly instruction manuals into connector-aware hierarchical graphs, elevating connectors (screws, pegs) to first-class planning primitives. AdaptPNP (WeI2I.184, arXiv 2511.11052) has a VLM emit a prehensile/non-prehensile plan skeleton refined against a digital-twin object-pose predictor. Seeing Farther and Smarter (ThI2I.132) adds value-guided multi-path reflection to VLM policy optimization for long-horizon reasoning, and TARAD (TuI1I.415) uses LLM-generated demonstrations to bootstrap an affordance-centric diffusion policy.
2. Integrated task-and-motion planning (TAMP)
Classic TAMP work targets the combinatorial blowup of long horizons. Learning Problem Decomposition for Efficient Sequential Multi-Object Manipulation Planning (ThBT2.7) attacks the exponential growth of TAMP solve time with object count via learned decomposition for fast replanning in dynamic scenes. SymSkill (WeBT1.8, arXiv 2510.01661) co-invents predicates, operators and skills from unsegmented play data — bridging IL's reactivity with TAMP's compositional generalization (85% single-step in RoboCasa; 11 operators from 5 min of Franka play). From CAD to POMDP (TuI1I.114) casts robotic disassembly sequence planning as a POMDP to handle uncertain, partially observable end-of-life products. MOASIC (WeI1I.256) and Uni-Skill (ThI1I.86) attack long-horizon planning over predefined / self-evolving skill libraries with physics simulation in the loop.
3. Learning-augmented motion planning for manipulation
Diffusion/flow/consistency models are increasingly used as fast, multi-modal trajectory generators inside otherwise classical pipelines. Accelerated Multi-Modal Motion Planning Using Context-Conditioned Diffusion Models (ThI2I.183, arXiv 2510.14615; "CAMPD") conditions a classifier-free diffusion U-Net on arbitrary context for 7-DoF planning at a fraction of baseline time. CAPE (ThI2I.63) expands diffusion-policy modes for collision avoidance; ConsistencyPlanner (TuI1I.298) uses fast-sampling consistency models for real-time closed-loop planning. Enhancing Classical Motion Planners Using RL with Safety Guarantees (TuI2I.143) keeps a classical planner's safety while RL-tuning its parameters online. DynDLO (TuI2I.201) learns trajectory planning for dynamic deformable-linear-object manipulation, and KAN Policy (TuI1I.47) uses Kolmogorov–Arnold networks for smooth trajectories.
4. Classical & batched motion planning with guarantees
A strong "fast + provably good" thread persists. AORRTC (WeI1I.339, arXiv 2505.10542) applies the AO-x meta-algorithm to RRT-Connect, getting RRT-Connect-speed initial solutions and almost-sure asymptotic optimality — solving hard high-DoF problems in milliseconds (Panda 7-DoF, Fetch 8-DoF on MotionBenchMaker). Global Tensor Motion Planning (WeI1I.68, arXiv 2411.19393) reformulates sampling-based planning as pure tensor ops over a random multipartite graph for GPU/TPU batch planning. GeoFIK (ThI2I.300, arXiv 2503.03992) is an analytical screw-theory IK solver for the 7-DoF Franka that enumerates redundancy solutions with free Jacobian. Safety-Critical Dynamic Motion Generation (TuI1I.244) uses differentiable configuration-space distance fields with CBFs; Optimal Dexterity Path Planning (WeI1I.187) maximizes workspace-density dexterity inside a sampling planner.
5. Constraint / contact / affordance-based planning (ReKep-style)
Echoing ReKep's "VLM-as-constraint-writer" paradigm, several papers plan over geometric/contact constraints rather than dense trajectories. A Closed-Chain Approach to Generating Affordance Joint Trajectories (ThI1I.358) extends screw-based affordance planning while avoiding singular/undesirable configurations. Screw Geometry Meets Bandits (WeI2I.259) incrementally acquires kinesthetic demonstrations (bandit-driven) to build screw-geometry manipulation plans. IMPACT (above) and the affordance/connector framing of Manual2Skill++ also fit this constraint-first lineage, as does A Contact-Driven Framework for Manipulating in the Blind (ThI2I.295), which plans from contact feedback when vision is inadequate.
6. Non-prehensile & sequential manipulation
Beyond pick-and-place, planners increasingly reason about pushing/sliding and contact-mode switches. H-MaP (TuI2I.15, arXiv 2403.10436) is a hybrid sequential planner decoupling object-trajectory from manipulation planning, handling tool use and contact-mode switches. AdaptPNP (above) unifies prehensile + non-prehensile skill selection. Robustness-Aware Tool Selection and Manipulation Planning (TuI2I.88) jointly picks tools and plans contact-rich motions under learned energy-informed robustness guidance. Pack It In (TuI1I.211) plans packing into partially filled containers through contact, and Peg-in-Hole (TuI2I.12) uses passive compliance for error-tolerant insertion.
7. Multi-object rearrangement
Rearrangement is its own hard combinatorial planning problem. MO-SeGMan (TuI2I.105, arXiv 2511.01476) is a multi-objective sequential/guided rearrangement planner with a Selective Guided Forward Search for non-monotone, cluttered scenes (feasible on all 9 benchmark tasks). Tidiness Score-Guided MCTS (ThI1I.29) plans tabletop tidying from RGB-D via a learned tidiness score guiding Monte Carlo tree search. Placeit! (WeI1I.124) learns object-placement skills with auto-generated training data.
8. Planning under uncertainty & long-horizon subgoals
Belief-space and subgoal planning close the loop with partial observability. Planning Using Belief Summaries (TuI1I.126) does goal-directed articulated-object manipulation from force/proprioception under belief uncertainty. From CAD to POMDP (above) is the disassembly instance. Not Throwing Away My Shot (ThI1I.72) plans long-horizon manipulation with dual subgoals (short-horizon + low-variance) to pick informative subgoals; Learning Composable Skills ("STACK", TuI2I.142) discovers spatial/temporal structure from foundation models for skill composition.
Standout deep-dives
AORRTC — WeI1I.339 · arXiv 2505.10542
The cleanest "classical planning still wins on guarantees" result in the cluster. By wrapping RRT-Connect in the AO-x meta-algorithm, AORRTC matches RRT-Connect's initial-solution speed yet converges almost-surely to the optimum in an anytime fashion. On MotionBenchMaker with the Panda (7-DoF) and Fetch (8-DoF), it finds solutions to hard high-DoF instances in milliseconds where prior a.s.a.o. planners couldn't reliably solve in seconds. A reminder that the bar learned planners must clear is high.
SymSkill — WeBT1.8 · arXiv 2510.01661
A genuine TAMP-meets-IL synthesis (UPenn GRASP): jointly co-invents predicates, operators, and skills from unlabeled, unsegmented demonstrations, getting TAMP's compositional generalization with IL's real-time reactivity and recovery. 85% single-step success in RoboCasa, composing to multi-step tasks with no extra data; on a real Franka it learns 11 operators from 5 minutes of play data and hits user-specified symbolic goals in real time. Directly addresses the symbol-grounding bottleneck that has limited classical TAMP.
IMPACT — WeI1I.64 · arXiv 2503.10110
A ReKep-style relaxation of the collision-free dogma. A VLM infers per-region contact tolerance from object semantics, producing an anisotropic 3D cost map encoding directional push safety; a contact-aware A* then plans semantically-acceptable contact-rich paths through clutter that pure collision-free planners cannot traverse. Builds conceptually on the VLM-as-cost/constraint-writer paradigm of ReKep. (USC LIRA Lab.)
MO-SeGMan — TuI2I.105 · arXiv 2511.01476
State-of-the-art constrained multi-object rearrangement (TU Munich / Toussaint & Oguz). A Selective Guided Forward Search relocates only critical obstacles, plus adaptive subgoal refinement removes redundant pick-and-place; lazy evaluation jointly minimizes per-object replanning and robot travel. Generates feasible plans on all 9 benchmark rearrangement tasks with faster solve times and better quality than baselines — important for non-monotone, highly cluttered scenes.
Manual2Skill++ — WeI2I.159 · arXiv 2510.16344
Treats connectors as first-class primitives: a VLM extracts structured hierarchical connection graphs (connector type, spec, quantity, placement) from assembly manuals, enabling millimeter-level pose alignment for robust execution. Ships a connector-annotated dataset and a multi-connector-modality simulation benchmark. A concrete instance of LLM/VLM-driven task planning grounded in real document structure rather than free-form prompting.
Accelerated Multi-Modal Motion Planning (CAMPD) — ThI2I.183 · arXiv 2510.14615
Representative of the learned-planner-as-generator thread. A classifier-free denoising diffusion U-Net with an attention mechanism conditions on an arbitrary number of sensor-agnostic context parameters, generalizing to unseen environments and producing high-quality multi-modal 7-DoF trajectories at a fraction of the time of state-of-the-art baselines — useful both for deployment and for generating diverse trajectory datasets.
Complete paper list (50)
| ID | Title | arXiv |
|---|---|---|
| ThAT1.6 | Robust Task Planning via Failure Detection Using Scene Graph from Multi-View Images | |
| ThBT2.7 | Learning Problem Decomposition for Efficient Sequential Multi-Object Manipulation Planning | 2408.06843 |
| ThI1I.242 | Find the Fruit: Zero-Shot Sim2Real RL for Occlusion-Aware Plant Manipulation | 2505.16547 |
| ThI1I.29 | Tidiness Score-Guided Monte Carlo Tree Search for Visual Tabletop Rearrangement | 2502.17235 |
| ThI1I.358 | A Closed-Chain Approach to Generating Affordance Joint Trajectories for Robotic Manipulators | |
| ThI1I.397 | Whole-Body Integrated Motion Planning for Aerial Manipulators | 2501.06493 |
| ThI1I.72 | Not Throwing Away My Shot: Planning Ahead with Dual Subgoals in Long-Horizon Robot Manipulation Tasks | |
| ThI1I.86 | Uni-Skill: Building Self-Evolving Skill Repository for Generalizable Robotic Manipulation | 2603.02623 |
| ThI2I.132 | Seeing Farther and Smarter: Value-Guided Multi-Path Reflection for VLM Policy Optimization | 2602.19372 |
| ThI2I.175 | The iMETRO Dynamic Simulation: An Open-Source Simulator for Intravehicular Space Robotics Research | |
| ThI2I.183 | Accelerated Multi-Modal Motion Planning Using Context-Conditioned Diffusion Models (CAMPD) | 2510.14615 |
| ThI2I.295 | A Contact-Driven Framework for Manipulating in the Blind | 2510.20177 |
| ThI2I.300 | GeoFIK: A Fast and Reliable Geometric Solver for the IK of the Franka Arm Based on Screw Theory | 2503.03992 |
| ThI2I.54 | Distracted Robot: How Visual Clutter Undermine Robotic Manipulation | 2511.22780 |
| ThI2I.63 | CAPE: Context-Aware Diffusion Policy via Proximal Mode Expansion for Collision Avoidance | 2511.22773 |
| TuAT3.4 | DYMO-Hair: Generalizable Volumetric Dynamics Modeling for Robot Hair Manipulation | 2510.06199 |
| TuI1I.114 | From CAD to POMDP: Probabilistic Planning for Robotic Disassembly of End-Of-Life Products | 2511.23407 |
| TuI1I.126 | Planning Using Belief Summaries for Goal-Directed Manipulation of Articulated Objects with Force and Proprioception | |
| TuI1I.211 | Pack It In: Packing into Partially Filled Containers through Contact | 2602.12095 |
| TuI1I.244 | Safety-Critical Dynamic Motion Generation for Manipulators Using Differentiable Distance Fields in Configuration Space | 2412.16456 |
| TuI1I.298 | ConsistencyPlanner: Real-Time Planning with Fast-Sampling Consistency Models | |
| TuI1I.415 | TARAD: Task-Aware Robot Affordance-Centric Diffusion Policy Learned from LLM-Generated Demonstrations | |
| TuI1I.47 | KAN Policy: Learning Efficient and Smooth Robotic Trajectories via Kolmogorov-Arnold Networks | |
| TuI2I.105 | MO-SeGMan: Rearrangement Planning Framework for Multi-Objective Sequential and Guided Manipulation in Constrained Environments | 2511.01476 |
| TuI2I.12 | Robust and Error-Tolerant Peg-In-Hole Assembly Using Simple Control | |
| TuI2I.142 | Learning Composable Skills by Discovering Spatial and Temporal Structure with Foundation Models (STACK) | |
| TuI2I.143 | Enhancing Classical Motion Planners Using RL with Safety Guarantees | 2403.18524 |
| TuI2I.15 | H-MaP: An Iterative and Hybrid Sequential Manipulation Planner | 2403.10436 |
| TuI2I.201 | DynDLO: Learning-Based Trajectory Planning for Dynamic Robotic Manipulation of Deformable Linear Objects | |
| TuI2I.272 | Run-Time Optimization of Overall Energy Consumption in Lightweight Collaborative Arms for Repetitive Tasks | |
| TuI2I.303 | Learning to Drive by Imitating Surrounding Vehicles | 2503.05997 |
| TuI2I.407 | A Differential Dynamic Programming Framework for Inverse Reinforcement Learning | 2407.19902 |
| TuI2I.88 | Robustness-Aware Tool Selection and Manipulation Planning with Learned Energy-Informed Guidance | 2506.03362 |
| WeBT1.8 | SymSkill: Symbol and Skill Co-Invention for Data-Efficient and Reactive Long-Horizon Manipulation | 2510.01661 |
| WeBT2.2 | Human2Nav: Learning Crowd Navigation from Human Videos across Robots via Feasibility-Guided Flow Matching | |
| WeBT2.4 | Shifted Flow Policy: Uncertainty-Aware Time Reparameterization for Visuomotor Learning | |
| WeBT2.5 | Closed-Loop Action Chunks with Dynamic Corrections for Training-Free Diffusion Policy (DCDP) | 2603.01953 |
| WeI1I.124 | Placeit! A Framework for Learning Robot Object Placement Skills | 2510.09267 |
| WeI1I.179 | Task Generalization with Pathwise Conditioning of Gaussian Process for Learning from Demonstration | |
| WeI1I.187 | Optimal Dexterity Path Planning for Robotic Manipulators Using Rapid Workspace Density Approximation | |
| WeI1I.221 | 3DFacePolicy: Speech-Driven 3D Facial Animation Based on Diffusion Policy | 2409.10848 |
| WeI1I.256 | MOASIC: Skill-Centric Manipulation Planning with Physics Simulation | 2504.16738 |
| WeI1I.339 | AORRTC: Almost-Surely Asymptotically Optimal Planning with RRT-Connect | 2505.10542 |
| WeI1I.64 | IMPACT: Intelligent Motion Planning with Acceptable Contact Trajectories via Vision-Language Models | 2503.10110 |
| WeI1I.68 | Global Tensor Motion Planning | 2411.19393 |
| WeI2I.116 | MetaDP: Meta-Manipulation Diffusion Policy for Robotic Manipulation | |
| WeI2I.159 | Manual2Skill++: Connector-Aware General Robotic Assembly from Instruction Manuals via Vision–Language Models | 2510.16344 |
| WeI2I.184 | AdaptPNP: Integrating Prehensile and Non-Prehensile Skills for Adaptive Robotic Manipulation | 2511.11052 |
| WeI2I.259 | Screw Geometry Meets Bandits: Incremental Acquisition of Demonstrations to Generate Manipulation Plans | 2410.18275 |
| WeI2I.300 | GPU-Accelerated Continuous-Time Successive Convexification for Contact-Implicit Legged Locomotion | 2604.09993 |
Related
- ICRA 2026 Survey
- ReKep — VLM-as-constraint-writer lineage referenced by IMPACT / affordance-constraint planners.
← Back to ICRA-2026-VLA-Manipulation-Survey