ICRA 2026 Topic Planning - Heungwoo/research GitHub Wiki

ICRA 2026 — Manipulation Planning & TAMP (Topic Analysis)

Venue: IEEE ICRA 2026 · Vienna, Austria · June 1–5, 2026 Compiled against the official ICRA 2026 PaperCept program. Paper IDs (e.g. ThI1I.72) are program session/slot codes. arXiv IDs were confirmed individually where listed; entries without an arXiv link were not located and are left blank rather than guessed.

This page covers the 50 papers grouped under Manipulation Planning & Task-and-Motion Planning in the ICRA 2026 program. The cluster spans the full planning stack: classical sampling/optimization-based motion planning (RRT-Connect variants, tensor/GPU-batched planners, screw-theoretic IK), learning-augmented planning (diffusion/consistency/flow planners, RL-tuned classical planners), LLM/VLM-driven task planning (failure detection, contact-acceptability reasoning, assembly-manual parsing), and long-horizon symbolic + skill planning (TAMP solvers, skill libraries, symbol/skill co-invention). Where the VLA survey tracks the policy frontier, this cluster tracks the deliberative frontier — how a robot decides what to do and in what order, and how to certify the resulting motions are feasible, optimal, or safe.

A recurring theme: foundation models are no longer only the policy. They increasingly sit above a classical planner — as a sequence proposer, a constraint/cost writer, a failure detector, or a demonstration generator — while the geometric/optimization machinery underneath retains the feasibility and optimality guarantees that learned policies still lack.

Sub-trends

1. LLM/VLM-driven task planning

The largest semantic shift is using VLMs as the deliberative layer rather than the action layer. Robust Task Planning via Failure Detection Using Scene Graph from Multi-View Images (ThAT1.6) argues that LLM/VLM failure detectors over-assume full scene understanding and grounds detection in an explicit multi-view scene graph. IMPACT (WeI1I.64, arXiv 2503.10110) uses a VLM to infer which surfaces tolerate contact, emitting an anisotropic cost map for a contact-aware A* — relaxing the collision-free assumption that makes classical planning brittle in clutter. Manual2Skill++ (WeI2I.159, arXiv 2510.16344) parses assembly instruction manuals into connector-aware hierarchical graphs, elevating connectors (screws, pegs) to first-class planning primitives. AdaptPNP (WeI2I.184, arXiv 2511.11052) has a VLM emit a prehensile/non-prehensile plan skeleton refined against a digital-twin object-pose predictor. Seeing Farther and Smarter (ThI2I.132) adds value-guided multi-path reflection to VLM policy optimization for long-horizon reasoning, and TARAD (TuI1I.415) uses LLM-generated demonstrations to bootstrap an affordance-centric diffusion policy.

2. Integrated task-and-motion planning (TAMP)

Classic TAMP work targets the combinatorial blowup of long horizons. Learning Problem Decomposition for Efficient Sequential Multi-Object Manipulation Planning (ThBT2.7) attacks the exponential growth of TAMP solve time with object count via learned decomposition for fast replanning in dynamic scenes. SymSkill (WeBT1.8, arXiv 2510.01661) co-invents predicates, operators and skills from unsegmented play data — bridging IL's reactivity with TAMP's compositional generalization (85% single-step in RoboCasa; 11 operators from 5 min of Franka play). From CAD to POMDP (TuI1I.114) casts robotic disassembly sequence planning as a POMDP to handle uncertain, partially observable end-of-life products. MOASIC (WeI1I.256) and Uni-Skill (ThI1I.86) attack long-horizon planning over predefined / self-evolving skill libraries with physics simulation in the loop.

3. Learning-augmented motion planning for manipulation

Diffusion/flow/consistency models are increasingly used as fast, multi-modal trajectory generators inside otherwise classical pipelines. Accelerated Multi-Modal Motion Planning Using Context-Conditioned Diffusion Models (ThI2I.183, arXiv 2510.14615; "CAMPD") conditions a classifier-free diffusion U-Net on arbitrary context for 7-DoF planning at a fraction of baseline time. CAPE (ThI2I.63) expands diffusion-policy modes for collision avoidance; ConsistencyPlanner (TuI1I.298) uses fast-sampling consistency models for real-time closed-loop planning. Enhancing Classical Motion Planners Using RL with Safety Guarantees (TuI2I.143) keeps a classical planner's safety while RL-tuning its parameters online. DynDLO (TuI2I.201) learns trajectory planning for dynamic deformable-linear-object manipulation, and KAN Policy (TuI1I.47) uses Kolmogorov–Arnold networks for smooth trajectories.

4. Classical & batched motion planning with guarantees

A strong "fast + provably good" thread persists. AORRTC (WeI1I.339, arXiv 2505.10542) applies the AO-x meta-algorithm to RRT-Connect, getting RRT-Connect-speed initial solutions and almost-sure asymptotic optimality — solving hard high-DoF problems in milliseconds (Panda 7-DoF, Fetch 8-DoF on MotionBenchMaker). Global Tensor Motion Planning (WeI1I.68, arXiv 2411.19393) reformulates sampling-based planning as pure tensor ops over a random multipartite graph for GPU/TPU batch planning. GeoFIK (ThI2I.300, arXiv 2503.03992) is an analytical screw-theory IK solver for the 7-DoF Franka that enumerates redundancy solutions with free Jacobian. Safety-Critical Dynamic Motion Generation (TuI1I.244) uses differentiable configuration-space distance fields with CBFs; Optimal Dexterity Path Planning (WeI1I.187) maximizes workspace-density dexterity inside a sampling planner.

5. Constraint / contact / affordance-based planning (ReKep-style)

Echoing ReKep's "VLM-as-constraint-writer" paradigm, several papers plan over geometric/contact constraints rather than dense trajectories. A Closed-Chain Approach to Generating Affordance Joint Trajectories (ThI1I.358) extends screw-based affordance planning while avoiding singular/undesirable configurations. Screw Geometry Meets Bandits (WeI2I.259) incrementally acquires kinesthetic demonstrations (bandit-driven) to build screw-geometry manipulation plans. IMPACT (above) and the affordance/connector framing of Manual2Skill++ also fit this constraint-first lineage, as does A Contact-Driven Framework for Manipulating in the Blind (ThI2I.295), which plans from contact feedback when vision is inadequate.

6. Non-prehensile & sequential manipulation

Beyond pick-and-place, planners increasingly reason about pushing/sliding and contact-mode switches. H-MaP (TuI2I.15, arXiv 2403.10436) is a hybrid sequential planner decoupling object-trajectory from manipulation planning, handling tool use and contact-mode switches. AdaptPNP (above) unifies prehensile + non-prehensile skill selection. Robustness-Aware Tool Selection and Manipulation Planning (TuI2I.88) jointly picks tools and plans contact-rich motions under learned energy-informed robustness guidance. Pack It In (TuI1I.211) plans packing into partially filled containers through contact, and Peg-in-Hole (TuI2I.12) uses passive compliance for error-tolerant insertion.

7. Multi-object rearrangement

Rearrangement is its own hard combinatorial planning problem. MO-SeGMan (TuI2I.105, arXiv 2511.01476) is a multi-objective sequential/guided rearrangement planner with a Selective Guided Forward Search for non-monotone, cluttered scenes (feasible on all 9 benchmark tasks). Tidiness Score-Guided MCTS (ThI1I.29) plans tabletop tidying from RGB-D via a learned tidiness score guiding Monte Carlo tree search. Placeit! (WeI1I.124) learns object-placement skills with auto-generated training data.

8. Planning under uncertainty & long-horizon subgoals

Belief-space and subgoal planning close the loop with partial observability. Planning Using Belief Summaries (TuI1I.126) does goal-directed articulated-object manipulation from force/proprioception under belief uncertainty. From CAD to POMDP (above) is the disassembly instance. Not Throwing Away My Shot (ThI1I.72) plans long-horizon manipulation with dual subgoals (short-horizon + low-variance) to pick informative subgoals; Learning Composable Skills ("STACK", TuI2I.142) discovers spatial/temporal structure from foundation models for skill composition.

Standout deep-dives

AORRTC — WeI1I.339 · arXiv 2505.10542

The cleanest "classical planning still wins on guarantees" result in the cluster. By wrapping RRT-Connect in the AO-x meta-algorithm, AORRTC matches RRT-Connect's initial-solution speed yet converges almost-surely to the optimum in an anytime fashion. On MotionBenchMaker with the Panda (7-DoF) and Fetch (8-DoF), it finds solutions to hard high-DoF instances in milliseconds where prior a.s.a.o. planners couldn't reliably solve in seconds. A reminder that the bar learned planners must clear is high.

SymSkill — WeBT1.8 · arXiv 2510.01661

A genuine TAMP-meets-IL synthesis (UPenn GRASP): jointly co-invents predicates, operators, and skills from unlabeled, unsegmented demonstrations, getting TAMP's compositional generalization with IL's real-time reactivity and recovery. 85% single-step success in RoboCasa, composing to multi-step tasks with no extra data; on a real Franka it learns 11 operators from 5 minutes of play data and hits user-specified symbolic goals in real time. Directly addresses the symbol-grounding bottleneck that has limited classical TAMP.

IMPACT — WeI1I.64 · arXiv 2503.10110

A ReKep-style relaxation of the collision-free dogma. A VLM infers per-region contact tolerance from object semantics, producing an anisotropic 3D cost map encoding directional push safety; a contact-aware A* then plans semantically-acceptable contact-rich paths through clutter that pure collision-free planners cannot traverse. Builds conceptually on the VLM-as-cost/constraint-writer paradigm of ReKep. (USC LIRA Lab.)

MO-SeGMan — TuI2I.105 · arXiv 2511.01476

State-of-the-art constrained multi-object rearrangement (TU Munich / Toussaint & Oguz). A Selective Guided Forward Search relocates only critical obstacles, plus adaptive subgoal refinement removes redundant pick-and-place; lazy evaluation jointly minimizes per-object replanning and robot travel. Generates feasible plans on all 9 benchmark rearrangement tasks with faster solve times and better quality than baselines — important for non-monotone, highly cluttered scenes.

Manual2Skill++ — WeI2I.159 · arXiv 2510.16344

Treats connectors as first-class primitives: a VLM extracts structured hierarchical connection graphs (connector type, spec, quantity, placement) from assembly manuals, enabling millimeter-level pose alignment for robust execution. Ships a connector-annotated dataset and a multi-connector-modality simulation benchmark. A concrete instance of LLM/VLM-driven task planning grounded in real document structure rather than free-form prompting.

Accelerated Multi-Modal Motion Planning (CAMPD) — ThI2I.183 · arXiv 2510.14615

Representative of the learned-planner-as-generator thread. A classifier-free denoising diffusion U-Net with an attention mechanism conditions on an arbitrary number of sensor-agnostic context parameters, generalizing to unseen environments and producing high-quality multi-modal 7-DoF trajectories at a fraction of the time of state-of-the-art baselines — useful both for deployment and for generating diverse trajectory datasets.

Complete paper list (50)

ID Title arXiv
ThAT1.6 Robust Task Planning via Failure Detection Using Scene Graph from Multi-View Images
ThBT2.7 Learning Problem Decomposition for Efficient Sequential Multi-Object Manipulation Planning 2408.06843
ThI1I.242 Find the Fruit: Zero-Shot Sim2Real RL for Occlusion-Aware Plant Manipulation 2505.16547
ThI1I.29 Tidiness Score-Guided Monte Carlo Tree Search for Visual Tabletop Rearrangement 2502.17235
ThI1I.358 A Closed-Chain Approach to Generating Affordance Joint Trajectories for Robotic Manipulators
ThI1I.397 Whole-Body Integrated Motion Planning for Aerial Manipulators 2501.06493
ThI1I.72 Not Throwing Away My Shot: Planning Ahead with Dual Subgoals in Long-Horizon Robot Manipulation Tasks
ThI1I.86 Uni-Skill: Building Self-Evolving Skill Repository for Generalizable Robotic Manipulation 2603.02623
ThI2I.132 Seeing Farther and Smarter: Value-Guided Multi-Path Reflection for VLM Policy Optimization 2602.19372
ThI2I.175 The iMETRO Dynamic Simulation: An Open-Source Simulator for Intravehicular Space Robotics Research
ThI2I.183 Accelerated Multi-Modal Motion Planning Using Context-Conditioned Diffusion Models (CAMPD) 2510.14615
ThI2I.295 A Contact-Driven Framework for Manipulating in the Blind 2510.20177
ThI2I.300 GeoFIK: A Fast and Reliable Geometric Solver for the IK of the Franka Arm Based on Screw Theory 2503.03992
ThI2I.54 Distracted Robot: How Visual Clutter Undermine Robotic Manipulation 2511.22780
ThI2I.63 CAPE: Context-Aware Diffusion Policy via Proximal Mode Expansion for Collision Avoidance 2511.22773
TuAT3.4 DYMO-Hair: Generalizable Volumetric Dynamics Modeling for Robot Hair Manipulation 2510.06199
TuI1I.114 From CAD to POMDP: Probabilistic Planning for Robotic Disassembly of End-Of-Life Products 2511.23407
TuI1I.126 Planning Using Belief Summaries for Goal-Directed Manipulation of Articulated Objects with Force and Proprioception
TuI1I.211 Pack It In: Packing into Partially Filled Containers through Contact 2602.12095
TuI1I.244 Safety-Critical Dynamic Motion Generation for Manipulators Using Differentiable Distance Fields in Configuration Space 2412.16456
TuI1I.298 ConsistencyPlanner: Real-Time Planning with Fast-Sampling Consistency Models
TuI1I.415 TARAD: Task-Aware Robot Affordance-Centric Diffusion Policy Learned from LLM-Generated Demonstrations
TuI1I.47 KAN Policy: Learning Efficient and Smooth Robotic Trajectories via Kolmogorov-Arnold Networks
TuI2I.105 MO-SeGMan: Rearrangement Planning Framework for Multi-Objective Sequential and Guided Manipulation in Constrained Environments 2511.01476
TuI2I.12 Robust and Error-Tolerant Peg-In-Hole Assembly Using Simple Control
TuI2I.142 Learning Composable Skills by Discovering Spatial and Temporal Structure with Foundation Models (STACK)
TuI2I.143 Enhancing Classical Motion Planners Using RL with Safety Guarantees 2403.18524
TuI2I.15 H-MaP: An Iterative and Hybrid Sequential Manipulation Planner 2403.10436
TuI2I.201 DynDLO: Learning-Based Trajectory Planning for Dynamic Robotic Manipulation of Deformable Linear Objects
TuI2I.272 Run-Time Optimization of Overall Energy Consumption in Lightweight Collaborative Arms for Repetitive Tasks
TuI2I.303 Learning to Drive by Imitating Surrounding Vehicles 2503.05997
TuI2I.407 A Differential Dynamic Programming Framework for Inverse Reinforcement Learning 2407.19902
TuI2I.88 Robustness-Aware Tool Selection and Manipulation Planning with Learned Energy-Informed Guidance 2506.03362
WeBT1.8 SymSkill: Symbol and Skill Co-Invention for Data-Efficient and Reactive Long-Horizon Manipulation 2510.01661
WeBT2.2 Human2Nav: Learning Crowd Navigation from Human Videos across Robots via Feasibility-Guided Flow Matching
WeBT2.4 Shifted Flow Policy: Uncertainty-Aware Time Reparameterization for Visuomotor Learning
WeBT2.5 Closed-Loop Action Chunks with Dynamic Corrections for Training-Free Diffusion Policy (DCDP) 2603.01953
WeI1I.124 Placeit! A Framework for Learning Robot Object Placement Skills 2510.09267
WeI1I.179 Task Generalization with Pathwise Conditioning of Gaussian Process for Learning from Demonstration
WeI1I.187 Optimal Dexterity Path Planning for Robotic Manipulators Using Rapid Workspace Density Approximation
WeI1I.221 3DFacePolicy: Speech-Driven 3D Facial Animation Based on Diffusion Policy 2409.10848
WeI1I.256 MOASIC: Skill-Centric Manipulation Planning with Physics Simulation 2504.16738
WeI1I.339 AORRTC: Almost-Surely Asymptotically Optimal Planning with RRT-Connect 2505.10542
WeI1I.64 IMPACT: Intelligent Motion Planning with Acceptable Contact Trajectories via Vision-Language Models 2503.10110
WeI1I.68 Global Tensor Motion Planning 2411.19393
WeI2I.116 MetaDP: Meta-Manipulation Diffusion Policy for Robotic Manipulation
WeI2I.159 Manual2Skill++: Connector-Aware General Robotic Assembly from Instruction Manuals via Vision–Language Models 2510.16344
WeI2I.184 AdaptPNP: Integrating Prehensile and Non-Prehensile Skills for Adaptive Robotic Manipulation 2511.11052
WeI2I.259 Screw Geometry Meets Bandits: Incremental Acquisition of Demonstrations to Generate Manipulation Plans 2410.18275
WeI2I.300 GPU-Accelerated Continuous-Time Successive Convexification for Contact-Implicit Legged Locomotion 2604.09993

Related

  • ICRA 2026 Survey
  • ReKep — VLM-as-constraint-writer lineage referenced by IMPACT / affordance-constraint planners.

← Back to ICRA-2026-VLA-Manipulation-Survey