ICRA 2026 Topic Assembly - Heungwoo/research GitHub Wiki
ICRA 2026 β Assembly & Contact-Rich Manipulation (Topic Analysis)
Venue: IEEE ICRA 2026 Β· Vienna, Austria Β· June 1β5, 2026 Compiled against the official PaperCept program. This page covers the 58 papers in the assembly / contact-rich cluster β i.e. work tagged with Assembly, Compliant Assembly, Compliance and Impedance Control, Force Control, and Contact Modeling keywords. The unifying problem is manipulation where contact is the task: precision insertion and assembly, compliant/impedance interaction, and the physics modeling and estimation that make contact tractable.
This is one of ICRA's most "robotics-native" clusters: where the VLA survey tracks learned vision-language policies, this group is dominated by control theory, contact dynamics, and sim-to-real for precise physical interaction. The two worlds meet at the edges β flow-matching/diffusion policies that emit impedance parameters, force-feedback RL, and LLM-driven assembly planners β but the center of gravity here is force, compliance, and contact.
The 58 papers split roughly into: (a) learned insertion/assembly policies and their sim-to-real transfer; (b) classical and adaptive compliance / impedance / admittance control; (c) contact modeling, differentiable contact, and collision detection; (d) force/tactile-grounded contact-rich policies; and (e) LLM/VLM-based assembly reasoning and planning. Many papers straddle several of these.
Sub-trends
1. Learning insertion & assembly policies (sim-trained, real-deployed)
The largest analytical thread is simulation-trained assembly policies and how to close the last mile to the real world. Refinery (arXiv 2510.11019) frames the gap explicitly: sim policies plateau at ~80% success, short of industry standards, and it actively fine-tunes on high-uncertainty initial states (Bayesian-optimization-guided) plus deployment-time initialization selection (GMM sampling). SPARR (arXiv 2602.23253) attacks the same gap with an asymmetric residual: a sim-trained base policy on low-level state + dense reward, plus a real-world residual learned from vision + sparse reward. Multimodal Variational DeepMDP targets high-mix, low-volume industrial insertion via a learned multimodal latent dynamics representation aimed at transfer without retooling. DiSPo (arXiv 2409.14719) learns coarse-to-fine action discretization with a Diffusion-SSM (Mamba) policy for granular, multi-scale assembly skills. Together these show the field converging on hybrid sim-base + real-residual / active-fine-tuning recipes rather than pure zero-shot transfer.
2. Compliance, impedance & admittance control (the classical core)
This is numerically the densest sub-theme. On impedance: an Impedance Control Design Framework Using Commutative Map between SE(3) and se(3) tackles 6-DoF coupling via Lie-group structure; Iterative Learning-Based Centre-of-Mass Impedance Control targets articulated-soft humanoids. On admittance: Torque-Bounded Task-Space Admittance Control for Redundant Manipulators extends Kikuuwe's TBAC with explicit joint-torque limits; Decentralized Admittance Control for a Multi-Manipulator System regulates external and internal wrenches without a central unit; Predictive Admittance Control for Aerial Physical Interaction fuses admittance with NMPC; and Adaptive pHRI via a Passivity-Aware Model Predictive Variable Admittance Control (MPVA) balances tracking, compliance, and safety under variable human behavior. On passivity / energy methods: Passive Multi-Task Compliance Control with Strict Priority through Energy Tanks and Harmonising Safety Paradigms (energy-aware active-response + passive-compliance) both use energy-budget formalisms for safe interaction. The recurring concern across this group is stability and safety guarantees under contact, not just task success.
3. Force-feedback RL & imitation learning
A distinct cluster makes force and compliance first-class signals in learned policies. Flow with the Force Field (arXiv 2510.02738) trains a 3D flow-matching policy that consumes point cloud + force and predicts actions including an impedance parameter, executed via a passive impedance controller β generated from a single human demo via force-informed simulation. Safe and Optimal Variable Impedance Control via Certified RL (arXiv 2511.16330) introduces Certified Gaussian Manifold Sampling (C-GMS), learning combined DMP + VIC policies with Lyapunov stability and actuator feasibility guaranteed by construction. CoTaP (arXiv 2509.25443) brings compliance modulation on the SPD manifold into a two-stage RL pipeline for humanoid whole-body control. CG-THWM uses a curriculum-guided temporal haptic world model for fine-tolerance peg-in-hole RL, and RoboMT learns human-like compliance for connector assembly via a bilateral-teleoperation Mamba-Transformer. Data-Efficient Constrained Robot Learning with Probabilistic Lagrangian Control and HMC (heterogeneous meta-control for contact-rich loco-manipulation) round out the learned-compliance set.
4. Contact modeling, differentiable contact & collision detection
A strong simulation/physics sub-theme supplies the contact primitives the policies depend on. SDRS (Shape-Differentiable Robot Simulator) addresses non-differentiable singularities in differentiable simulation. A Convex Formulation of Compliant Contact between Filaments and Rigid Bodies handles codimensional 1D structures. Differentiable Contact Dynamics for Stable Object Placement under Geometric Uncertainties and Hybrid Contact Dynamics and Residual-RL for Multi-Point Object Pushing combine analytic contact models with learning. On collision detection: Amortized NeuralSDF-Mesh Collision Detection speeds non-convex queries, and Robust Differentiable Collision Detection for General Objects makes GJK+EPA-style witness-point computation differentiable for gradient-based optimization. Few-Shot Neural Differentiable Simulator does real-to-sim rigid-contact modeling from minimal real data β the inverse direction of the sim-to-real cluster.
5. Peg-in-hole, connectors & precision insertion tasks
The canonical contact-rich benchmark recurs throughout. CG-THWM explicitly targets fine-tolerance peg-in-hole under nonsmooth dynamics with irregular geometries and tight clearances. RoboMT focuses on electronic connector assembly where force regulation is paramount. Estimation of the Caged Object's Posture under Forces proposes a caging-based pinβhole assembly strategy that adjusts allowable relative pose via stepwise geometric calculation. Multimodal Variational DeepMDP and SPARR both use insertion as their core evaluation. This sub-trend underscores that insertion remains the field's stress test for precision + compliance + perception jointly.
6. Tactile / visuo-tactile & force for contact-rich perception
Perception-side work grounds control in contact sensing. M-VTOP (Modular Visuo-Tactile Object Pose Estimation) targets high-precision pose for small/intricate parts under occlusion and noise. TwinTrack (arXiv 2505.22882) bridges vision and contact physics for real-time 6-DoF tracking of unknown objects in contact-rich scenes via a Real2Sim/Sim2Real loop on a GPU-accelerated physics engine. GaussTwin uses Gaussian-splatting digital twins to close the real-to-sim gap for dynamic interactions. TIGeR (arXiv 2506.00953) does template-free hand-object reconstruction via text-instructed shape priors. These show contact perception fusing vision with physics priors rather than relying on either alone.
7. Sim-to-real & digital twins for contact
Beyond individual policies, several papers attack the transfer problem structurally. SPARR and Refinery (above) are the policy-side; Few-Shot Neural Differentiable Simulator, GaussTwin, and TwinTrack are the model-side, all trying to make simulated contact match real contact with little real data. The convergent insight is that contact physics is the hardest thing to transfer, so the community is investing in differentiable, correctable, few-shot-tunable simulators rather than betting on one-shot domain randomization.
8. LLM/VLM-driven assembly reasoning & planning
A newer, language-centric cluster treats assembly as a reasoning and planning problem. ActionReasoning uses an LLM orchestrator that decomposes brick-stacking into agents generating waypoints, updating a world model from 3D-scene changes. AssemMate is a graph-based LLM for assembly assistance; IDfRA adds LLM self-verification for iterative Design-for-Robotic-Assembly; GLaMP is a grounded multi-agent LLM for long-horizon industrial planning; MICA is a speech-interactive multi-agent industrial coordination assistant; GPT-PDDL and Compositional Context Fine-Tuning of VLMs target executable task plans and assembly-action understanding from video. This cluster is where the assembly group most directly touches the VLA wave covered in the main survey.
Standout deep-dives
Refinery β closing the sim-policy "last mile" for assembly (arXiv 2510.11019)
Sim-trained contact-rich policies stall around ~80% success, below industry needs. Refinery does two things: (1) active fine-tuning that identifies high-uncertainty initial states and re-trains on them via Bayesian-optimization-guided sampling; (2) deployment-time optimization that uses a GMM over initializations to prioritize high-success starts. Reported gains: +10.98% mean success over prior SOTA, reaching 91.51% in simulation, and fine-tuned policies chain to assemble up to 8 parts without explicit multi-step training. Authors include Bingjie Tang, Iretiayo Akinola, Jie Xu (NVIDIA-adjacent assembly lineage). This is the clearest articulation of the "sim is good, deployment is the bottleneck" thesis.
SPARR β asymmetric real-world residuals for assembly (arXiv 2602.23253)
SPARR pairs a sim base policy (low-level state, dense reward) with a real-world residual policy (vision, sparse reward) β the asymmetry is deliberate: sim gives strong priors cheaply, the residual corrects dynamics/sensor mismatch with zero human supervision. Reported: 95β100% success on real assembly tasks, +38.4% over SOTA zero-shot sim-to-real, and β29.7% cycle time. Authors include Yijie Guo, Iretiayo Akinola, Yashraj Narang. A clean counterpoint to Refinery: residual-RL vs. active-fine-tuning as two routes across the same gap.
Flow with the Force Field β compliant flow-matching policies (arXiv 2510.02738)
A 3D flow-matching policy that takes point cloud + force and predicts actions including an impedance parameter, synthesized into a state-velocity field and run through a Passive Impedance Controller. Training data is force-informed simulation generated from a single human demonstration. Result: zero-shot deployment on real Franka arms for contact-rich tasks with no real-world training data. This is the bridge paper between the VLA/flow-policy world and classical impedance control β compliance becomes a predicted output, not a fixed gain.
Safe and Optimal Variable Impedance Control via Certified RL (arXiv 2511.16330)
Combines DMPs (motion) + Variable Impedance Control (compliance) and learns both with Certified Gaussian Manifold Sampling (C-GMS) β policies are sampled via Gaussian perturbations confined to a manifold where Lyapunov stability and actuator feasibility are enforced analytically, eliminating penalty terms, barriers, or post-hoc projection. The model-free-learning-with-model-based-guarantees framing is exactly the safety-under-contact concern that defines sub-trend 2.
TwinTrack β vision + contact physics for tracking unknown objects (arXiv 2505.22882)
Real-time 6-DoF pose tracking of unseen dynamic objects in contact-rich scenes, where pure vision fails under occlusion and motion blur. Real2Sim jointly estimates geometry and physical parameters (mass, inertia, friction) from vision + contact-dynamics consistency; Sim2Real fuses visual tracking with contact simulation on a GPU-accelerated physics engine for real-time performance. A strong example of contact perception that refuses to choose between vision and physics.
CoTaP β compliance modulation for humanoid control (arXiv 2509.25443)
A Compliant Task Pipeline that injects compliance into learning-based humanoid control despite human datasets lacking measured force. Two-stage dual-agent RL: a position-based base policy, then a distillation stage where upper-body policy is combined with model-based compliance control on the SPD manifold (with the lower body guided by the base policy). Shows how the compliance-control community is absorbing learning while keeping stability guarantees.
Complete paper list (58)
| Code | Title | arXiv |
|---|---|---|
| ThBT1.9 | A Kinesthetic Teaching Framework for Tasks with Contact Transitions and Time-Optimized Execution | β |
| ThBT3.9 | A Gripper for Flap Separation and Opening of Sealed Bags | 2603.10890 |
| ThI1I.108 | PaiP: An Operational Aware Interactive Planner for Unknown Cabinet Environments | 2509.11516 |
| ThI1I.129 | Refinery: Active Fine-Tuning and Deployment-Time Optimization for Contact-Rich Policies | 2510.11019 |
| ThI1I.209 | Adaptive Physical HumanβRobot Interaction Via a Passivity-Aware Model Predictive Variable Admittance Control | β |
| ThI1I.220 | M-VTOP: Modular Visuo-Tactile Object Pose Estimation for High-Precision Robotic Manipulation | β |
| ThI1I.266 | SPARR: Simulation-Based Policies with Asymmetric Real-World Residuals for Assembly | 2602.23253 |
| ThI1I.281 | Compositional Context Fine-Tuning Vision-Language Model for Complex Assembly Action Understanding from Videos | β |
| ThI1I.385 | Hybrid Contact Dynamics and Residual-RL Framework for Multi-Point Object Pushing | β |
| ThI1I.398 | Impedance Control Design Framework Using Commutative Map between SE(3) and se(3) | β |
| ThI1I.405 | Predictive Admittance Control for Aerial Physical Interaction | β |
| ThI1LB.15 | Grasping Point Estimation for EA Suction Cup Grippers on Curved Objects | β |
| ThI2I.1 | Multimodal Variational DeepMDP: An Efficient Approach for Industrial Assembly in High-Mix, Low-Volume Production | β |
| ThI2I.205 | Decentralized Admittance Control for a Multiβmanipulator System: Theory and Experiments | β |
| ThI2I.282 | CG-THWM: Curriculum-Guided Temporal Haptic World Modeling for Peg-In-Hole Tasks | β |
| ThI2I.304 | GaussTwin: Unified Simulation and Correction with Gaussian Splatting for Robotic Digital Twins | 2603.05108 |
| ThI2I.325 | Passive Multi-Task Compliance Control with Strict Priority through Energy Tanks | β |
| ThI2I.357 | SDRS: Shape-Differentiable Robot Simulator | 2412.19127 |
| ThI2I.395 | Data-Efficient Constrained Robot Learning with Probabilistic Lagrangian Control | β |
| ThI2I.407 | Harmonising Safety Paradigms: Energy-Aware Control of Active Response and Passive Compliance for Safety-Critical Robotic Tasks | β |
| ThI2I.45 | TIGeR: Text-Instructed Generation and Refinement for Template-Free Hand-Object Interaction | 2506.00953 |
| TuAT1.5 | RCM Constraint-Consistent Dynamic Control in Surgical Robots | 2509.14075 |
| TuBT3.9 | Differentiable Contact Dynamics for Stable Object Placement under Geometric Uncertainties | 2409.17725 |
| TuBT4.1 | Morphogenetic Assembly and Adaptive Control for Heterogeneous Modular Robots | 2602.10561 |
| TuI1I.122 | Bending Perception-Based Variable Stiffness Control for Snake Robots in Pipe Navigation | β |
| TuI1I.14 | Limiting Kinetic Energy through Control Barrier Functions: Analysis and Experimental Validation | 2411.02186 |
| TuI1I.163 | DiSPo: Diffusion-SSM Based Policy Learning for Coarse-To-Fine Action Discretization | 2409.14719 |
| TuI1I.177 | ActionReasoning: Robot Action Reasoning in 3D Space with LLM for Robotic Brick Stacking | 2602.21157 |
| TuI1I.216 | A Convex Formulation of Compliant Contact between Filaments and Rigid Bodies | 2509.13434 |
| TuI1I.231 | IDfRA: Self-Verification for Iterative Design in Robotic Assembly | 2509.16998 |
| TuI1I.291 | Sym-Servo: Disambiguate Symmetric Object Pose by End-To-End Optimal Visual Servo | β |
| TuI1I.360 | Torque-Bounded Task-Space Admittance Control for Redundant Manipulators | β |
| TuI1LB.11 | Suppressing Initial Force Overshoot Using Admittance Filter and ASMC under Contact Location Uncertainty | β |
| TuI2I.226 | Iterative Learning-Based Centre-Of-Mass Impedance Control for Articulated-Soft Humanoid Robots | β |
| TuI2I.233 | Amortized NeuralSDF-Mesh Collision Detection for Robotic Contact Simulation | β |
| TuI2I.248 | TwinTrack: Bridging Vision and Contact Physics for Real-Time Tracking of Unknown Objects in Contact-Rich Scenes | 2505.22882 |
| TuI2I.283 | Stroke-Based Variable-Damping with Force Attenuation for Capturing Large-Momentum Objects under Non-Zero Contact Velocity | β |
| TuI2I.291 | Safe and Optimal Variable Impedance Control Via Certified Reinforcement Learning | 2511.16330 |
| TuI2I.406 | Physics-Informed Passive Motion Paradigm for Parallel Robots: A High-Precision Motor-Primitives Framework | β |
| TuI2I.5 | Semi-Autonomous Teleoperation Using Differential Flatness of a Crane Robot for Aircraft In-Wing Inspection | 2412.10973 |
| TuI2LB.23 | GPT-PDDL: Towards Executable Robot Task Planning | β |
| TuI2LB.4 | GLaMP: A Grounded Language Model-Based Multi-Agent System for Long-Horizon Robotic Task Planning in Industrial Settings | β |
| WeAT4.1 | On Robust Coordinated Compliant Control Design for Space Manipulators under Flexible and Uncertain Dynamics | β |
| WeAT4.9 | Astrobee: Free-Flying Robots for the International Space Station (I) | β |
| WeI1I.121 | AssemMate: Graph-Based LLM for Robotic Assembly Assistance | 2509.11617 |
| WeI1I.172 | Bipedal-Walking-Dynamics Model on Granular Terrains | 2604.11981 |
| WeI1I.375 | Flexible-Link Velocity-Bounding Proxy Based Sliding Mode Control | β |
| WeI1I.391 | Nullspace Optimization of Redundant Robots for Dynamics Decoupling in Motion Force Control | β |
| WeI1I.436 | Estimation of the Caged Object's Posture under Forces Using Stepwise Geometric Calculations | β |
| WeI1I.9 | Stable Object Placement Planning from Contact Point Robustness | 2410.12483 |
| WeI2I.121 | CoTaP: Compliant Task Pipeline and Reinforcement Learning of Its Controller with Compliance Modulation | 2509.25443 |
| WeI2I.236 | Few-Shot Neural Differentiable Simulator: Real-To-Sim Rigid-Contact Modeling | 2603.06218 |
| WeI2I.248 | MICA: Multi-Agent Industrial Coordination Assistant | 2509.15237 |
| WeI2I.263 | Robust Differentiable Collision Detection for General Objects | 2511.06267 |
| WeI2I.319 | Flow with the Force Field: Learning 3D Compliant Flow Matching Policies from Force and Demonstration-Guided Simulation Data | 2510.02738 |
| WeI2I.373 | RoboMT: Human-Like Compliance Control for Assembly Via a Bilateral Robotic Teleoperation and Hybrid Mamba-Transformer Framework | β |
| WeI2I.397 | Robotic Harvesting of Delicate Fruit: Design and Implementation of an Under-Actuated Disturbance-Resistant Gripper | β |
| WeI2I.88 | HMC: Learning Heterogeneous Meta-Control for Contact-Rich Loco-Manipulation | 2511.14756 |
Related
β Back to ICRA-2026-VLA-Manipulation-Survey