ICRA 2026 Topic Mobile - Heungwoo/research GitHub Wiki
ICRA 2026 — Mobile Manipulation (Topic Analysis)
Venue: IEEE ICRA 2026 · Vienna, Austria · June 1–5, 2026 Compiled against the official PaperCept program. This page covers the 34 papers that carry Mobile Manipulation as a topic keyword in the technical program.
Mobile manipulation is the part of ICRA 2026 where the field's two halves — a base that moves and an arm that manipulates — have to be solved together. The 34 papers here are united less by a single method than by a single hard problem: coordinating base and arm degrees of freedom under one objective, so that a robot can do long-horizon household, service, and industrial tasks that no fixed-base manipulator could reach. The cluster is strikingly systems-first — most papers report real-robot results on holonomic bases, differential-drive carts, quadrupeds, humanoids, ROVs, or aerial platforms — and it spans the full stack from torque-level whole-body control through learned visuomotor policies, navigation-for-manipulation, and TAMP. Compared with the broader VLA wave (see the ICRA 2026 Survey), this cluster's center of gravity is embodiment and control: how to map a desired end-effector behavior onto a redundant, often nonholonomic, sometimes floating base — and how to collect data for that without prohibitively expensive whole-body teleoperation.
Sub-trends
1. Whole-body base–arm coordination (control & MPC)
The largest and most classical sub-theme: treat the base+arm as one redundant system and optimize the whole thing. Reactive Whole-Body Control of Mobile Manipulators does optimization-based eye-in-hand tracking of moving targets in an unbounded workspace via adaptive-predictive visual servoing; SM²ITH (arXiv 2511.17798) layers interactive human prediction onto Hierarchical-Task MPC through a bilevel optimization that jointly accounts for robot and human dynamics, validated on Stretch 3 and Ridgeback-UR10. Robust Nonprehensile Object Transportation (arXiv 2411.07079) tackles the "waiter's problem" with moment-relaxation robust constraints so a tray-carried object with uncertain inertia never slips. Inverse Reachability Map Guided Motion Planning chooses base poses that keep the arm in high-manipulability configurations, and Generating and Optimizing Topologically Distinct Guesses (arXiv 2410.20635) escapes local optima in constrained mobile-manipulator path planning by optimizing homotopically distinct seeds.
2. Learned visuomotor policies for mobile manipulation
A fast-growing cluster putting diffusion/imitation policies on a moving base. HoMeR (arXiv 2506.01185) combines a kinematics-based whole-body controller with hybrid absolute/relative action modes, hitting 79.17% on in-the-wild home tasks from just 20 demos/task. M⁴Diffuser (arXiv 2509.14980) pairs a multi-view diffusion policy with a manipulability-aware reduced QP (ReM-QP) controller, reporting 7–56% higher success and fewer collisions. MIMO uses exoskeleton-VR teleoperation plus multi-receptive-field visual fusion for long-horizon imitation; LeGO-MM distills a hierarchical policy for goal-oriented MM. These differ from full VLAs (covered in the survey) by keeping a learned high level + analytic low level split rather than predicting joint actions end-to-end.
3. Learning from human / teleop data — escaping the teleoperation bottleneck
A pointed response to the fact that mobile teleoperation is expensive. EMMA (arXiv 2509.04443) co-trains egocentric human full-body data (captured with Project Aria glasses) with static robot data, matching Mobile-ALOHA-style baselines without any mobile teleoperation and scaling with hours of human data. UMI-on-Air (arXiv 2510.02614) takes embodiment-agnostic policies trained on handheld-gripper (UMI) human demos and deploys them on a constrained aerial manipulator via an Embodiment-Aware Diffusion Policy that couples the high-level policy to a low-level embodiment-specific controller. MIMO (above) is the teleop-hardware answer to the same problem.
4. Navigation-for-manipulation, search & open-vocabulary interaction
Mobile manipulation where the hard part is finding and reaching the object. BINDER (arXiv 2511.22364) is a dual-process open-vocabulary system: a deliberative MLLM planner with structured 3D scene updates guides an instant Video-LLM monitor that corrects actions and triggers replanning as the scene changes. Searching in Space and Time / STAR (arXiv 2511.14004) unifies memory retrieval ("the mug that was here yesterday") with embodied search actions for open-world object retrieval, with the STARBench benchmark. CMAR-Search adds commonsense + memory-augmented reasoning over 3D scene graphs to find functional storage areas; VLION does vision-language-guided interactive object navigation for occluded targets (behind doors, inside containers); and the Mobile Manipulation Instruction Generation paper (arXiv 2501.17022) learns to generate free-form instructions from target+receptacle images with metric-augmented training.
5. Loco-manipulation on legged & humanoid bases
A distinct hardware family — the "base" is a quadruped or biped, so locomotion stability and manipulation must be co-optimized. RAMBO (arXiv 2504.06662) fuses model-based whole-body QP feedforward torques with an RL feedback policy on a Unitree Go2, doing cart-pushing, plate-balancing in quadrupedal and bipedal modes. Whole-Body Inverse Dynamics MPC (arXiv 2511.19709) optimizes joint torques through full-order inverse dynamics at 80 Hz on a B2+Z1, pulling loads and wiping whiteboards. SEEC (arXiv 2509.21231) uses model-enhanced residual learning to stabilize a humanoid arm end-effector against lower-body disturbances. CAIMAN (arXiv 2502.00835) uses causal action-influence as an intrinsic reward for sample-efficient legged non-prehensile pushing.
6. Cooperative & multi-robot mobile manipulation
Several papers move from one robot to teams sharing a payload. Multi-Quadruped Cooperative Object Transport (arXiv 2509.14342) learns decentralized pinch-lift-move with a constellation reward; a single 2-robot-trained policy transfers to teams of 2–10. Cooperative Grasping for Collective Object Transport (arXiv 2509.03638) uses a Conditional Embedding model to pick feasible two-robot grasp configurations in constrained spaces. Heterogeneous Skill Learning for Asynchronous Multi-Robot Relay Pushing builds a room/corridor/helper skill library for relay transport, and Virtual-Force Based Visual Servo (arXiv 2407.10570) couples multiple manipulators for multi-peg-in-hole assembly (85% at 0.2 mm clearance).
7. Robustness, uncertainty & non-rigid / non-prehensile manipulation
A control-theory thread on acting under uncertainty and partial observability. CURA-PPO (arXiv 2602.01731) treats object-induced sensor occlusion as a distribution over collision risk, using uncertainty to drive active perception (up to 3× higher success). RAPiD (arXiv 2603.18246) rapidly adapts a learned particle-dynamics model for deformable-object MM (+65% success on unseen 1D/2D objects). Uncertainty-Aware Adaptive Dynamics (arXiv 2603.06548) does physically-consistent online parameter estimation for an underwater vehicle-manipulator (BlueROV2 + 4-DOF arm). SHOPPER (arXiv 2504.12512) is the field-test reality check — hundreds of grocery-store pick attempts and their failure modes.
8. Design, software & systems for mobile manipulators
The enabling infrastructure. Task-Driven Co-Design of Mobile Manipulators (arXiv 2412.16635) jointly optimizes arm-mounting parameters with an RL policy via BOHB. TASP / Beyond Task and Motion Planning (arXiv 2504.17901) wraps black-box skills as Composable Interaction Primitives for hierarchical planning. From Composable Models to Correct-by-Construction Software gives a graph-structured interchange format generating correct-by-construction code for contact-rich MM. TopAY (arXiv 2507.02761) is an efficient differential-drive MM trajectory planner; Serving Innovation (MOMO) reconfigures serving robots into mobile manipulators; and the Nonlinear Predictive Control of a Suspended Deformable Cable (arXiv 2602.17199) extends MM to aerial pick-and-place with a PDE/ROM cable model.
Standout deep-dives
HoMeR — hybrid imitation + whole-body control (arXiv 2506.01185)
A clean recipe for in-the-wild home MM: a fast kinematics-based whole-body controller maps desired end-effector poses to coordinated base+arm motion, while the learned policy switches between absolute poses for long-range moves and relative poses for fine manipulation. On a holonomic 7-DoF mobile manipulator across 3 sim + 3 real household tasks (opening cabinets, sweeping trash, rearranging pillows), HoMeR reaches 79.17% success from 20 demos/task, beating the next-best baseline by 29.17% on average, and is VLM-compatible for novel-object generalization.
EMMA — scaling MM via egocentric human data (arXiv 2509.04443)
Directly attacks the mobile-teleoperation bottleneck. EMMA co-trains egocentric human mobile-manipulation data (Project Aria glasses) with static robot teleop data, sidestepping costly mobile-robot teleoperation entirely. Across three real-world tasks it matches or beats Mobile-ALOHA-style baselines trained on teleoperated mobile-robot data, generalizes to new spatial configurations/scenes, and shows positive scaling as human-data hours grow — a data-source argument as much as a method.
RAMBO — RL-augmented model-based whole-body control (arXiv 2504.06662)
The cleanest statement of the "model-based + RL" loco-manipulation design point. A model-based module solves a QP for feedforward torques; an RL policy supplies feedback corrective terms for robustness to unmodeled dynamics. Validated on a Unitree Go2 across pushing a shopping cart, balancing a plate, and holding soft objects in both quadrupedal and bipedal modes. Code at github.com/catachiii/rambo.
M⁴Diffuser — multi-view diffusion + manipulability-aware QP (arXiv 2509.14980)
Addresses the perception/control split for MM: a multi-view diffusion policy consumes proprioception + complementary camera views to emit world-frame end-effector goals, executed by a Reduced and Manipulability-aware QP (ReM-QP) that drops slack variables for speed and biases away from singularities. Reports 7–56% higher success and 3–31% fewer collisions than baselines in sim and real.
CURA-PPO — uncertainty-aware non-prehensile MM under occlusion (arXiv 2602.01731)
When the pushed object blocks the onboard sensor, occluded regions cause collisions. CURA-PPO predicts collision possibility as a distribution, extracting both risk and uncertainty; the uncertainty term rewards active perception, so the robot simultaneously manipulates and gathers information to resolve occlusion — up to 3× higher success than baselines across object sizes and obstacle layouts.
UMI-on-Air — embodiment-aware guidance for agnostic policies (arXiv 2510.02614)
A transfer story: train an embodiment-agnostic visuomotor policy on unconstrained handheld-gripper (UMI) human demos, then deploy on a constrained aerial manipulator. The Embodiment-Aware Diffusion Policy (EADP) couples the high-level UMI policy to a low-level embodiment-specific controller at inference time, avoiding the OOD/poor-execution failure of naive transfer — improving success, efficiency, and robustness under disturbance on long-horizon, high-precision aerial tasks.
Complete paper list (34)
| Code | Title | arXiv |
|---|---|---|
| ThI1I.105 | MIMO: A Multimodal Imitation Learning Framework for Mobile Manipulation with Exoskeleton-VR Teleoperation | — |
| ThI1I.313 | CMAR-Search: Commonsense and Memory Augmented Reasoning for Object Search in Dynamic Interactive Environments | — |
| ThI1I.361 | Generating and Optimizing Topologically Distinct Guesses for Mobile Manipulator Path Planning with Path Constraints | 2410.20635 |
| ThI1I.408 | From Composable Models to Correct-By-Construction Software for Contact-Rich Robotic Mobile-Manipulation Tasks | — |
| ThI2I.19 | RAMBO: RL-Augmented Model-Based Whole-Body Control for Loco-Manipulation | 2504.06662 |
| ThI2I.221 | HoMeR: Learning In-The-Wild Mobile Manipulation Via Hybrid Imitation and Whole-Body Control | 2506.01185 |
| ThI2I.319 | Task and Skill Planning: Hierarchical Robot Planning with Black-Box Skills | 2504.17901 |
| ThI2I.7 | Task-Driven Co-Design of Mobile Manipulators | 2412.16635 |
| ThI2LB.12 | Heterogeneous Skill Learning for Asynchronous Multi-Robot Relay Pushing in Complex Environments | — |
| ThI2LB.14 | Inverse Reachability Map Guided Motion Planning of Mobile Manipulator | — |
| TuBT2.4 | Nonlinear Predictive Control of the Continuum and Hybrid Dynamics of a Suspended Deformable Cable for Aerial Pick and Place | 2602.17199 |
| TuI1I.251 | SEEC: Stable End-Effector Control with Model-Enhanced Residual Learning for Humanoid Loco-Manipulation | 2509.21231 |
| TuI1I.268 | Uncertainty-Aware Adaptive Dynamics for Underwater Vehicle–Manipulator Robots | 2603.06548 |
| TuI1I.318 | BINDER: Instantly Adaptive Mobile Manipulation with Open-Vocabulary Commands | 2511.22364 |
| TuI1I.34 | Rapid Adaptation of Particle Dynamics for Generalized Deformable Object Mobile Manipulation | 2603.18246 |
| TuI1I.423 | Virtual-Force Based Visual Servo for Multiple Peg-In-Hole Assembly with Tightly Coupled Multi-Manipulator | 2407.10570 |
| TuI1I.58 | TopAY: Efficient Trajectory Planning for Differential Drive Mobile Manipulators Via Topological Paths Search and Arc Length-Yaw Parameterization | 2507.02761 |
| TuI2I.166 | M4Diffuser: Multi-View Diffusion Policy with Manipulability-Aware Control for Robust Mobile Manipulation | 2509.14980 |
| TuI2I.2 | Mobile Manipulation Instruction Generation from Multiple Images with Automatic Metric Enhancement | 2501.17022 |
| TuI2I.258 | Searching in Space and Time: Unified Memory-Action Loops for Open-World Object Retrieval | 2511.14004 |
| TuI2I.419 | Cooperative Grasping for Collective Object Transport in Constrained Environments | 2509.03638 |
| WeBT2.6 | EMMA: Scaling Mobile Manipulation Via Egocentric Human Data | 2509.04443 |
| WeI1I.1 | Robust Nonprehensile Object Transportation with Uncertain Inertial Parameters | 2411.07079 |
| WeI1I.104 | SHOPPER: Practical Insights on Grasp Strategies for Mobile Manipulation in the Wild | 2504.12512 |
| WeI1I.175 | VLION: Vision-Language Guided Interactive Object Navigation with Mobile Manipulation | — |
| WeI1I.220 | UMI-On-Air: Embodiment-Aware Guidance for Embodiment-Agnostic Visuomotor Policies | 2510.02614 |
| WeI1I.267 | Multi-Quadruped Cooperative Object Transport: Learning Decentralized Pinch-Lift-Move | 2509.14342 |
| WeI1I.324 | LeGO-MM: Learning Navigation for Goal-Oriented Mobile Manipulation Via Hierarchical Policy Distillation | — |
| WeI1I.67 | CAIMAN: Causal Action Influence Detection for Sample-Efficient Loco-Manipulation | 2502.00835 |
| WeI1I.76 | Reactive Whole-Body Control of Mobile Manipulators for Dynamic Target Tracking Via Adaptive-Predictive Visual Servoing | — |
| WeI2I.280 | Uncertainty-Aware Non-Prehensile Manipulation with Mobile Manipulators under Object-Induced Occlusion (CURA-PPO) | 2602.01731 |
| WeI2I.293 | SM²ITH: Safe Mobile Manipulation with Interactive Human Prediction Via Task-Hierarchical Bilevel Model Predictive Control | 2511.17798 |
| WeI2I.407 | Serving Innovation: Seamless Service by Advancing Food Runners with Mobile Manipulation (MOMO) | — |
| WeI2I.426 | Whole-Body Inverse Dynamics MPC for Legged Loco-Manipulation | 2511.19709 |
Dashes mark papers for which no public arXiv preprint was confirmed (several are IEEE RA-L / Xplore-only or too recent to be indexed). Numbers in the deep-dives and sub-trends are quoted from the corresponding arXiv/program abstracts; figures are not asserted for papers without a confirmed source.
Related
← Back to ICRA-2026-VLA-Manipulation-Survey