ICRA 2026 Topic Mobile - Heungwoo/research GitHub Wiki

ICRA 2026 — Mobile Manipulation (Topic Analysis)

Venue: IEEE ICRA 2026 · Vienna, Austria · June 1–5, 2026 Compiled against the official PaperCept program. This page covers the 34 papers that carry Mobile Manipulation as a topic keyword in the technical program.

Mobile manipulation is the part of ICRA 2026 where the field's two halves — a base that moves and an arm that manipulates — have to be solved together. The 34 papers here are united less by a single method than by a single hard problem: coordinating base and arm degrees of freedom under one objective, so that a robot can do long-horizon household, service, and industrial tasks that no fixed-base manipulator could reach. The cluster is strikingly systems-first — most papers report real-robot results on holonomic bases, differential-drive carts, quadrupeds, humanoids, ROVs, or aerial platforms — and it spans the full stack from torque-level whole-body control through learned visuomotor policies, navigation-for-manipulation, and TAMP. Compared with the broader VLA wave (see the ICRA 2026 Survey), this cluster's center of gravity is embodiment and control: how to map a desired end-effector behavior onto a redundant, often nonholonomic, sometimes floating base — and how to collect data for that without prohibitively expensive whole-body teleoperation.

Sub-trends

1. Whole-body base–arm coordination (control & MPC)

The largest and most classical sub-theme: treat the base+arm as one redundant system and optimize the whole thing. Reactive Whole-Body Control of Mobile Manipulators does optimization-based eye-in-hand tracking of moving targets in an unbounded workspace via adaptive-predictive visual servoing; SM²ITH (arXiv 2511.17798) layers interactive human prediction onto Hierarchical-Task MPC through a bilevel optimization that jointly accounts for robot and human dynamics, validated on Stretch 3 and Ridgeback-UR10. Robust Nonprehensile Object Transportation (arXiv 2411.07079) tackles the "waiter's problem" with moment-relaxation robust constraints so a tray-carried object with uncertain inertia never slips. Inverse Reachability Map Guided Motion Planning chooses base poses that keep the arm in high-manipulability configurations, and Generating and Optimizing Topologically Distinct Guesses (arXiv 2410.20635) escapes local optima in constrained mobile-manipulator path planning by optimizing homotopically distinct seeds.

2. Learned visuomotor policies for mobile manipulation

A fast-growing cluster putting diffusion/imitation policies on a moving base. HoMeR (arXiv 2506.01185) combines a kinematics-based whole-body controller with hybrid absolute/relative action modes, hitting 79.17% on in-the-wild home tasks from just 20 demos/task. M⁴Diffuser (arXiv 2509.14980) pairs a multi-view diffusion policy with a manipulability-aware reduced QP (ReM-QP) controller, reporting 7–56% higher success and fewer collisions. MIMO uses exoskeleton-VR teleoperation plus multi-receptive-field visual fusion for long-horizon imitation; LeGO-MM distills a hierarchical policy for goal-oriented MM. These differ from full VLAs (covered in the survey) by keeping a learned high level + analytic low level split rather than predicting joint actions end-to-end.

3. Learning from human / teleop data — escaping the teleoperation bottleneck

A pointed response to the fact that mobile teleoperation is expensive. EMMA (arXiv 2509.04443) co-trains egocentric human full-body data (captured with Project Aria glasses) with static robot data, matching Mobile-ALOHA-style baselines without any mobile teleoperation and scaling with hours of human data. UMI-on-Air (arXiv 2510.02614) takes embodiment-agnostic policies trained on handheld-gripper (UMI) human demos and deploys them on a constrained aerial manipulator via an Embodiment-Aware Diffusion Policy that couples the high-level policy to a low-level embodiment-specific controller. MIMO (above) is the teleop-hardware answer to the same problem.

4. Navigation-for-manipulation, search & open-vocabulary interaction

Mobile manipulation where the hard part is finding and reaching the object. BINDER (arXiv 2511.22364) is a dual-process open-vocabulary system: a deliberative MLLM planner with structured 3D scene updates guides an instant Video-LLM monitor that corrects actions and triggers replanning as the scene changes. Searching in Space and Time / STAR (arXiv 2511.14004) unifies memory retrieval ("the mug that was here yesterday") with embodied search actions for open-world object retrieval, with the STARBench benchmark. CMAR-Search adds commonsense + memory-augmented reasoning over 3D scene graphs to find functional storage areas; VLION does vision-language-guided interactive object navigation for occluded targets (behind doors, inside containers); and the Mobile Manipulation Instruction Generation paper (arXiv 2501.17022) learns to generate free-form instructions from target+receptacle images with metric-augmented training.

5. Loco-manipulation on legged & humanoid bases

A distinct hardware family — the "base" is a quadruped or biped, so locomotion stability and manipulation must be co-optimized. RAMBO (arXiv 2504.06662) fuses model-based whole-body QP feedforward torques with an RL feedback policy on a Unitree Go2, doing cart-pushing, plate-balancing in quadrupedal and bipedal modes. Whole-Body Inverse Dynamics MPC (arXiv 2511.19709) optimizes joint torques through full-order inverse dynamics at 80 Hz on a B2+Z1, pulling loads and wiping whiteboards. SEEC (arXiv 2509.21231) uses model-enhanced residual learning to stabilize a humanoid arm end-effector against lower-body disturbances. CAIMAN (arXiv 2502.00835) uses causal action-influence as an intrinsic reward for sample-efficient legged non-prehensile pushing.

6. Cooperative & multi-robot mobile manipulation

Several papers move from one robot to teams sharing a payload. Multi-Quadruped Cooperative Object Transport (arXiv 2509.14342) learns decentralized pinch-lift-move with a constellation reward; a single 2-robot-trained policy transfers to teams of 2–10. Cooperative Grasping for Collective Object Transport (arXiv 2509.03638) uses a Conditional Embedding model to pick feasible two-robot grasp configurations in constrained spaces. Heterogeneous Skill Learning for Asynchronous Multi-Robot Relay Pushing builds a room/corridor/helper skill library for relay transport, and Virtual-Force Based Visual Servo (arXiv 2407.10570) couples multiple manipulators for multi-peg-in-hole assembly (85% at 0.2 mm clearance).

7. Robustness, uncertainty & non-rigid / non-prehensile manipulation

A control-theory thread on acting under uncertainty and partial observability. CURA-PPO (arXiv 2602.01731) treats object-induced sensor occlusion as a distribution over collision risk, using uncertainty to drive active perception (up to 3× higher success). RAPiD (arXiv 2603.18246) rapidly adapts a learned particle-dynamics model for deformable-object MM (+65% success on unseen 1D/2D objects). Uncertainty-Aware Adaptive Dynamics (arXiv 2603.06548) does physically-consistent online parameter estimation for an underwater vehicle-manipulator (BlueROV2 + 4-DOF arm). SHOPPER (arXiv 2504.12512) is the field-test reality check — hundreds of grocery-store pick attempts and their failure modes.

8. Design, software & systems for mobile manipulators

The enabling infrastructure. Task-Driven Co-Design of Mobile Manipulators (arXiv 2412.16635) jointly optimizes arm-mounting parameters with an RL policy via BOHB. TASP / Beyond Task and Motion Planning (arXiv 2504.17901) wraps black-box skills as Composable Interaction Primitives for hierarchical planning. From Composable Models to Correct-by-Construction Software gives a graph-structured interchange format generating correct-by-construction code for contact-rich MM. TopAY (arXiv 2507.02761) is an efficient differential-drive MM trajectory planner; Serving Innovation (MOMO) reconfigures serving robots into mobile manipulators; and the Nonlinear Predictive Control of a Suspended Deformable Cable (arXiv 2602.17199) extends MM to aerial pick-and-place with a PDE/ROM cable model.

Standout deep-dives

HoMeR — hybrid imitation + whole-body control (arXiv 2506.01185)

A clean recipe for in-the-wild home MM: a fast kinematics-based whole-body controller maps desired end-effector poses to coordinated base+arm motion, while the learned policy switches between absolute poses for long-range moves and relative poses for fine manipulation. On a holonomic 7-DoF mobile manipulator across 3 sim + 3 real household tasks (opening cabinets, sweeping trash, rearranging pillows), HoMeR reaches 79.17% success from 20 demos/task, beating the next-best baseline by 29.17% on average, and is VLM-compatible for novel-object generalization.

EMMA — scaling MM via egocentric human data (arXiv 2509.04443)

Directly attacks the mobile-teleoperation bottleneck. EMMA co-trains egocentric human mobile-manipulation data (Project Aria glasses) with static robot teleop data, sidestepping costly mobile-robot teleoperation entirely. Across three real-world tasks it matches or beats Mobile-ALOHA-style baselines trained on teleoperated mobile-robot data, generalizes to new spatial configurations/scenes, and shows positive scaling as human-data hours grow — a data-source argument as much as a method.

RAMBO — RL-augmented model-based whole-body control (arXiv 2504.06662)

The cleanest statement of the "model-based + RL" loco-manipulation design point. A model-based module solves a QP for feedforward torques; an RL policy supplies feedback corrective terms for robustness to unmodeled dynamics. Validated on a Unitree Go2 across pushing a shopping cart, balancing a plate, and holding soft objects in both quadrupedal and bipedal modes. Code at github.com/catachiii/rambo.

M⁴Diffuser — multi-view diffusion + manipulability-aware QP (arXiv 2509.14980)

Addresses the perception/control split for MM: a multi-view diffusion policy consumes proprioception + complementary camera views to emit world-frame end-effector goals, executed by a Reduced and Manipulability-aware QP (ReM-QP) that drops slack variables for speed and biases away from singularities. Reports 7–56% higher success and 3–31% fewer collisions than baselines in sim and real.

CURA-PPO — uncertainty-aware non-prehensile MM under occlusion (arXiv 2602.01731)

When the pushed object blocks the onboard sensor, occluded regions cause collisions. CURA-PPO predicts collision possibility as a distribution, extracting both risk and uncertainty; the uncertainty term rewards active perception, so the robot simultaneously manipulates and gathers information to resolve occlusion — up to 3× higher success than baselines across object sizes and obstacle layouts.

UMI-on-Air — embodiment-aware guidance for agnostic policies (arXiv 2510.02614)

A transfer story: train an embodiment-agnostic visuomotor policy on unconstrained handheld-gripper (UMI) human demos, then deploy on a constrained aerial manipulator. The Embodiment-Aware Diffusion Policy (EADP) couples the high-level UMI policy to a low-level embodiment-specific controller at inference time, avoiding the OOD/poor-execution failure of naive transfer — improving success, efficiency, and robustness under disturbance on long-horizon, high-precision aerial tasks.

Complete paper list (34)

Code Title arXiv
ThI1I.105 MIMO: A Multimodal Imitation Learning Framework for Mobile Manipulation with Exoskeleton-VR Teleoperation —
ThI1I.313 CMAR-Search: Commonsense and Memory Augmented Reasoning for Object Search in Dynamic Interactive Environments —
ThI1I.361 Generating and Optimizing Topologically Distinct Guesses for Mobile Manipulator Path Planning with Path Constraints 2410.20635
ThI1I.408 From Composable Models to Correct-By-Construction Software for Contact-Rich Robotic Mobile-Manipulation Tasks —
ThI2I.19 RAMBO: RL-Augmented Model-Based Whole-Body Control for Loco-Manipulation 2504.06662
ThI2I.221 HoMeR: Learning In-The-Wild Mobile Manipulation Via Hybrid Imitation and Whole-Body Control 2506.01185
ThI2I.319 Task and Skill Planning: Hierarchical Robot Planning with Black-Box Skills 2504.17901
ThI2I.7 Task-Driven Co-Design of Mobile Manipulators 2412.16635
ThI2LB.12 Heterogeneous Skill Learning for Asynchronous Multi-Robot Relay Pushing in Complex Environments —
ThI2LB.14 Inverse Reachability Map Guided Motion Planning of Mobile Manipulator —
TuBT2.4 Nonlinear Predictive Control of the Continuum and Hybrid Dynamics of a Suspended Deformable Cable for Aerial Pick and Place 2602.17199
TuI1I.251 SEEC: Stable End-Effector Control with Model-Enhanced Residual Learning for Humanoid Loco-Manipulation 2509.21231
TuI1I.268 Uncertainty-Aware Adaptive Dynamics for Underwater Vehicle–Manipulator Robots 2603.06548
TuI1I.318 BINDER: Instantly Adaptive Mobile Manipulation with Open-Vocabulary Commands 2511.22364
TuI1I.34 Rapid Adaptation of Particle Dynamics for Generalized Deformable Object Mobile Manipulation 2603.18246
TuI1I.423 Virtual-Force Based Visual Servo for Multiple Peg-In-Hole Assembly with Tightly Coupled Multi-Manipulator 2407.10570
TuI1I.58 TopAY: Efficient Trajectory Planning for Differential Drive Mobile Manipulators Via Topological Paths Search and Arc Length-Yaw Parameterization 2507.02761
TuI2I.166 M4Diffuser: Multi-View Diffusion Policy with Manipulability-Aware Control for Robust Mobile Manipulation 2509.14980
TuI2I.2 Mobile Manipulation Instruction Generation from Multiple Images with Automatic Metric Enhancement 2501.17022
TuI2I.258 Searching in Space and Time: Unified Memory-Action Loops for Open-World Object Retrieval 2511.14004
TuI2I.419 Cooperative Grasping for Collective Object Transport in Constrained Environments 2509.03638
WeBT2.6 EMMA: Scaling Mobile Manipulation Via Egocentric Human Data 2509.04443
WeI1I.1 Robust Nonprehensile Object Transportation with Uncertain Inertial Parameters 2411.07079
WeI1I.104 SHOPPER: Practical Insights on Grasp Strategies for Mobile Manipulation in the Wild 2504.12512
WeI1I.175 VLION: Vision-Language Guided Interactive Object Navigation with Mobile Manipulation —
WeI1I.220 UMI-On-Air: Embodiment-Aware Guidance for Embodiment-Agnostic Visuomotor Policies 2510.02614
WeI1I.267 Multi-Quadruped Cooperative Object Transport: Learning Decentralized Pinch-Lift-Move 2509.14342
WeI1I.324 LeGO-MM: Learning Navigation for Goal-Oriented Mobile Manipulation Via Hierarchical Policy Distillation —
WeI1I.67 CAIMAN: Causal Action Influence Detection for Sample-Efficient Loco-Manipulation 2502.00835
WeI1I.76 Reactive Whole-Body Control of Mobile Manipulators for Dynamic Target Tracking Via Adaptive-Predictive Visual Servoing —
WeI2I.280 Uncertainty-Aware Non-Prehensile Manipulation with Mobile Manipulators under Object-Induced Occlusion (CURA-PPO) 2602.01731
WeI2I.293 SM²ITH: Safe Mobile Manipulation with Interactive Human Prediction Via Task-Hierarchical Bilevel Model Predictive Control 2511.17798
WeI2I.407 Serving Innovation: Seamless Service by Advancing Food Runners with Mobile Manipulation (MOMO) —
WeI2I.426 Whole-Body Inverse Dynamics MPC for Legged Loco-Manipulation 2511.19709

Dashes mark papers for which no public arXiv preprint was confirmed (several are IEEE RA-L / Xplore-only or too recent to be indexed). Numbers in the deep-dives and sub-trends are quoted from the corresponding arXiv/program abstracts; figures are not asserted for papers without a confirmed source.

Related

← Back to ICRA-2026-VLA-Manipulation-Survey