ICLR 2026 MoMaGen - Heungwoo/research GitHub Wiki
MoMaGen β Generating Demonstrations under Soft and Hard Constraints for Multi-Step Bimanual Mobile Manipulation
Venue: ICLR 2026 Category: Data Generation β Bimanual Mobile Manipulation Trend tag: Synthetic data / sim-to-real
flowchart LR
Demo[Single source demo<br/>1-3 min teleop, ~45% base motion] --> Anno[Annotate subtasks:<br/>o_target, o_held, t_pregrasp, t_end, retraction]
Anno --> Loop[For each subtask]
Loop --> Reach[Hard: reachability<br/>arm IK feasible at sampled base]
Loop --> VisM[Hard: visibility during manipulation<br/>head cam unoccluded]
Loop --> VisN[Soft: visibility during navigation<br/>cost biases head toward target]
Loop --> Retr[Soft: retraction<br/>compact pose for next nav]
Reach --> Sample[Random sample T_base, T_cam<br/>cuRobo IK + motion plan]
VisM --> Sample
VisN --> Sample
Sample --> Sim[OmniGibson rollout<br/>verify task success]
Sim --> Data[Synthetic demo set<br/>~1000 demos, D0/D1/D2]
Data --> Pol[Train WB-VIMA / Ο0 LoRA]
Pol --> RT[Real-robot fine-tune<br/>+40 demos]
Imitation learning for bimanual mobile manipulation requires controlling a mobile base, two 7-DoF arms, and a head camera simultaneously β a teleoperator overload that drives data collection cost up steeply. Prior X-Gen demonstration generators (MimicGen, SkillMimicGen, DexMimicGen, DemoGen, PhysicsGen) focus on fixed-base manipulation and miss two concerns specific to the mobile setting:
- Reachability: if the base pose is naively replayed from a source demo, randomized objects fall outside the arm workspace.
- Visibility: training a visuomotor policy requires the head camera to actually see task-relevant objects; replayed camera trajectories do not guarantee this once base/object randomisation kicks in.
Table 1 in the paper directly contrasts methods:
| Method | Bimanual | Mobile | Obstacles | Base rand. | Active perception |
|---|---|---|---|---|---|
| MimicGen | β | β | β | β | β |
| SkillMimicGen | β | β | β | β | β |
| DexMimicGen | β | β | β | β | β |
| DemoGen | β | β | β | β | β |
| PhysicsGen | β | β | β | β | β |
| MoMaGen | β | β | β | β | β |
Each task is an MDP. Given source demos D_src, the goal is generating new successful demos D_gen subject to a constraint set {G_i}:
arg min_{a_t} L(Β·) s.t. s_{t+1} = f(s_t, a_t) (dynamics) G_kin(s_t, a_t) β€ 0 (kinematic feasibility) G_coll(s_t, a_t) β₯ 0 (collision-free) G_vis(s_t, a_t, o_{i(t)}) β€ 0 (visibility) T^{E_k}_W = T^{o_i}_W (T^{o_i,src}_W)β»ΒΉ T^{E_k}_W (β contact subtask, β k) β t : s_t β D_success (task success)
Each demonstration is decomposed into N subtasks (free-space or contact-rich); end-effectorβtarget relative poses are preserved across regenerated demos for the contact-rich segments.
Hard constraints must hold:
- Reachability: sampled base must allow IK for all required end-effector poses in the subtask.
- Object visibility during manipulation: head camera must see the target object unoccluded (using torso articulation if needed).
Soft constraints are penalised in the cost:
- Object visibility during navigation: head camera biased toward the target during base motion.
- Retraction: torso/arms tucked into a compact configuration after each manipulation phase to keep the robot's footprint small for subsequent navigation.
For each subtask:
- Verify held object is in hand; abort if not (typically a failed grasp).
- Compute transformed end-effector pose T_eef using new target object pose.
- Check visibility of target with current T_cam.
- Solve IK for arm trajectory {q_t^arm} with current T_base, T_cam.
- If visibility fails or no IK exists: sample new T_base, new T_cam, solve IK for both arm and torso {q_t^torso}.
- Plan torso/base motion to sampled T_base with soft visibility cost.
- Plan arm motion to pregrasp T_eef.
- Replay contact-rich segment in task space.
- Attempt retraction.
Motion planning + IK runs on cuRobo (GPU-accelerated). The authors decompose configuration into torso + arm subspaces for efficient conditional sampling β a TAMP-style trick.
- OmniGibson simulator (BEHAVIOR-1K).
- Four tasks (each in three randomization levels D0 / D1 / D2):
- Pick Cup β navigate to table, lift cup.
- Tidy Table β move cup from countertop to sink (long-range mobile manipulation).
- Put Dishes Away β stack two plates on a shelf with two arms (uncoordinated bimanual).
- Clean Frying Pan β scrub pan with brush using both arms (contact-rich coordinated bimanual).
- Domain randomization:
- D0: Β±15 cm and Β±15Β° object pose perturbation, same furniture.
- D1: task objects anywhere on furniture, unrestricted orientation (~1.3 m Γ 0.8 m for Pick Cup).
- D2: D1 + obstacles on furniture and floor.
A single human teleoperated demo per task (1β3 min, ~45 % base motion).
| Rand. | Method | Pick Cup | Tidy Table | Put Dishes | Clean Pan |
|---|---|---|---|---|---|
| D0 | MoMaGen | 0.86 | 0.80 | 0.38 | 0.51 |
| SkillMimicGen | 1.00 | 0.69 | 0.38 | 0.40 | |
| DexMimicGen | 1.00 | 0.72 | 0.38 | 0.35 | |
| MoMaGen w/o soft vis. | 0.88 | 0.78 | 0.50 | 0.46 | |
| MoMaGen w/o hard vis. | 0.97 | 0.59 | 0.29 | 0.24 | |
| MoMaGen w/o vis. | 0.97 | 0.74 | 0.29 | 0.36 | |
| D1 | MoMaGen | 0.60 | 0.64 | 0.34 | 0.20 |
| MoMaGen w/o vis. | 0.66 | 0.48 | 0.23 | 0.13 | |
| D2 | MoMaGen | 0.47 | 0.22 | 0.07 | 0.16 |
| MoMaGen w/o vis. | 0.50 | 0.16 | 0.05 | 0.12 |
Baselines (SkillMimicGen, DexMimicGen) fail outright on D1/D2 because replayed base motion cannot reach randomised objects; MoMaGen still produces valid demos at D2 with success rates in the 7β47 % range.
Aggregate: ~63 % average data-generation success on D0.
| Rand. | Method | Pick Cup | Tidy Table | Put Dishes | Clean Pan |
|---|---|---|---|---|---|
| D0 | MoMaGen | 1.00 | 0.86 | 0.79 | 0.69 |
| SkillMimicGen | 1.00 | 0.40 | 0.71 | 0.65 | |
| DexMimicGen | 1.00 | 0.39 | 0.71 | 0.67 | |
| MoMaGen w/o soft vis. | 1.00 | 0.63 | 0.62 | 0.56 | |
| MoMaGen w/o hard vis. | 0.98 | 0.63 | 0.68 | 0.55 | |
| MoMaGen w/o vis. | 0.90 | 0.46 | 0.40 | 0.35 | |
| D1 | MoMaGen | 0.93 | 0.89 | 0.78 | 0.80 |
| MoMaGen w/o vis. | 0.71 | 0.46 | 0.40 | 0.43 | |
| D2 | MoMaGen | 0.94 | 0.79 | 0.75 | 0.81 |
| MoMaGen w/o vis. | 0.73 | 0.48 | 0.40 | 0.44 |
MoMaGen often doubles the visibility ratio relative to baselines/ablations, and stays above 75 % even in D1/D2.
- 1000 generated demos per (task, randomization).
- WB-VIMA = whole-body VIMA, trained from scratch per task; Ο0 LoRA-finetuned (rank 32) on top of pretrained Ο0.
- WB-VIMA on Pick Cup D0 matches baselines (small randomization range makes replay sufficient); on Tidy Table D0 WB-VIMA on MoMaGen data significantly outperforms baselines that overfit to long replayed base trajectories.
- Pick Cup D1: only MoMaGen-trained WB-VIMA achieves any success (0.25); baseline-trained policies fail completely.
- Ο0 fine-tuning on MoMaGen data matches WB-VIMA across all three settings, showing transferability across imitation methods.
- WB-VIMA on Pick Cup D0: full MoMaGen 0.75 vs ablations 0.45β0.65.
- WB-VIMA on Tidy Table D0: full MoMaGen 0.40 vs ablations peak at 0.05 β the gap is even larger.
Ο0 fine-tuned with 500 / 1000 / 2000 demos for 50k steps; performance scales smoothly across tasks and randomization levels, with biggest gains under D1.
40 real demos collected; pretraining on 1000 MoMaGen synthetic demos vs from-scratch:
- WB-VIMA: synthetic-pretrained then fine-tuned = 10 % real success vs 0 % real-only. Even at low success, the pretrained policy reaches the cup; baseline doesn't progress.
- Ο0: synthetic-pretrained + 40 real demos = 60 % vs 0 % for real-only Ο0 fine-tune.
Generated Pick Cup demos for a TIAGo robot using only a single source demo from a Galexea R1, by replaying dense end-effector trajectories in task space (largely embodiment-agnostic).
- w/o soft visibility: drops navigation visibility (Tidy Table D0: 0.86 β 0.63) and policy success.
- w/o hard visibility: drops data-generation success on multi-step tasks (Clean Pan D0: 0.51 β 0.24) because torso/cam configuration is no longer optimised for reachability.
- w/o all visibility: worst across the board; mid-cluttered scenes especially affected.
- Failure analysis (Figure 8). Across all three randomization levels, simulation instabilities (controller inaccuracies / stochastic simulator effects) account for ~35 % of failures on average; arm-level motion planning is the largest planner-related failure source at ~40 % on average, exceeding base-level planning at ~26 %. In D2, navigation-related failures (base sampling, base IK, base TrajOpt) rise sharply because of floor obstacles in an already tight navigation space.
- Generation cost. The paper reports that each successful demonstration takes 0.1β1.3 GPU-hours to generate, ranging from the cheapest task (Pick Cup) to the most expensive (Put Dishes Away), on a single NVIDIA TITAN RTX. cuRobo-based GPU motion generation is the dominant cost.
- Compute-cost breakdown (Figures 9β10). Simulation execution dominates total compute, greatly exceeding the corresponding planning durations β e.g. base motion planning averages ~18 s, whereas executing that motion in simulation takes ~100 s. Base sampling shows high variance (ring-shaped random sampling around the target) and becomes increasingly expensive from D0 β D2 as scene complexity rises and feasible base poses become scarcer.
The authors explicitly state:
- Privileged information. Generation assumes ground-truth object poses and geometry β easy in simulation, hard in the real world. They suggest combining with vision foundation models (SAM2) for relative-pose estimation in real.
- No whole-body manipulation. Current method alternates discrete navigation and manipulation phases; tasks like opening doors that need simultaneous base+arm coordination are out of scope but stated to be a natural extension.
- GPU-heavy. GPU-accelerated motion generators (cuRobo) are computationally expensive during data generation.
Implicit limitations seen in the numbers:
- Sim-to-real gap is still large. WB-VIMA only reaches 10 % real success even with 1000 synthetic + 40 real demos; only Ο0's strong pretrained backbone closes the gap to 60 %.
- D2 success rates collapse for multi-step tasks (Put Dishes Away: 0.07; Clean Pan: 0.16), suggesting the constraint sampler still struggles with cluttered floor space.
MoMaGen is the first general data-generation method to simultaneously handle the mobile base, two arms, and the head camera in a unified constraint-optimization framework. Its contribution is twofold:
- Theoretical unification. The paper recasts the entire X-Gen family (MimicGen, SkillMimicGen, DexMimicGen, DemoGen, PhysicsGen) under a single constrained-optimization formulation, where each prior method corresponds to a specific (insufficient) hard/soft constraint set. This is a clarifying contribution for the field.
- Empirical breakthrough. It is the only automated generator that produces successful demos under D1/D2 randomization for mobile-manipulation tasks; baselines collapse to 0 %.
Compared to neighbours in the 2026 wiki:
- VLBiMan tackles bimanual one-shot generalisation at test time via VLM grounding; MoMaGen tackles the same scarcity problem at training time via diverse synthetic data. The two are complementary.
- WholeBodyVLA focuses on whole-body humanoid VLA architecture; MoMaGen focuses on the data substrate that such policies need.
- The Ο0-LoRA result (60 % real success after 1000 synthetic + 40 real demos) is direct evidence that synthetic-data pretraining is now a load-bearing component of generalist policy fine-tuning recipes alongside Ο0.5, Ο0.6 / Ο0.7.
- Among CoRL-25 / NeurIPS-25 data-generation peers, MoMaGen is the first to enable active perception as a first-class constraint, anticipating the visibility-aware data-generation trend likely to continue at RSS / IROS 2026.
- arXiv: https://arxiv.org/abs/2510.18316
- OpenReview: https://openreview.net/forum?id=bGPDviEtZ1
- Project: https://momagen.github.io
Authors / affiliations: Chengshu Li, Mengdi Xu, Arpit Bahety, Hang Yin, Yunfan Jiang, Huang Huang, Josiah Wong, Sujay Garlanka, Cem Gokmen, Ruohan Zhang, Weiyu Liu, Jiajun Wu, Roberto MartΓn-MartΓn, Li Fei-Fei (Stanford University; UT Austin). ICLR 2026 Poster.
β Back to ICLR-2026