ICLR 2026 MoMaGen - Heungwoo/research GitHub Wiki

MoMaGen β€” Generating Demonstrations under Soft and Hard Constraints for Multi-Step Bimanual Mobile Manipulation

Venue: ICLR 2026 Category: Data Generation β€” Bimanual Mobile Manipulation Trend tag: Synthetic data / sim-to-real

Approach diagram

flowchart LR
  Demo[Single source demo<br/>1-3 min teleop, ~45% base motion] --> Anno[Annotate subtasks:<br/>o_target, o_held, t_pregrasp, t_end, retraction]
  Anno --> Loop[For each subtask]
  Loop --> Reach[Hard: reachability<br/>arm IK feasible at sampled base]
  Loop --> VisM[Hard: visibility during manipulation<br/>head cam unoccluded]
  Loop --> VisN[Soft: visibility during navigation<br/>cost biases head toward target]
  Loop --> Retr[Soft: retraction<br/>compact pose for next nav]
  Reach --> Sample[Random sample T_base, T_cam<br/>cuRobo IK + motion plan]
  VisM --> Sample
  VisN --> Sample
  Sample --> Sim[OmniGibson rollout<br/>verify task success]
  Sim --> Data[Synthetic demo set<br/>~1000 demos, D0/D1/D2]
  Data --> Pol[Train WB-VIMA / Ο€0 LoRA]
  Pol --> RT[Real-robot fine-tune<br/>+40 demos]
Loading

Problem

Imitation learning for bimanual mobile manipulation requires controlling a mobile base, two 7-DoF arms, and a head camera simultaneously β€” a teleoperator overload that drives data collection cost up steeply. Prior X-Gen demonstration generators (MimicGen, SkillMimicGen, DexMimicGen, DemoGen, PhysicsGen) focus on fixed-base manipulation and miss two concerns specific to the mobile setting:

  1. Reachability: if the base pose is naively replayed from a source demo, randomized objects fall outside the arm workspace.
  2. Visibility: training a visuomotor policy requires the head camera to actually see task-relevant objects; replayed camera trajectories do not guarantee this once base/object randomisation kicks in.

Table 1 in the paper directly contrasts methods:

Method Bimanual Mobile Obstacles Base rand. Active perception
MimicGen βœ— βœ“ βœ— βœ— βœ—
SkillMimicGen βœ— βœ— βœ“ βœ— βœ—
DexMimicGen βœ“ βœ— βœ— βœ— βœ—
DemoGen βœ“ βœ— βœ“ βœ— βœ—
PhysicsGen βœ“ βœ— βœ— βœ— βœ—
MoMaGen βœ“ βœ“ βœ“ βœ“ βœ“

Detailed Method

Constrained-optimization formulation

Each task is an MDP. Given source demos D_src, the goal is generating new successful demos D_gen subject to a constraint set {G_i}:

arg min_{a_t} L(Β·) s.t. s_{t+1} = f(s_t, a_t) (dynamics) G_kin(s_t, a_t) ≀ 0 (kinematic feasibility) G_coll(s_t, a_t) β‰₯ 0 (collision-free) G_vis(s_t, a_t, o_{i(t)}) ≀ 0 (visibility) T^{E_k}_W = T^{o_i}_W (T^{o_i,src}_W)⁻¹ T^{E_k}_W (βˆ€ contact subtask, βˆ€ k) βˆƒ t : s_t ∈ D_success (task success)

Each demonstration is decomposed into N subtasks (free-space or contact-rich); end-effector–target relative poses are preserved across regenerated demos for the contact-rich segments.

Hard / soft constraint instantiation

Hard constraints must hold:

  • Reachability: sampled base must allow IK for all required end-effector poses in the subtask.
  • Object visibility during manipulation: head camera must see the target object unoccluded (using torso articulation if needed).

Soft constraints are penalised in the cost:

  • Object visibility during navigation: head camera biased toward the target during base motion.
  • Retraction: torso/arms tucked into a compact configuration after each manipulation phase to keep the robot's footprint small for subsequent navigation.

Algorithm 1 (data generation)

For each subtask:

  1. Verify held object is in hand; abort if not (typically a failed grasp).
  2. Compute transformed end-effector pose T_eef using new target object pose.
  3. Check visibility of target with current T_cam.
  4. Solve IK for arm trajectory {q_t^arm} with current T_base, T_cam.
  5. If visibility fails or no IK exists: sample new T_base, new T_cam, solve IK for both arm and torso {q_t^torso}.
  6. Plan torso/base motion to sampled T_base with soft visibility cost.
  7. Plan arm motion to pregrasp T_eef.
  8. Replay contact-rich segment in task space.
  9. Attempt retraction.

Motion planning + IK runs on cuRobo (GPU-accelerated). The authors decompose configuration into torso + arm subspaces for efficient conditional sampling β€” a TAMP-style trick.

Tasks and randomization

  • OmniGibson simulator (BEHAVIOR-1K).
  • Four tasks (each in three randomization levels D0 / D1 / D2):
    • Pick Cup β€” navigate to table, lift cup.
    • Tidy Table β€” move cup from countertop to sink (long-range mobile manipulation).
    • Put Dishes Away β€” stack two plates on a shelf with two arms (uncoordinated bimanual).
    • Clean Frying Pan β€” scrub pan with brush using both arms (contact-rich coordinated bimanual).
  • Domain randomization:
    • D0: Β±15 cm and Β±15Β° object pose perturbation, same furniture.
    • D1: task objects anywhere on furniture, unrestricted orientation (~1.3 m Γ— 0.8 m for Pick Cup).
    • D2: D1 + obstacles on furniture and floor.

A single human teleoperated demo per task (1–3 min, ~45 % base motion).

Comprehensive Results

Data-generation success rate (Table 2)

Rand. Method Pick Cup Tidy Table Put Dishes Clean Pan
D0 MoMaGen 0.86 0.80 0.38 0.51
SkillMimicGen 1.00 0.69 0.38 0.40
DexMimicGen 1.00 0.72 0.38 0.35
MoMaGen w/o soft vis. 0.88 0.78 0.50 0.46
MoMaGen w/o hard vis. 0.97 0.59 0.29 0.24
MoMaGen w/o vis. 0.97 0.74 0.29 0.36
D1 MoMaGen 0.60 0.64 0.34 0.20
MoMaGen w/o vis. 0.66 0.48 0.23 0.13
D2 MoMaGen 0.47 0.22 0.07 0.16
MoMaGen w/o vis. 0.50 0.16 0.05 0.12

Baselines (SkillMimicGen, DexMimicGen) fail outright on D1/D2 because replayed base motion cannot reach randomised objects; MoMaGen still produces valid demos at D2 with success rates in the 7–47 % range.

Aggregate: ~63 % average data-generation success on D0.

Object visibility during navigation (Table 3)

Rand. Method Pick Cup Tidy Table Put Dishes Clean Pan
D0 MoMaGen 1.00 0.86 0.79 0.69
SkillMimicGen 1.00 0.40 0.71 0.65
DexMimicGen 1.00 0.39 0.71 0.67
MoMaGen w/o soft vis. 1.00 0.63 0.62 0.56
MoMaGen w/o hard vis. 0.98 0.63 0.68 0.55
MoMaGen w/o vis. 0.90 0.46 0.40 0.35
D1 MoMaGen 0.93 0.89 0.78 0.80
MoMaGen w/o vis. 0.71 0.46 0.40 0.43
D2 MoMaGen 0.94 0.79 0.75 0.81
MoMaGen w/o vis. 0.73 0.48 0.40 0.44

MoMaGen often doubles the visibility ratio relative to baselines/ablations, and stays above 75 % even in D1/D2.

Policy learning (WB-VIMA + Ο€0 LoRA)

  • 1000 generated demos per (task, randomization).
  • WB-VIMA = whole-body VIMA, trained from scratch per task; Ο€0 LoRA-finetuned (rank 32) on top of pretrained Ο€0.
  • WB-VIMA on Pick Cup D0 matches baselines (small randomization range makes replay sufficient); on Tidy Table D0 WB-VIMA on MoMaGen data significantly outperforms baselines that overfit to long replayed base trajectories.
  • Pick Cup D1: only MoMaGen-trained WB-VIMA achieves any success (0.25); baseline-trained policies fail completely.
  • Ο€0 fine-tuning on MoMaGen data matches WB-VIMA across all three settings, showing transferability across imitation methods.

Visibility ablation on policy (Figure 6d)

  • WB-VIMA on Pick Cup D0: full MoMaGen 0.75 vs ablations 0.45–0.65.
  • WB-VIMA on Tidy Table D0: full MoMaGen 0.40 vs ablations peak at 0.05 β€” the gap is even larger.

Data scaling (Figure 7)

Ο€0 fine-tuned with 500 / 1000 / 2000 demos for 50k steps; performance scales smoothly across tasks and randomization levels, with biggest gains under D1.

Real-world deployment (Pick Cup, Franka mobile platform)

40 real demos collected; pretraining on 1000 MoMaGen synthetic demos vs from-scratch:

  • WB-VIMA: synthetic-pretrained then fine-tuned = 10 % real success vs 0 % real-only. Even at low success, the pretrained policy reaches the cup; baseline doesn't progress.
  • Ο€0: synthetic-pretrained + 40 real demos = 60 % vs 0 % for real-only Ο€0 fine-tune.

Cross-embodiment data generation

Generated Pick Cup demos for a TIAGo robot using only a single source demo from a Galexea R1, by replaying dense end-effector trajectories in task space (largely embodiment-agnostic).

Ablation Studies

  • w/o soft visibility: drops navigation visibility (Tidy Table D0: 0.86 β†’ 0.63) and policy success.
  • w/o hard visibility: drops data-generation success on multi-step tasks (Clean Pan D0: 0.51 β†’ 0.24) because torso/cam configuration is no longer optimised for reachability.
  • w/o all visibility: worst across the board; mid-cluttered scenes especially affected.
  • Failure analysis (Figure 8). Across all three randomization levels, simulation instabilities (controller inaccuracies / stochastic simulator effects) account for ~35 % of failures on average; arm-level motion planning is the largest planner-related failure source at ~40 % on average, exceeding base-level planning at ~26 %. In D2, navigation-related failures (base sampling, base IK, base TrajOpt) rise sharply because of floor obstacles in an already tight navigation space.
  • Generation cost. The paper reports that each successful demonstration takes 0.1–1.3 GPU-hours to generate, ranging from the cheapest task (Pick Cup) to the most expensive (Put Dishes Away), on a single NVIDIA TITAN RTX. cuRobo-based GPU motion generation is the dominant cost.
  • Compute-cost breakdown (Figures 9–10). Simulation execution dominates total compute, greatly exceeding the corresponding planning durations β€” e.g. base motion planning averages ~18 s, whereas executing that motion in simulation takes ~100 s. Base sampling shows high variance (ring-shaped random sampling around the target) and becomes increasingly expensive from D0 β†’ D2 as scene complexity rises and feasible base poses become scarcer.

Limitations

The authors explicitly state:

  1. Privileged information. Generation assumes ground-truth object poses and geometry β€” easy in simulation, hard in the real world. They suggest combining with vision foundation models (SAM2) for relative-pose estimation in real.
  2. No whole-body manipulation. Current method alternates discrete navigation and manipulation phases; tasks like opening doors that need simultaneous base+arm coordination are out of scope but stated to be a natural extension.
  3. GPU-heavy. GPU-accelerated motion generators (cuRobo) are computationally expensive during data generation.

Implicit limitations seen in the numbers:

  • Sim-to-real gap is still large. WB-VIMA only reaches 10 % real success even with 1000 synthetic + 40 real demos; only Ο€0's strong pretrained backbone closes the gap to 60 %.
  • D2 success rates collapse for multi-step tasks (Put Dishes Away: 0.07; Clean Pan: 0.16), suggesting the constraint sampler still struggles with cluttered floor space.

Significance & Positioning

MoMaGen is the first general data-generation method to simultaneously handle the mobile base, two arms, and the head camera in a unified constraint-optimization framework. Its contribution is twofold:

  1. Theoretical unification. The paper recasts the entire X-Gen family (MimicGen, SkillMimicGen, DexMimicGen, DemoGen, PhysicsGen) under a single constrained-optimization formulation, where each prior method corresponds to a specific (insufficient) hard/soft constraint set. This is a clarifying contribution for the field.
  2. Empirical breakthrough. It is the only automated generator that produces successful demos under D1/D2 randomization for mobile-manipulation tasks; baselines collapse to 0 %.

Compared to neighbours in the 2026 wiki:

  • VLBiMan tackles bimanual one-shot generalisation at test time via VLM grounding; MoMaGen tackles the same scarcity problem at training time via diverse synthetic data. The two are complementary.
  • WholeBodyVLA focuses on whole-body humanoid VLA architecture; MoMaGen focuses on the data substrate that such policies need.
  • The Ο€0-LoRA result (60 % real success after 1000 synthetic + 40 real demos) is direct evidence that synthetic-data pretraining is now a load-bearing component of generalist policy fine-tuning recipes alongside Ο€0.5, Ο€0.6 / Ο€0.7.
  • Among CoRL-25 / NeurIPS-25 data-generation peers, MoMaGen is the first to enable active perception as a first-class constraint, anticipating the visibility-aware data-generation trend likely to continue at RSS / IROS 2026.

Links

Authors / affiliations: Chengshu Li, Mengdi Xu, Arpit Bahety, Hang Yin, Yunfan Jiang, Huang Huang, Josiah Wong, Sujay Garlanka, Cem Gokmen, Ruohan Zhang, Weiyu Liu, Jiajun Wu, Roberto MartΓ­n-MartΓ­n, Li Fei-Fei (Stanford University; UT Austin). ICLR 2026 Poster.

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️