RSS 2026 One Shot Real World Demonstration Synthesis for - Heungwoo/research GitHub Wiki
One-Shot Real-World Demonstration Synthesis for Scalable Bimanual Manipulation
Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Manipulation 1 · paper #1 Authors: Huayi Zhou, Kui Jia arXiv: 2512.09297 · program page
Summary compiled from the arXiv paper (v3); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Figure 1 uses the dual-arm pouring task to show the "From One to Many" idea: from a single kinesthetic demonstration, the left panel synthesizes fans of pre-grasping/lifting trajectories of the left arm for new bottle placements, the right panel does the same for the cup with the right arm, and the center column shows the demonstration decomposed into variant (adaptable) and invariable coordination blocks (grasp bottle → lift → pour → put down; grasp cup → move → stabilize → put down).
Problem
Scaling bimanual imitation learning is bottlenecked by data: teleoperation (e.g., ALOHA-style rigs) yields physically grounded demonstrations but is prohibitively labor-intensive, while simulation-based synthesis (MimicGen, RoboGen, RoboCasa) scales cheaply but inherits sim-to-real gaps in contact dynamics and rendering. The paper asks how to synthesize thousands of contact-rich, physically feasible dual-arm demonstrations directly in the real world from a single example.
Method
BiDemoSyn is a three-stage, simulator-free pipeline. (1) Deconstruction of one-shot teaching: the single demonstration is segmented into bimanual execution blocks and refined into Atomic Execution Primitives, then categorized into invariant blocks (task-semantic primitives like screwing, pressing) and variable, object-dependent blocks. (2) Vision-based initial-frame alignment: open-vocabulary detection/segmentation (YOLO-World/YOLOE or Florence2+SAM2), image-moment + PCA instance-level 6-DoF pose estimation, and a rigid-transform adaptation of the demonstrated grasp pose to the novel scene. (3) Trajectory modulation and optimization: IK reachability and motion-planner collision checks per block, plus instance-level endpoint offsets scaled by object bounding-box differences (λ ≈ 0.8–1.0). Policies (bimanual variants of DP, DP3, EquiBot) are trained on the synthesized data with a coordination-weighted diffusion loss and a simplified 6-DoF end-effector action space that enables cross-embodiment transfer.
Results
On six real dual-arm tasks (plugpen, inserting, unscrew, pouring, pressing, reorient), synthesis takes about 5 s per demonstration versus 42 s for YOTO auto-replay and 91 s for teleoperation, enabling thousands of demos per task. With EquiBot on point clouds, BiDemoSyn data reaches 86.7% average in-distribution success and 66.7% out-of-distribution, versus 71.1%/47.8% for YOTO and 68.9%/40.6% for DemoGen data; DP3 gets 81.1%/54.4% and RGB DP 67.8%/42.2%. Training-free baselines (ReKep, ReKep+, ODIL, MAGIC) top out at 55.6% ID. The paper also reports few-shot extension and zero-shot cross-embodiment transfer to an auxiliary humanoid dual-arm platform.
Significance
A real-world alternative to MimicGen-style simulation augmentation: real2real data generation keeps physical fidelity while approaching scripted-synthesis throughput. Relevant to the demonstration-scaling thread in Review-Dexterous-Manipulation and the data-efficiency discussion in Review-LBM-Cotraining.
← Back to RSS 2026 survey · RSS-2026-Papers · Home