ICLR 2026 VLBiMan - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Authors: Huayi Zhou, Kui Jia (CUHK Shenzhen; DexForce) Category: Bimanual Manipulation — One-Shot Demonstration Trend tag: Data efficiency / VLM-grounded modular pipelines
flowchart LR
Demo[Single human demo<br/>kinesthetic teaching, 10Hz] --> Seg[Spatiotemporal segmentation<br/>velocity / gripper waypoints<br/>+ human-in-the-loop refinement]
Seg --> Atom[Atomic skill extraction<br/>invariant vs. adaptable<br/>via bind/o,r,t/ + geometry]
T[Task description] --> VLM[Florence-2 + SAM2<br/>2D semantic masks]
VLM --> Adapt[Geometric adaptation<br/>centroid Δx, principal-axis Δθ,<br/>z-extent Δh]
Atom --> Comp[Trajectory composition<br/>progressive IK + δ_base/δ_z<br/>collision compensation]
Adapt --> Comp
Comp --> Exec[Bimanual execution<br/>fixed-base Aubo-i5 + binocular Kingfisher R-6000<br/>or Rokae xMate humanoid]
Bimanual manipulation policies trained via imitation (RDT-1B, π0, ALOHA series) demand thousands of teleoperated demonstrations per task and fail to generalize to instance variations, novel placements, or different embodiments without retraining. Zero-shot pipelines such as ReKep and MOKA require fragile prompt engineering and produce unreliable trajectories. VLBiMan targets the middle ground: one demonstration plus VLM grounding for adaptation, with no policy retraining at deployment.
The pipeline has three stages, formalized as F: (T, D, S_new) → {A_t_new}.
1. Task-Aware Bimanual Decomposition.
-
Spatiotemporal segmentation: trajectories sampled at 10 FPS produce a sequence of
(O_t, A_t)withA_t ∈ R^14(6-DoF + gripper for each arm). Waypointsw_iare scripted from velocity discontinuities, acceleration spikes, or gripper-state transitions. IK feasibility (MoveIt) validates each candidate segment. A human-in-the-loop pass refines the segmentation. -
Atomic skill extraction: a segment
M_iis labeled invariant ifbind(o_k, r, t) = 1for the whole interval (object rigidly grasped) andgeometry(o_k) ≈ geometry(o_k^demo)within toleranceε_g; otherwise it requires adaptation. Output:D ⇒ {M_i^inv} ∪ {M_j^var}.
2. Vision-Language Anchored Adaptation.
-
VLM scene understanding: task prompts
p_kparsed fromTare passed to Florence-2 (detection) + SAM2 (segmentation) to obtain 2D masksM_k^2D. No 6-DoF pose estimator and no learned grasp net (e.g., AnyGrasp) is used. -
Geometric feasibility: (i) position shift
Δx = p_new − p_demofrom back-projected representative points (mask centroid or table-contact point — Fig. 3 of the paper); (ii) rotationΔθfrom the principal axis of second-order image moments (Algorithm 1 in appendix) — applied only when object orientation matters (pens, spoons, lying-down bottles); (iii) scaleΔh_k = max_z(P_k^3D) − min_z(P_k^3D)for category-level shape change.
3. Autonomous Trajectory Composition.
-
Progressive IK refinement: for grasping segments, iteratively solve IK on a spline interpolation of the target pose with interpolation density
n=6. Initial pose updated continuously to give closed-loop correction under disturbances. -
Dynamic collision compensation:
x̃_goal = x_goal + δ_base · u_∥ + δ_z · u_zreduces early contact risk; a one-time physical replay refinesδvalues, reusable across deployments of the same object.
The framework natively supports both synchronous and asynchronous bimanual control because invariant modules are stored per-arm and reused.
Hardware. Two opposite-side fixed-base Aubo-i5 6-DoF arms (880 mm reach) with DH-Robotics parallel grippers (80 mm max width); binocular Kingfisher R-6000 stereo camera at 960×540 (third-person, 100 cm above the long edge). Cross-embodiment platform: two Rokae xMate CR7 humanoid arms with Jodell RG75-300 grippers.
Six basic bimanual tasks, 20 trials each, success rates under "new placements + same objects" / "new placements + novel instances", with and without external interference (Table 1 of the paper).
| Method | No-interf., same objs | No-interf., novel insts | With-interf., same objs | With-interf., novel insts |
|---|---|---|---|---|
| Mechanisms (Mao et al. 2023) | 28.3% | 12.5% | 13.3% | 3.3% |
| MAGIC (Liu et al. 2025b) | 45.8% | 28.3% | 24.2% | 11.7% |
| Robot-ABC (Ju et al. 2024) | 36.7% | 24.2% | 18.3% | 6.7% |
| ReKep (Huang et al. 2024b) | 43.3% | 30.8% | 23.3% | 14.2% |
| ReKep+ (oracle grasp) | 65.0% | 43.3% | 37.5% | 25.0% |
| VLBiMan | 85.0% | 78.3% | 70.0% | 59.2% |
Four long-horizon multi-stage tasks (Table 2): reorient+unscrew, unscrew+pouring, tool-use:spoon, tool-use:funnel. Composed tasks have no additional one-shot demo — they are inferred from constituent base skills.
| Method | No-interf., same objs | No-interf., novel insts | With-interf., same objs | With-interf., novel insts |
|---|---|---|---|---|
| ReKep+ | 35.0% | 20.0% | 25.0% | 12.5% |
| VLBiMan | 52.5% | 41.3% | 38.8% | 25.0% |
Cross-embodiment transfer: VLBiMan executed on the humanoid Rokae platform (inserting, pouring, unscrew, reorient) without retraining, demonstrated qualitatively (Fig. 6 of the paper).
Conducted under the hardest setting (new placements + novel instances + interference), six basic tasks averaged (Table 3 of paper):
| VLM | Initial grasp | IK refine | Collision avoid | Avg SR |
|---|---|---|---|---|
| SAM+DINOv2 | ours | yes | yes | 35.8% |
| ours (Florence-2+SAM2) | AnyGrasp | yes | yes | 31.7% |
| ours | ours | no | yes | 29.2% |
| ours | ours | yes | no | 34.2% |
| ours | ours | yes | yes | 59.2% |
Notable findings:
- IK refinement is the single most important component (29.2% → 59.2%).
- Replacing the demo-anchored grasp with AnyGrasp proposals drops 28 pts, despite AnyGrasp being a stronger general-purpose grasp predictor — task-grounding from the demo dominates.
- SAM+DINOv2 vs. Florence-2+SAM2 costs ~24 pts: VLM choice is non-trivial.
Error breakdown (Fig. 7): Initial grasp execution (45%) dominates failures, followed by dual-arm coordination (21%), initial grasp computing (12%), VLM perception (6%), trajectory optimization (6%), other (10%).
Appendix additionally reports robustness to uneven illumination (70.0% → 67.5%, marginal degradation) and cluttered scenes, as well as ablations on pre-grasp interpolation density n.
- Restricted to rigid objects — does not handle deformables (cloth, rope).
- No runtime anomaly detection or recovery — slippage or occlusion mid-execution causes failures.
- Hardware-bounded: fixed-base platform limits workspace and the grippers lack force/tactile sensing; future work mentions mobile base and tactile end-effectors.
VLBiMan is a clean counter-example to the "scale teleop data" hypothesis for bimanual manipulation: with a single kinesthetic demonstration and off-the-shelf VLMs (Florence-2 + SAM2) used purely as perception modules (not as planners or controllers), it beats the strongest VLM-pipeline baseline (ReKep+, which gets an oracle grasp) by 35+ points absolute on novel-instance settings. The architectural commitment is that what to achieve matters more than how to execute — invariant primitives (lift, carry, align, pour) are preserved verbatim, and only the contact-affecting components are re-grounded.
Compared to:
- MoMaGen and WholeBodyVLA — VLBiMan does not learn a policy; it composes.
- ReKep — VLBiMan is demonstration-conditioned, avoiding the brittle prompt-engineering of pure zero-shot VLM pipelines.
- TwinVLA — both attack bimanual data scarcity, but TwinVLA composes two single-arm policies while VLBiMan composes primitives extracted from one demo.
Aligned with the broader 2026 trend of using VLM priors as the adaptation layer rather than as the policy itself.
- OpenReview: https://openreview.net/forum?id=he86smZzRk
- Project page: https://hnuzhy.github.io/projects/VLBiMan
- TwinVLA — modular composition of two single-arm VLAs
- MoMaGen
- WholeBodyVLA
- Sim2Real-VLA — sibling CUHK Shenzhen / DexForce work
← Back to ICLR-2026