ICLR 2026 VLBiMan - Heungwoo/research GitHub Wiki

VLBiMan — Vision-Language Anchored One-Shot Bimanual Manipulation

Venue: ICLR 2026 Authors: Huayi Zhou, Kui Jia (CUHK Shenzhen; DexForce) Category: Bimanual Manipulation — One-Shot Demonstration Trend tag: Data efficiency / VLM-grounded modular pipelines

Approach diagram

flowchart LR
  Demo[Single human demo<br/>kinesthetic teaching, 10Hz] --> Seg[Spatiotemporal segmentation<br/>velocity / gripper waypoints<br/>+ human-in-the-loop refinement]
  Seg --> Atom[Atomic skill extraction<br/>invariant vs. adaptable<br/>via bind/o,r,t/ + geometry]
  T[Task description] --> VLM[Florence-2 + SAM2<br/>2D semantic masks]
  VLM --> Adapt[Geometric adaptation<br/>centroid Δx, principal-axis Δθ,<br/>z-extent Δh]
  Atom --> Comp[Trajectory composition<br/>progressive IK + δ_base/δ_z<br/>collision compensation]
  Adapt --> Comp
  Comp --> Exec[Bimanual execution<br/>fixed-base Aubo-i5 + binocular Kingfisher R-6000<br/>or Rokae xMate humanoid]
Loading

Problem

Bimanual manipulation policies trained via imitation (RDT-1B, π0, ALOHA series) demand thousands of teleoperated demonstrations per task and fail to generalize to instance variations, novel placements, or different embodiments without retraining. Zero-shot pipelines such as ReKep and MOKA require fragile prompt engineering and produce unreliable trajectories. VLBiMan targets the middle ground: one demonstration plus VLM grounding for adaptation, with no policy retraining at deployment.

Detailed Method

The pipeline has three stages, formalized as F: (T, D, S_new) → {A_t_new}.

1. Task-Aware Bimanual Decomposition.

  • Spatiotemporal segmentation: trajectories sampled at 10 FPS produce a sequence of (O_t, A_t) with A_t ∈ R^14 (6-DoF + gripper for each arm). Waypoints w_i are scripted from velocity discontinuities, acceleration spikes, or gripper-state transitions. IK feasibility (MoveIt) validates each candidate segment. A human-in-the-loop pass refines the segmentation.
  • Atomic skill extraction: a segment M_i is labeled invariant if bind(o_k, r, t) = 1 for the whole interval (object rigidly grasped) and geometry(o_k) ≈ geometry(o_k^demo) within tolerance ε_g; otherwise it requires adaptation. Output: D ⇒ {M_i^inv} ∪ {M_j^var}.

2. Vision-Language Anchored Adaptation.

  • VLM scene understanding: task prompts p_k parsed from T are passed to Florence-2 (detection) + SAM2 (segmentation) to obtain 2D masks M_k^2D. No 6-DoF pose estimator and no learned grasp net (e.g., AnyGrasp) is used.
  • Geometric feasibility: (i) position shift Δx = p_new − p_demo from back-projected representative points (mask centroid or table-contact point — Fig. 3 of the paper); (ii) rotation Δθ from the principal axis of second-order image moments (Algorithm 1 in appendix) — applied only when object orientation matters (pens, spoons, lying-down bottles); (iii) scale Δh_k = max_z(P_k^3D) − min_z(P_k^3D) for category-level shape change.

3. Autonomous Trajectory Composition.

  • Progressive IK refinement: for grasping segments, iteratively solve IK on a spline interpolation of the target pose with interpolation density n=6. Initial pose updated continuously to give closed-loop correction under disturbances.
  • Dynamic collision compensation: x̃_goal = x_goal + δ_base · u_∥ + δ_z · u_z reduces early contact risk; a one-time physical replay refines δ values, reusable across deployments of the same object.

The framework natively supports both synchronous and asynchronous bimanual control because invariant modules are stored per-arm and reused.

Hardware. Two opposite-side fixed-base Aubo-i5 6-DoF arms (880 mm reach) with DH-Robotics parallel grippers (80 mm max width); binocular Kingfisher R-6000 stereo camera at 960×540 (third-person, 100 cm above the long edge). Cross-embodiment platform: two Rokae xMate CR7 humanoid arms with Jodell RG75-300 grippers.

Comprehensive Results

Six basic bimanual tasks, 20 trials each, success rates under "new placements + same objects" / "new placements + novel instances", with and without external interference (Table 1 of the paper).

Method No-interf., same objs No-interf., novel insts With-interf., same objs With-interf., novel insts
Mechanisms (Mao et al. 2023) 28.3% 12.5% 13.3% 3.3%
MAGIC (Liu et al. 2025b) 45.8% 28.3% 24.2% 11.7%
Robot-ABC (Ju et al. 2024) 36.7% 24.2% 18.3% 6.7%
ReKep (Huang et al. 2024b) 43.3% 30.8% 23.3% 14.2%
ReKep+ (oracle grasp) 65.0% 43.3% 37.5% 25.0%
VLBiMan 85.0% 78.3% 70.0% 59.2%

Four long-horizon multi-stage tasks (Table 2): reorient+unscrew, unscrew+pouring, tool-use:spoon, tool-use:funnel. Composed tasks have no additional one-shot demo — they are inferred from constituent base skills.

Method No-interf., same objs No-interf., novel insts With-interf., same objs With-interf., novel insts
ReKep+ 35.0% 20.0% 25.0% 12.5%
VLBiMan 52.5% 41.3% 38.8% 25.0%

Cross-embodiment transfer: VLBiMan executed on the humanoid Rokae platform (inserting, pouring, unscrew, reorient) without retraining, demonstrated qualitatively (Fig. 6 of the paper).

Ablation Studies

Conducted under the hardest setting (new placements + novel instances + interference), six basic tasks averaged (Table 3 of paper):

VLM Initial grasp IK refine Collision avoid Avg SR
SAM+DINOv2 ours yes yes 35.8%
ours (Florence-2+SAM2) AnyGrasp yes yes 31.7%
ours ours no yes 29.2%
ours ours yes no 34.2%
ours ours yes yes 59.2%

Notable findings:

  • IK refinement is the single most important component (29.2% → 59.2%).
  • Replacing the demo-anchored grasp with AnyGrasp proposals drops 28 pts, despite AnyGrasp being a stronger general-purpose grasp predictor — task-grounding from the demo dominates.
  • SAM+DINOv2 vs. Florence-2+SAM2 costs ~24 pts: VLM choice is non-trivial.

Error breakdown (Fig. 7): Initial grasp execution (45%) dominates failures, followed by dual-arm coordination (21%), initial grasp computing (12%), VLM perception (6%), trajectory optimization (6%), other (10%).

Appendix additionally reports robustness to uneven illumination (70.0% → 67.5%, marginal degradation) and cluttered scenes, as well as ablations on pre-grasp interpolation density n.

Limitations stated by authors

  1. Restricted to rigid objects — does not handle deformables (cloth, rope).
  2. No runtime anomaly detection or recovery — slippage or occlusion mid-execution causes failures.
  3. Hardware-bounded: fixed-base platform limits workspace and the grippers lack force/tactile sensing; future work mentions mobile base and tactile end-effectors.

Significance & Positioning

VLBiMan is a clean counter-example to the "scale teleop data" hypothesis for bimanual manipulation: with a single kinesthetic demonstration and off-the-shelf VLMs (Florence-2 + SAM2) used purely as perception modules (not as planners or controllers), it beats the strongest VLM-pipeline baseline (ReKep+, which gets an oracle grasp) by 35+ points absolute on novel-instance settings. The architectural commitment is that what to achieve matters more than how to execute — invariant primitives (lift, carry, align, pour) are preserved verbatim, and only the contact-affecting components are re-grounded.

Compared to:

  • MoMaGen and WholeBodyVLA — VLBiMan does not learn a policy; it composes.
  • ReKep — VLBiMan is demonstration-conditioned, avoiding the brittle prompt-engineering of pure zero-shot VLM pipelines.
  • TwinVLA — both attack bimanual data scarcity, but TwinVLA composes two single-arm policies while VLBiMan composes primitives extracted from one demo.

Aligned with the broader 2026 trend of using VLM priors as the adaptation layer rather than as the policy itself.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️