ICRA 2026 Topic Bimanual - Heungwoo/research GitHub Wiki

ICRA 2026 — Bimanual & Dual-Arm Manipulation (Topic Analysis)

Venue: IEEE ICRA 2026 · Vienna, Austria · June 1–5, 2026 Compiled against the official PaperCept program. This page covers the 50 papers that the program tags with Bimanual Manipulation or Dual Arm Manipulation as a primary/secondary keyword. ← Back to ICRA-2026-VLA-Manipulation-Survey

Overview

Bimanual and dual-arm manipulation is one of ICRA 2026's fastest-growing manipulation sub-fields: the program's topic-keyword counts put Bimanual at 36 and Dual Arm Manipulation as a distinct large cluster, and once both tags are pooled the result is the 50-paper group analyzed here. The growth tracks the broader move toward two-arm hardware platforms — ALOHA/Mobile-ALOHA, Unitree G1, and assorted dual-arm industrial cells — and the recognition that most genuinely useful household and industrial tasks (folding, assembly, handover, deformable-object handling, surgery) are intrinsically two-handed.

The unifying technical problem is coordination: two arms share a workspace, a kinematic budget, and often a single object, so naive single-arm policies, planners, and grasp generators do not transfer. The 50 papers attack this from every direction at once — mechanism design, optimal control, TAMP, imitation/diffusion policies, dexterous teleoperation, and bimanual VLA. A notable feature of the ICRA cohort (versus ML-venue work) is its systems orientation: many entries are about hardware (wrists, grippers, surgical tools), force/tactile control, and real-robot data collection rather than pure architecture. Where the ICLR/CVPR bimanual wave is dominated by VLA backbones, ICRA's center of gravity sits on planning, control, and deformable-object physics.

Sub-trends

1. Coordination frameworks, role assignment & dynamic division of labor

The defining bimanual challenge — who does what, when — gets explicit treatment. PA-BiCoop (ThI1I.268) frames coordination as a primary–auxiliary cooperative structure with dynamic division of labor, addressing methods that either lack inter-arm interaction or fix the labor split. DAG-Plan (ThI2I.248, arXiv 2406.09953) uses LLMs to decompose long-horizon tasks into a directed acyclic graph with explicit temporal dependencies, then dynamically assigns sub-tasks to arms based on real-time observations — modeling parallelism that linear LLM plans cannot. VLM-SFD (TuI2I.365, arXiv 2506.13428) similarly uses a pre-trained VLM to adaptively assign the optimal object-centric motion flow to each arm over time. Consensus Driven Dynamical Systems Control for Dual-Arm Handover (TuI2I.235) handles the tight spatial-and-temporal coordination of object transfer between two arms. Impact-Aware Dual-Arm Manipulation (WeI1I.395) targets logistics depalletizing where the two arms must grab/place under impact dynamics.

2. Bimanual / dual-arm task & motion planning

A robust planning cluster tackles the combinatorial blow-up of two-arm action spaces. ScheduleStream (WeI1I.203, arXiv 2511.04758, NVIDIA) is the first general framework for planning and scheduling with samplers, using GPU acceleration and hybrid durative actions so arms can move in parallel rather than one-at-a-time. High-Performance Dual-Arm TAMP for Tabletop Rearrangement (WeI2I.224) presents the Synchronous Dual-Arm Rearrangement Planner (SDAR) for close-proximity rearrangement. TOCALib (ThI1I.116, arXiv 2504.07708) builds a library of optimal two-arm trajectories with DCOL-based symbolic collision expressions inside the FROST framework (demonstrated on Mobile ALOHA). An Efficient Learning-Based Task Planning Approach (TuI2LB.2) introduces a bio-inspired action context-free grammar to curb the TAMP combinatorial explosion, while Connectivity-Aware Representations (ThI1I.128) uses multi-scale contrastive learning to connect disconnected constraint regions for bimanual motion planning.

3. Bimanual diffusion / flow policies & constraint-aware learning

Diffusion and flow models adapted to two-arm coordination form a dense methodological cluster. Planning-Guided Diffusion Policy Learning (ThI1I.71) targets contact-rich bimanual object reorientation. RoTri-Diff (ThI2I.181) is a spatial robot–object triadic interaction-guided diffusion model that captures the dynamic geometry robot-centric and object-centric methods miss. Adaptive Diffusion Constrained Sampling (WeI1I.112) enforces multiple simultaneous geometric constraints across high-DoF configuration spaces during sampling. VLM-SFD's SFDNet (see above) is a Siamese flow-diffusion network. Notably, How Well Do Diffusion Policies Learn Kinematic Constraint Manifolds? (WeI2I.233) is an analysis paper arguing that task success alone does not certify that a diffusion policy has actually learned the kinematic equality constraints in its data — a useful skeptical counterpoint to the cluster.

4. Teleoperation, data collection & data augmentation for two arms

Because two-arm data is scarce and expensive, a sizable group builds collection/augmentation pipelines. DexTele (TuI1I.123) is a dual-arm dexterous teleoperation system using motion retargeting plus adaptive force control for cross-platform generalization. CaFe-TeleVision (TuI1I.395) is a coarse-to-fine immersive teleoperation system focused on ergonomics. On the synthetic side, ROPA (WeI1I.57) generates synthetic robot poses for RGB-D bimanual data augmentation to broaden coverage over poses/contacts. MonoDuo (WeI1I.262) is a striking data-efficiency idea: learn bimanual policies using widely-available single-arm robots, sidestepping the scarcity of bimanual hardware (see human/single-arm-to-bimanual transfer below). ALOHA Lightning (WeI2I.216) attacks the speed gap, providing a high-speed demonstration interface and training recipe so learned policies run fast and precise rather than far slower than humans.

5. Active perception, gaze & multi-view fusion

Several papers note that two arms create occlusion and viewpoint problems and respond with active or selective vision. Observer–Actor (ObAct) (ThI1I.284, arXiv 2511.18140) dynamically assigns observer/actor roles: the observer arm builds a sparse-view 3D Gaussian-Splatting scene, virtually explores it for an optimal camera pose, moves there, then the actor arm executes — keeping observations near the occlusion-free training distribution. Look, Focus, Act (ThI1I.303, arXiv 2507.15833) brings human gaze and foveated ViT tokenization to ALOHA-style bimanual learning, cutting tokens/compute. Enhancing Reusability of Learned Skills via Gaze Information and Motion Bottlenecks (WeI1I.33) and BFA (WeI2I.18, arXiv 2502.11161) round this out — BFA's best-feature-aware fusion dynamically reweights multi-view cameras per task stage, reporting a 22–46% success improvement while cutting compute. Towards Exploratory and Focused Manipulation with Bimanual Active Perception (WeI2I.281) frames active perception as a new problem with its own benchmark.

6. Deformable & cloth/cable bimanual manipulation

Deformable-object handling is where two hands are most indispensable, and it forms one of the largest sub-clusters. Garments: Right-Side-Out (ThI2I.118, arXiv 2509.15953) is a zero-shot sim-to-real framework for turning garments right-side-out via depth-keypoint bimanual primitives; SIS (TuI1I.409) uses a seam-informed strategy for T-shirt unfolding; Transformer Driven Visual Servoing (TuI2I.382) matches fabric textures with a dual-arm manipulator. Cables / deformable linear objects: CRAFT (TuI2I.228) routes cables around fixtures with two caging grippers; Adaptive Curvature-Aware Routing (WeI2I.295) and Time-Series 3D Shape Control of DLOs (WeI2I.436) control stiff/flexible linear objects; Dual Arm Steering of Flexible Linear Objects (ThI2I.31) uses Euler's elastica closed-form solutions. Others: DiffDef (TuI1I.256) generates multimodal goal shapes for shape servoing; Dual Quaternion Compliant Movement Primitives (ThI1I.253) and Iterative Shaping of Multi-Particle Aggregates (TuI1I.7, with a VLM action-tree planner) extend the deformable theme to compliant motion and granular media.

7. Tactile, force & contact-rich dual-arm control

A systems cluster grounds two-arm manipulation in physical contact. TactileAloha (ThI2I.18) mounts a tactile sensor on an ALOHA gripper, fusing ResNet-encoded tactile signals with vision/proprioception in a transformer action-chunking policy for texture-dependent tasks (zip-tie insertion, Velcro fastening). Stereo-Based Vision and Tactile Sensing (TuI2LB.7) does robust dual-arm connector assembly of deformable wires. On the control side, Velocity-Based Admittance-Impedance Control (TuI2I.203) achieves stable force closure on position/velocity-only manipulators, and Bimanual Regrasp Planning (TuI2I.375) actively reduces object-pose uncertainty through regrasping.

8. Bimanual VLA, whole-body & cross-embodiment policies

The VLA wave reaches the bimanual cohort through whole-body and cross-embodiment work. TrajBooster (ThI1I.144, arXiv 2509.11839) extracts 6D dual-arm end-effector trajectories from wheeled humanoids, retargets them to a Unitree G1 whole-body controller, and post-pre-trains a VLA with only ~10 minutes of target-robot teleoperation — enabling squatting and cross-height bimanual manipulation. DSPv2 (TuI1I.157, arXiv 2509.16063) extends the Dense Policy paradigm to whole-body mobile manipulation by aligning 3D spatial with multi-view 2D semantic features. Residual Off-Policy RL (TuI2I.230, arXiv 2509.19301) fine-tunes behavior-cloning policies with sample-efficient off-policy RL and reports the first successful real-world RL training on a humanoid with dexterous hands (bimanual Vega platform, 29-D action space). See also the dedicated Dexora (36-DoF dual-arm/dual-hand VLA) and the cross-arm TwinVLA for the broader bimanual-VLA landscape.

9. Grasping, assembly & domain systems

Finally, several papers target specific two-arm capabilities. BiGraspFormer (WeI1I.292, arXiv 2509.19142) is an end-to-end transformer generating coordinated bimanual grasps from point clouds via a single-guided-bimanual strategy (<0.05 s inference). Bi-Adapt (WeAT1.3) does few-shot bimanual adaptation to novel 3D-object categories via semantic correspondence. Assembly: Leveraging Two Robotic Arms for Tight Assembly (TuI1I.159) and Search Strategy for Layered Peg-In-Hole (TuI2I.114). Domain/medical: Give Me Scissors (ThI1I.93, collision-free surgical instrument delivery) and A Transendoscopic Telerobotic System (WeI2I.332, bimanual endoscopic submucosal dissection). Hardware: ByteWrist (ThBT3.2, arXiv 2509.18084, ByteDance) is a compact three-stage parallel anthropomorphic wrist for confined-space dual-arm cooperation. Representation/eval: HeRO (TuI2I.167, hierarchical 3D semantic representation) and RoboEval (ThI1I.341, arXiv 2507.00435), an evaluation framework with eight bimanual tasks, 3000+ demos, and metrics for coordination/efficiency/safety beyond binary success.

Standout deep-dives

TrajBooster — cross-embodiment trajectory transfer to bipedal humanoid bimanual VLA

arXiv 2509.11839 (ThI1I.144). The headline number: after retargeting wheeled-humanoid dual-arm trajectories to a Unitree G1 whole-body controller, the VLA needs only ~10 minutes of target-domain teleoperation to enable beyond-tabletop tasks including squatting and coordinated cross-height bimanual motion. Uses end-effector trajectories as a morphology-agnostic interface and heterogeneous triplets (source vision/language + target actions). A clean answer to bimanual-humanoid data scarcity.

Residual Off-Policy RL for Finetuning BC Policies — first real-world RL on a dexterous humanoid

arXiv 2509.19301 (TuI2I.230; Ankile, Jiang, Duan, Shi, Abbeel, Nagabandi). Learns lightweight per-step residual corrections on top of a black-box BC policy via sample-efficient off-policy RL using only sparse binary rewards. Demonstrated on the bimanual wheeled Vega humanoid (two 7-DoF arms + two 6-DoF dexterous hands, 29-D action space) — claimed as the first successful real-world RL training on a humanoid robot with dexterous hands. A bridge between the BC and RL paradigms for high-DoF bimanual systems.

Right-Side-Out — zero-shot sim-to-real garment reversal

arXiv 2509.15953 (ThI2I.118). Decomposes the highly-dynamic, occlusion-heavy garment-reversal task into Drag/Fling (create + stabilize an opening) then Insert&Pull (invert), each a depth-keypoint-parameterized bimanual primitive. Trained entirely in a custom GPU-parallel MPM thin-shell simulator and deployed zero-shot on real hardware at up to 81.3% success. A strong example of structured bimanual primitives + physics-accurate sim closing the cloth sim-to-real gap.

Observer–Actor (ObAct) — active-vision role assignment via 3D Gaussian Splatting

arXiv 2511.18140 (ThI1I.284). Turns the two-arm occlusion problem into an asset: at test time one arm becomes an observer that builds a sparse-view 3DGS scene, virtually searches it for an optimal viewpoint, and physically moves there before the actor arm executes. The result keeps policy observations near the occlusion-free training distribution and supports ambidextrous role-swapping — a distinctive use of dual-arm hardware for perception rather than just manipulation.

ScheduleStream — GPU-accelerated multi-arm planning and scheduling

arXiv 2511.04758 (WeI1I.203; Garrett & Ramos, NVIDIA). Standard TAMP typically produces plans where only one arm moves at a time; ScheduleStream instead models hybrid durative actions that can start asynchronously and persist, producing schedules with genuine parallel arm motion, and uses GPU acceleration inside samplers to make this tractable. Billed as the first general-purpose framework for planning-and-scheduling with sampling operations.

RoboEval — a structured bimanual evaluation benchmark

arXiv 2507.00435 (ThI1I.341). Argues that collapsing performance into binary success counts hides execution-quality and failure structure. Provides eight bimanual tasks with controlled variations, 3000+ expert demos, and a modular sim platform, instrumenting every task with efficiency, bimanual-coordination, and safety/stability metrics plus stagewise outcome tracing. A useful shared yardstick as the bimanual learning cluster grows.

Complete paper list (50)

Code Title arXiv
ThBT3.2 ByteWrist: A Parallel Robotic Wrist Enabling Flexible and Anthropomorphic Motion for Confined Spaces 2509.18084
ThI1I.116 TOCALib: Optimal Control Library with Interpolation for Bimanual Manipulation and Obstacles Avoidance 2504.07708
ThI1I.128 Connectivity-Aware Representations for Constrained Motion Planning Via Multi-Scale Contrastive Learning 2603.25298
ThI1I.144 TrajBooster: Boosting Humanoid Whole-Body Manipulation Via Trajectory-Centric Learning 2509.11839
ThI1I.253 Dual Quaternion Based Compliant Movement Primitives for Deformable Object Manipulation —
ThI1I.268 PA-BiCoop: A Primary-Auxiliary Cooperative Framework for General Bimanual Manipulation —
ThI1I.284 Observer–Actor: Active Vision Imitation Learning with Sparse-View Gaussian Splatting 2511.18140
ThI1I.303 Look, Focus, Act: Efficient and Robust Robot Learning Via Human Gaze and Foveated Vision Transformers 2507.15833
ThI1I.341 RoboEval: Where Robotic Manipulation Meets Structured and Scalable Evaluation 2507.00435
ThI1I.71 Planning-Guided Diffusion Policy Learning for Contact-Rich Bimanual Object Reorientation 2412.02676
ThI1I.93 Give Me Scissors: Collision-Free Dual-Arm Surgical Assistive Robot for Instrument Delivery 2603.02553
ThI2I.118 Right-Side-Out: Learning Zero-Shot Sim-To-Real Garment Reversal 2509.15953
ThI2I.18 TactileAloha: Learning Bimanual Manipulation with Tactile Sensing RA-L 2025
ThI2I.181 RoTri-Diff: A Spatial Robot–Object Triadic Interaction-Guided Diffusion Model for Bimanual Manipulation 2603.07165
ThI2I.248 DAG-Plan: Generating Directed Acyclic Dependency Graphs for Dual-Arm Cooperative Planning 2406.09953
ThI2I.31 Dual Arm Steering of Flexible Linear Objects in 2-D and 3-D Environments Using Euler's Elastica Solutions 2502.07509
TuI1I.123 DexTele: A Dual-Arm Dexterous Teleoperation System Based on Motion Retargeting and Adaptive Force Control —
TuI1I.157 DSPv2: Improved Dense Policy for Effective and Generalizable Whole-Body Mobile Manipulation 2509.16063
TuI1I.159 Leveraging Two Robotic Arms for Tight Assembly Performance Gains —
TuI1I.256 DiffDef: A Diffusion Model for Generating Multimodal Goal Shapes from Demonstrations for Deformable Object Manipulation 2506.18779
TuI1I.395 CaFe-TeleVision: A Coarse-To-Fine Teleoperation System with Immersive Situated Visualization for Enhanced Ergonomics 2512.14270
TuI1I.409 SIS: Seam-Informed Strategy for T-Shirt Unfolding 2409.06990
TuI1I.7 Iterative Shaping of Multi-Particle Aggregates Based on Action Trees and VLM 2501.13507
TuI2I.114 Search Strategy for Layered Peg-In-Hole Using Dual Manipulator System —
TuI2I.167 HeRO: Hierarchical 3D Semantic Representation for Pose-Aware Object Manipulation 2602.18817
TuI2I.203 Velocity-Based Admittance-Impedance Control with Contact Compliance Modeling for Robust Dual-Arm Manipulation —
TuI2I.228 CRAFT: Long-Horizon Cable Routing Algorithm and Low-Friction Caging Gripper —
TuI2I.230 Residual Off-Policy RL for Finetuning Behavior Cloning Policies 2509.19301
TuI2I.235 Consensus Driven Dynamical Systems Control for Dual-Arm Handover —
TuI2I.365 VLM-SFD: VLM-Assisted Siamese Flow Diffusion Framework for Dual-Arm Cooperative Manipulation 2506.13428
TuI2I.375 Bimanual Regrasp Planning and Control for Active Reduction of Object Pose Uncertainty 2503.22240
TuI2I.382 Transformer Driven Visual Servoing for Fabric Texture Matching Using Dual-Arm Manipulator 2511.21203
TuI2LB.2 An Efficient Learning-Based Task Planning Approach Using a Bio-Inspired Action Context-Free Grammar for Bimanual Manipulation —
TuI2LB.7 Stereo-Based Vision and Tactile Sensing for Robust Dual-Arm Robotic Connector Assembly —
WeAT1.3 Bi-Adapt: Few-Shot Bimanual Adaptation for Novel Categories of 3D Objects Via Semantic Correspondence 2602.08425
WeI1I.112 Adaptive Diffusion Constrained Sampling for Bimanual Robot Manipulation 2505.13667
WeI1I.203 ScheduleStream: Temporal Planning with Samplers for GPU-Accelerated Multi-Arm Task and Motion Planning & Scheduling 2511.04758
WeI1I.262 MonoDuo: Using One Robot Arm to Learn Bimanual Policies —
WeI1I.292 BiGraspFormer: End-To-End Bimanual Grasp Transformer 2509.19142
WeI1I.33 Enhancing Reusability of Learned Skills for Robot Manipulation Via Gaze Information and Motion Bottlenecks 2502.18121
WeI1I.395 Impact-Aware Dual-Arm Manipulation —
WeI1I.57 ROPA: Synthetic Robot Pose Generation for RGB-D Bimanual Data Augmentation 2509.19454
WeI2I.18 BFA: Best-Feature-Aware Fusion for Multi-View Fine-Grained Manipulation 2502.11161
WeI2I.216 ALOHA Lightning: Learning Fast and Precise Manipulation —
WeI2I.224 High-Performance Dual-Arm Task and Motion Planning for Tabletop Rearrangement 2512.08206
WeI2I.233 How Well Do Diffusion Policies Learn Kinematic Constraint Manifolds? 2510.01404
WeI2I.281 Towards Exploratory and Focused Manipulation with Bimanual Active Perception: A New Problem, Benchmark and Strategy 2602.01939
WeI2I.295 Adaptive Curvature-Aware Routing for Stiff Cable Control Via Dual Manipulation —
WeI2I.332 A Transendoscopic Telerobotic System Using Heterogeneous Flexible Manipulators for Bimanual Endoscopic Submucosal Dissection —
WeI2I.436 Time-Series Data-Driven Three Dimensional Shape Control of Deformable Linear Objects Using a Dual-Arm Robot with Dynamic Model Updating —

Related

← Back to ICRA-2026-VLA-Manipulation-Survey