ICRA 2026 Topic Bimanual - Heungwoo/research GitHub Wiki
ICRA 2026 — Bimanual & Dual-Arm Manipulation (Topic Analysis)
Venue: IEEE ICRA 2026 · Vienna, Austria · June 1–5, 2026 Compiled against the official PaperCept program. This page covers the 50 papers that the program tags with Bimanual Manipulation or Dual Arm Manipulation as a primary/secondary keyword. ← Back to ICRA-2026-VLA-Manipulation-Survey
Overview
Bimanual and dual-arm manipulation is one of ICRA 2026's fastest-growing manipulation sub-fields: the program's topic-keyword counts put Bimanual at 36 and Dual Arm Manipulation as a distinct large cluster, and once both tags are pooled the result is the 50-paper group analyzed here. The growth tracks the broader move toward two-arm hardware platforms — ALOHA/Mobile-ALOHA, Unitree G1, and assorted dual-arm industrial cells — and the recognition that most genuinely useful household and industrial tasks (folding, assembly, handover, deformable-object handling, surgery) are intrinsically two-handed.
The unifying technical problem is coordination: two arms share a workspace, a kinematic budget, and often a single object, so naive single-arm policies, planners, and grasp generators do not transfer. The 50 papers attack this from every direction at once — mechanism design, optimal control, TAMP, imitation/diffusion policies, dexterous teleoperation, and bimanual VLA. A notable feature of the ICRA cohort (versus ML-venue work) is its systems orientation: many entries are about hardware (wrists, grippers, surgical tools), force/tactile control, and real-robot data collection rather than pure architecture. Where the ICLR/CVPR bimanual wave is dominated by VLA backbones, ICRA's center of gravity sits on planning, control, and deformable-object physics.
Sub-trends
1. Coordination frameworks, role assignment & dynamic division of labor
The defining bimanual challenge — who does what, when — gets explicit treatment. PA-BiCoop (ThI1I.268) frames coordination as a primary–auxiliary cooperative structure with dynamic division of labor, addressing methods that either lack inter-arm interaction or fix the labor split. DAG-Plan (ThI2I.248, arXiv 2406.09953) uses LLMs to decompose long-horizon tasks into a directed acyclic graph with explicit temporal dependencies, then dynamically assigns sub-tasks to arms based on real-time observations — modeling parallelism that linear LLM plans cannot. VLM-SFD (TuI2I.365, arXiv 2506.13428) similarly uses a pre-trained VLM to adaptively assign the optimal object-centric motion flow to each arm over time. Consensus Driven Dynamical Systems Control for Dual-Arm Handover (TuI2I.235) handles the tight spatial-and-temporal coordination of object transfer between two arms. Impact-Aware Dual-Arm Manipulation (WeI1I.395) targets logistics depalletizing where the two arms must grab/place under impact dynamics.
2. Bimanual / dual-arm task & motion planning
A robust planning cluster tackles the combinatorial blow-up of two-arm action spaces. ScheduleStream (WeI1I.203, arXiv 2511.04758, NVIDIA) is the first general framework for planning and scheduling with samplers, using GPU acceleration and hybrid durative actions so arms can move in parallel rather than one-at-a-time. High-Performance Dual-Arm TAMP for Tabletop Rearrangement (WeI2I.224) presents the Synchronous Dual-Arm Rearrangement Planner (SDAR) for close-proximity rearrangement. TOCALib (ThI1I.116, arXiv 2504.07708) builds a library of optimal two-arm trajectories with DCOL-based symbolic collision expressions inside the FROST framework (demonstrated on Mobile ALOHA). An Efficient Learning-Based Task Planning Approach (TuI2LB.2) introduces a bio-inspired action context-free grammar to curb the TAMP combinatorial explosion, while Connectivity-Aware Representations (ThI1I.128) uses multi-scale contrastive learning to connect disconnected constraint regions for bimanual motion planning.
3. Bimanual diffusion / flow policies & constraint-aware learning
Diffusion and flow models adapted to two-arm coordination form a dense methodological cluster. Planning-Guided Diffusion Policy Learning (ThI1I.71) targets contact-rich bimanual object reorientation. RoTri-Diff (ThI2I.181) is a spatial robot–object triadic interaction-guided diffusion model that captures the dynamic geometry robot-centric and object-centric methods miss. Adaptive Diffusion Constrained Sampling (WeI1I.112) enforces multiple simultaneous geometric constraints across high-DoF configuration spaces during sampling. VLM-SFD's SFDNet (see above) is a Siamese flow-diffusion network. Notably, How Well Do Diffusion Policies Learn Kinematic Constraint Manifolds? (WeI2I.233) is an analysis paper arguing that task success alone does not certify that a diffusion policy has actually learned the kinematic equality constraints in its data — a useful skeptical counterpoint to the cluster.
4. Teleoperation, data collection & data augmentation for two arms
Because two-arm data is scarce and expensive, a sizable group builds collection/augmentation pipelines. DexTele (TuI1I.123) is a dual-arm dexterous teleoperation system using motion retargeting plus adaptive force control for cross-platform generalization. CaFe-TeleVision (TuI1I.395) is a coarse-to-fine immersive teleoperation system focused on ergonomics. On the synthetic side, ROPA (WeI1I.57) generates synthetic robot poses for RGB-D bimanual data augmentation to broaden coverage over poses/contacts. MonoDuo (WeI1I.262) is a striking data-efficiency idea: learn bimanual policies using widely-available single-arm robots, sidestepping the scarcity of bimanual hardware (see human/single-arm-to-bimanual transfer below). ALOHA Lightning (WeI2I.216) attacks the speed gap, providing a high-speed demonstration interface and training recipe so learned policies run fast and precise rather than far slower than humans.
5. Active perception, gaze & multi-view fusion
Several papers note that two arms create occlusion and viewpoint problems and respond with active or selective vision. Observer–Actor (ObAct) (ThI1I.284, arXiv 2511.18140) dynamically assigns observer/actor roles: the observer arm builds a sparse-view 3D Gaussian-Splatting scene, virtually explores it for an optimal camera pose, moves there, then the actor arm executes — keeping observations near the occlusion-free training distribution. Look, Focus, Act (ThI1I.303, arXiv 2507.15833) brings human gaze and foveated ViT tokenization to ALOHA-style bimanual learning, cutting tokens/compute. Enhancing Reusability of Learned Skills via Gaze Information and Motion Bottlenecks (WeI1I.33) and BFA (WeI2I.18, arXiv 2502.11161) round this out — BFA's best-feature-aware fusion dynamically reweights multi-view cameras per task stage, reporting a 22–46% success improvement while cutting compute. Towards Exploratory and Focused Manipulation with Bimanual Active Perception (WeI2I.281) frames active perception as a new problem with its own benchmark.
6. Deformable & cloth/cable bimanual manipulation
Deformable-object handling is where two hands are most indispensable, and it forms one of the largest sub-clusters. Garments: Right-Side-Out (ThI2I.118, arXiv 2509.15953) is a zero-shot sim-to-real framework for turning garments right-side-out via depth-keypoint bimanual primitives; SIS (TuI1I.409) uses a seam-informed strategy for T-shirt unfolding; Transformer Driven Visual Servoing (TuI2I.382) matches fabric textures with a dual-arm manipulator. Cables / deformable linear objects: CRAFT (TuI2I.228) routes cables around fixtures with two caging grippers; Adaptive Curvature-Aware Routing (WeI2I.295) and Time-Series 3D Shape Control of DLOs (WeI2I.436) control stiff/flexible linear objects; Dual Arm Steering of Flexible Linear Objects (ThI2I.31) uses Euler's elastica closed-form solutions. Others: DiffDef (TuI1I.256) generates multimodal goal shapes for shape servoing; Dual Quaternion Compliant Movement Primitives (ThI1I.253) and Iterative Shaping of Multi-Particle Aggregates (TuI1I.7, with a VLM action-tree planner) extend the deformable theme to compliant motion and granular media.
7. Tactile, force & contact-rich dual-arm control
A systems cluster grounds two-arm manipulation in physical contact. TactileAloha (ThI2I.18) mounts a tactile sensor on an ALOHA gripper, fusing ResNet-encoded tactile signals with vision/proprioception in a transformer action-chunking policy for texture-dependent tasks (zip-tie insertion, Velcro fastening). Stereo-Based Vision and Tactile Sensing (TuI2LB.7) does robust dual-arm connector assembly of deformable wires. On the control side, Velocity-Based Admittance-Impedance Control (TuI2I.203) achieves stable force closure on position/velocity-only manipulators, and Bimanual Regrasp Planning (TuI2I.375) actively reduces object-pose uncertainty through regrasping.
8. Bimanual VLA, whole-body & cross-embodiment policies
The VLA wave reaches the bimanual cohort through whole-body and cross-embodiment work. TrajBooster (ThI1I.144, arXiv 2509.11839) extracts 6D dual-arm end-effector trajectories from wheeled humanoids, retargets them to a Unitree G1 whole-body controller, and post-pre-trains a VLA with only ~10 minutes of target-robot teleoperation — enabling squatting and cross-height bimanual manipulation. DSPv2 (TuI1I.157, arXiv 2509.16063) extends the Dense Policy paradigm to whole-body mobile manipulation by aligning 3D spatial with multi-view 2D semantic features. Residual Off-Policy RL (TuI2I.230, arXiv 2509.19301) fine-tunes behavior-cloning policies with sample-efficient off-policy RL and reports the first successful real-world RL training on a humanoid with dexterous hands (bimanual Vega platform, 29-D action space). See also the dedicated Dexora (36-DoF dual-arm/dual-hand VLA) and the cross-arm TwinVLA for the broader bimanual-VLA landscape.
9. Grasping, assembly & domain systems
Finally, several papers target specific two-arm capabilities. BiGraspFormer (WeI1I.292, arXiv 2509.19142) is an end-to-end transformer generating coordinated bimanual grasps from point clouds via a single-guided-bimanual strategy (<0.05 s inference). Bi-Adapt (WeAT1.3) does few-shot bimanual adaptation to novel 3D-object categories via semantic correspondence. Assembly: Leveraging Two Robotic Arms for Tight Assembly (TuI1I.159) and Search Strategy for Layered Peg-In-Hole (TuI2I.114). Domain/medical: Give Me Scissors (ThI1I.93, collision-free surgical instrument delivery) and A Transendoscopic Telerobotic System (WeI2I.332, bimanual endoscopic submucosal dissection). Hardware: ByteWrist (ThBT3.2, arXiv 2509.18084, ByteDance) is a compact three-stage parallel anthropomorphic wrist for confined-space dual-arm cooperation. Representation/eval: HeRO (TuI2I.167, hierarchical 3D semantic representation) and RoboEval (ThI1I.341, arXiv 2507.00435), an evaluation framework with eight bimanual tasks, 3000+ demos, and metrics for coordination/efficiency/safety beyond binary success.
Standout deep-dives
TrajBooster — cross-embodiment trajectory transfer to bipedal humanoid bimanual VLA
arXiv 2509.11839 (ThI1I.144). The headline number: after retargeting wheeled-humanoid dual-arm trajectories to a Unitree G1 whole-body controller, the VLA needs only ~10 minutes of target-domain teleoperation to enable beyond-tabletop tasks including squatting and coordinated cross-height bimanual motion. Uses end-effector trajectories as a morphology-agnostic interface and heterogeneous triplets (source vision/language + target actions). A clean answer to bimanual-humanoid data scarcity.
Residual Off-Policy RL for Finetuning BC Policies — first real-world RL on a dexterous humanoid
arXiv 2509.19301 (TuI2I.230; Ankile, Jiang, Duan, Shi, Abbeel, Nagabandi). Learns lightweight per-step residual corrections on top of a black-box BC policy via sample-efficient off-policy RL using only sparse binary rewards. Demonstrated on the bimanual wheeled Vega humanoid (two 7-DoF arms + two 6-DoF dexterous hands, 29-D action space) — claimed as the first successful real-world RL training on a humanoid robot with dexterous hands. A bridge between the BC and RL paradigms for high-DoF bimanual systems.
Right-Side-Out — zero-shot sim-to-real garment reversal
arXiv 2509.15953 (ThI2I.118). Decomposes the highly-dynamic, occlusion-heavy garment-reversal task into Drag/Fling (create + stabilize an opening) then Insert&Pull (invert), each a depth-keypoint-parameterized bimanual primitive. Trained entirely in a custom GPU-parallel MPM thin-shell simulator and deployed zero-shot on real hardware at up to 81.3% success. A strong example of structured bimanual primitives + physics-accurate sim closing the cloth sim-to-real gap.
Observer–Actor (ObAct) — active-vision role assignment via 3D Gaussian Splatting
arXiv 2511.18140 (ThI1I.284). Turns the two-arm occlusion problem into an asset: at test time one arm becomes an observer that builds a sparse-view 3DGS scene, virtually searches it for an optimal viewpoint, and physically moves there before the actor arm executes. The result keeps policy observations near the occlusion-free training distribution and supports ambidextrous role-swapping — a distinctive use of dual-arm hardware for perception rather than just manipulation.
ScheduleStream — GPU-accelerated multi-arm planning and scheduling
arXiv 2511.04758 (WeI1I.203; Garrett & Ramos, NVIDIA). Standard TAMP typically produces plans where only one arm moves at a time; ScheduleStream instead models hybrid durative actions that can start asynchronously and persist, producing schedules with genuine parallel arm motion, and uses GPU acceleration inside samplers to make this tractable. Billed as the first general-purpose framework for planning-and-scheduling with sampling operations.
RoboEval — a structured bimanual evaluation benchmark
arXiv 2507.00435 (ThI1I.341). Argues that collapsing performance into binary success counts hides execution-quality and failure structure. Provides eight bimanual tasks with controlled variations, 3000+ expert demos, and a modular sim platform, instrumenting every task with efficiency, bimanual-coordination, and safety/stability metrics plus stagewise outcome tracing. A useful shared yardstick as the bimanual learning cluster grows.
Complete paper list (50)
| Code | Title | arXiv |
|---|---|---|
| ThBT3.2 | ByteWrist: A Parallel Robotic Wrist Enabling Flexible and Anthropomorphic Motion for Confined Spaces | 2509.18084 |
| ThI1I.116 | TOCALib: Optimal Control Library with Interpolation for Bimanual Manipulation and Obstacles Avoidance | 2504.07708 |
| ThI1I.128 | Connectivity-Aware Representations for Constrained Motion Planning Via Multi-Scale Contrastive Learning | 2603.25298 |
| ThI1I.144 | TrajBooster: Boosting Humanoid Whole-Body Manipulation Via Trajectory-Centric Learning | 2509.11839 |
| ThI1I.253 | Dual Quaternion Based Compliant Movement Primitives for Deformable Object Manipulation | — |
| ThI1I.268 | PA-BiCoop: A Primary-Auxiliary Cooperative Framework for General Bimanual Manipulation | — |
| ThI1I.284 | Observer–Actor: Active Vision Imitation Learning with Sparse-View Gaussian Splatting | 2511.18140 |
| ThI1I.303 | Look, Focus, Act: Efficient and Robust Robot Learning Via Human Gaze and Foveated Vision Transformers | 2507.15833 |
| ThI1I.341 | RoboEval: Where Robotic Manipulation Meets Structured and Scalable Evaluation | 2507.00435 |
| ThI1I.71 | Planning-Guided Diffusion Policy Learning for Contact-Rich Bimanual Object Reorientation | 2412.02676 |
| ThI1I.93 | Give Me Scissors: Collision-Free Dual-Arm Surgical Assistive Robot for Instrument Delivery | 2603.02553 |
| ThI2I.118 | Right-Side-Out: Learning Zero-Shot Sim-To-Real Garment Reversal | 2509.15953 |
| ThI2I.18 | TactileAloha: Learning Bimanual Manipulation with Tactile Sensing | RA-L 2025 |
| ThI2I.181 | RoTri-Diff: A Spatial Robot–Object Triadic Interaction-Guided Diffusion Model for Bimanual Manipulation | 2603.07165 |
| ThI2I.248 | DAG-Plan: Generating Directed Acyclic Dependency Graphs for Dual-Arm Cooperative Planning | 2406.09953 |
| ThI2I.31 | Dual Arm Steering of Flexible Linear Objects in 2-D and 3-D Environments Using Euler's Elastica Solutions | 2502.07509 |
| TuI1I.123 | DexTele: A Dual-Arm Dexterous Teleoperation System Based on Motion Retargeting and Adaptive Force Control | — |
| TuI1I.157 | DSPv2: Improved Dense Policy for Effective and Generalizable Whole-Body Mobile Manipulation | 2509.16063 |
| TuI1I.159 | Leveraging Two Robotic Arms for Tight Assembly Performance Gains | — |
| TuI1I.256 | DiffDef: A Diffusion Model for Generating Multimodal Goal Shapes from Demonstrations for Deformable Object Manipulation | 2506.18779 |
| TuI1I.395 | CaFe-TeleVision: A Coarse-To-Fine Teleoperation System with Immersive Situated Visualization for Enhanced Ergonomics | 2512.14270 |
| TuI1I.409 | SIS: Seam-Informed Strategy for T-Shirt Unfolding | 2409.06990 |
| TuI1I.7 | Iterative Shaping of Multi-Particle Aggregates Based on Action Trees and VLM | 2501.13507 |
| TuI2I.114 | Search Strategy for Layered Peg-In-Hole Using Dual Manipulator System | — |
| TuI2I.167 | HeRO: Hierarchical 3D Semantic Representation for Pose-Aware Object Manipulation | 2602.18817 |
| TuI2I.203 | Velocity-Based Admittance-Impedance Control with Contact Compliance Modeling for Robust Dual-Arm Manipulation | — |
| TuI2I.228 | CRAFT: Long-Horizon Cable Routing Algorithm and Low-Friction Caging Gripper | — |
| TuI2I.230 | Residual Off-Policy RL for Finetuning Behavior Cloning Policies | 2509.19301 |
| TuI2I.235 | Consensus Driven Dynamical Systems Control for Dual-Arm Handover | — |
| TuI2I.365 | VLM-SFD: VLM-Assisted Siamese Flow Diffusion Framework for Dual-Arm Cooperative Manipulation | 2506.13428 |
| TuI2I.375 | Bimanual Regrasp Planning and Control for Active Reduction of Object Pose Uncertainty | 2503.22240 |
| TuI2I.382 | Transformer Driven Visual Servoing for Fabric Texture Matching Using Dual-Arm Manipulator | 2511.21203 |
| TuI2LB.2 | An Efficient Learning-Based Task Planning Approach Using a Bio-Inspired Action Context-Free Grammar for Bimanual Manipulation | — |
| TuI2LB.7 | Stereo-Based Vision and Tactile Sensing for Robust Dual-Arm Robotic Connector Assembly | — |
| WeAT1.3 | Bi-Adapt: Few-Shot Bimanual Adaptation for Novel Categories of 3D Objects Via Semantic Correspondence | 2602.08425 |
| WeI1I.112 | Adaptive Diffusion Constrained Sampling for Bimanual Robot Manipulation | 2505.13667 |
| WeI1I.203 | ScheduleStream: Temporal Planning with Samplers for GPU-Accelerated Multi-Arm Task and Motion Planning & Scheduling | 2511.04758 |
| WeI1I.262 | MonoDuo: Using One Robot Arm to Learn Bimanual Policies | — |
| WeI1I.292 | BiGraspFormer: End-To-End Bimanual Grasp Transformer | 2509.19142 |
| WeI1I.33 | Enhancing Reusability of Learned Skills for Robot Manipulation Via Gaze Information and Motion Bottlenecks | 2502.18121 |
| WeI1I.395 | Impact-Aware Dual-Arm Manipulation | — |
| WeI1I.57 | ROPA: Synthetic Robot Pose Generation for RGB-D Bimanual Data Augmentation | 2509.19454 |
| WeI2I.18 | BFA: Best-Feature-Aware Fusion for Multi-View Fine-Grained Manipulation | 2502.11161 |
| WeI2I.216 | ALOHA Lightning: Learning Fast and Precise Manipulation | — |
| WeI2I.224 | High-Performance Dual-Arm Task and Motion Planning for Tabletop Rearrangement | 2512.08206 |
| WeI2I.233 | How Well Do Diffusion Policies Learn Kinematic Constraint Manifolds? | 2510.01404 |
| WeI2I.281 | Towards Exploratory and Focused Manipulation with Bimanual Active Perception: A New Problem, Benchmark and Strategy | 2602.01939 |
| WeI2I.295 | Adaptive Curvature-Aware Routing for Stiff Cable Control Via Dual Manipulation | — |
| WeI2I.332 | A Transendoscopic Telerobotic System Using Heterogeneous Flexible Manipulators for Bimanual Endoscopic Submucosal Dissection | — |
| WeI2I.436 | Time-Series Data-Driven Three Dimensional Shape Control of Deformable Linear Objects Using a Dual-Arm Robot with Dynamic Model Updating | — |
Related
- ICRA 2026 Survey
- Dexterous Manipulation review
- Dexora — open-source 36-DoF dual-arm/dual-hand VLA
- TwinVLA — cross-arm bimanual VLA
← Back to ICRA-2026-VLA-Manipulation-Survey