ICRA 2026 Dexora - Heungwoo/research GitHub Wiki
Venue: ICRA 2026 · Authors: Zongzheng Zhang, Jingrui Pang, Zhuo Yang, Kun Li, Minwen Liao, Saining Zhang, Guoxuan Chi, et al. (Tsinghua / BAAI / PKU) · arXiv: 2605.18722 Category: Dexterous / bimanual VLA Trend tag: Open-source high-DoF dexterity
flowchart LR
subgraph Teleop[Hybrid teleoperation]
EXO[Exoskeleton backpack<br/>gross arm kinematics] --> PLT[Dual-arm dual-hand<br/>platform + MuJoCo twin]
AVP[Apple Vision Pro<br/>markerless finger tracking] --> PLT
end
PLT --> DATA[(Data: 100K sim<br/>+ 10K real episodes)]
DATA --> DISC[Data-quality discriminator<br/>PU objective → loss weights]
V[Multi-view RGB] --> ENC[SigLIP vision + T5 text]
L[Language] --> ENC
ENC --> BB[Decoder-only Transformer<br/>28 layers · 1024 hidden · 16 heads]
DISC -. weighted loss .-> BB
BB --> HEAD[Diffusion transformer head<br/>DDPM train · DPMSolver++ infer]
HEAD --> ACT[36-DoF bimanual action<br/>2×6 arm + 2×12 hand]
Most open VLAs target single-arm grippers; dexterous, bimanual, high-DoF control is gated behind closed hardware and proprietary data. Two coupled obstacles: (1) collecting high-quality demonstrations for many-DoF hands is hard, and naive teleoperation conflates gross arm motion with fine finger motion; (2) real demonstration sets are noisy, so uniform imitation wastes capacity on bad trajectories. Dexora aims to be a fully open-source stack — hardware, data, and policy — for dual-arm dual-hand dexterity.
Platform. AIRBOT 6-DoF arms paired with XHAND dexterous hands (12 fully actuated joints each), giving a 36-DoF bimanual system (2×6 arm + 2×12 hand), mirrored by an identical MuJoCo digital twin.
Hybrid teleoperation. A custom exoskeleton backpack captures gross arm kinematics while an Apple Vision Pro provides markerless finger tracking — decoupling coarse arm motion from fine hand motion so high-DoF demonstrations stay clean.
Data. Pretrain on 100K simulated bimanual-hand trajectories; post-train on 10K real teleoperated episodes (the released real-world set: 12.2K episodes, 2.92M frames, 40.5 hours; 347 objects across 17 categories).
Architecture. SigLIP encodes multi-view RGB and T5 encodes language into a decoder-only Transformer backbone (28 layers, hidden 1024, 16 heads); a diffusion-transformer head predicts action chunks (DDPM in training, DPMSolver++ for fast inference).
Discriminator-guided training. A learned data-quality discriminator scores each demonstration via a positive–unlabeled objective; scores become per-sample weights on the diffusion loss, so the policy prioritizes high-quality trajectories and down-weights poor ones. This is a discriminator-weighted imitation variant tailored to noisy high-DoF data.
This sits in the VLM + separate diffusion action expert family (see VLA Architectures review).
On basic tasks Dexora reaches 89.6 average success, ahead of GR00T N1 (82.1), π₀ (50.4), and Diffusion Policy (34.2). On dexterous tasks it averages 66.7% vs. GR00T N1 51.7%, π₀ 26.7%, and DP 6.7%. Per-task dexterous results include Fetch Book 80%, Cut Leek 80%, Rough Dough 80%, Place Plates 70%, Use Pen 65%, Twist Cap 25%. The paper reports robust out-of-distribution and cross-embodiment generalization.
Dexora is presented as the first open-source VLA natively targeting dual-arm, dual-hand high-DoF manipulation, lowering the barrier for dexterous manipulation research. Two transferable ideas: hybrid arm/finger teleoperation for clean high-DoF data, and discriminator-weighted diffusion training that turns data-quality into a learning signal rather than relying on manual curation.
- Dexterous Manipulation review
- VLA Architectures review
- DexVLA (sibling plug-in diffusion action expert)
- ICRA 2026 Survey
← Back to ICRA-2026