CVPR 2026 UniDex - Heungwoo/research GitHub Wiki
Venue: CVPR 2026 Category: Egocentric + Dexterous VLA Trend tag: Trend 5 (egocentric → manipulation)
flowchart LR
EGO["50k+ egocentric<br/>human-hand videos"] --> RETARGET["retarget to 8 dex hands"]
RETARGET --> FAAS["Function-Actuator-Aligned Space"]
FAAS --> VLA["3D VLA backbone"]
ROBOT["unseen dex hand"] --> VLA
VLA --> ACT["action chunk"]
Dexterous manipulation needs data per-hand because of the actuator-count mismatch (humans: ~20 DoF; robot hands: 6–24 active DoF, different topology). Existing dex datasets are tiny relative to web video. Egocentric human-hand video is abundant but cannot be directly used without bridging the actuator gap.
- Collect / curate 50 000+ retargeted trajectories (~9M paired image–pointcloud–action frames) across 8 dexterous hands (Inspire, Leap, Shadow, Allegro, Ability, Oymotion, Xhand, Wuji), via human-in-the-loop retargeting from egocentric human video.
- Define the Function-Actuator-Aligned Space (FAAS) — a unified action space that maps functionally similar actuators to shared coordinates regardless of the underlying actuator topology, enabling cross-hand transfer.
- Train a 3D VLA (UniDex-VLA) on the unified FAAS with multi-hand training: a Uni3D point-cloud encoder (replacing the SigLIP 2D encoder) on a PaliGemma backbone, trained with a conditional flow-matching objective.
- Ship UniDex-Cap, a portable capture setup for human–robot data co-training.
81 % average task progress across five real-world tool-use tasks (Make Coffee, Sweep Objects, Water Flowers, Cut Bags, Use Mouse) on two hands (Inspire, Wuji), outperforming prior VLA baselines by a large margin. Separately, a policy trained on the Inspire Hand exhibits zero-shot cross-hand transfer, reaching 60 % success on Oymotion and 40 % on Wuji without any fine-tuning, alongside spatial and object generalization.
UniDex is the dex-hand counterpart to X-VLA's soft-prompt cross-embodiment thesis: the right factorization (FAAS for dex hands; soft prompts for arms) lets one policy generalize across mechanically different end-effectors. Likely to become a baseline for subsequent dex-VLA work. Thematically related (concurrent ego-video work, different team): EgoScale (data scaling for ego video).
From Tsinghua University, Shanghai Qizhi Institute, Sun Yat-sen University, and UNC Chapel Hill.
- arXiv: 2603.22264
- Code:
unidex-ai/UniDex
- Dexterous Manipulation review · Cross-Embodiment review
- DexUMI · EgoDex · Human-Video Pretraining
- CVPR 2026 survey
← Back to CVPR-2026