Review Dexterous Hand Data Pyramid - Heungwoo/research GitHub Wiki
Goal: classify the data-acquisition/generation methods 2024β2026 papers & companies propose for training policies on a humanoid five-finger dexterous hand, as a data pyramid (broad-abundant-far-from-hardware base β scarce-high-fidelity-on-hardware tip). Two hand-specific axes cut across every tier: the retargeting gap (human 5-finger β robot 5-finger) and force/tactile (the signal lost first as you descend). Companions: Dexterous Manipulation Β· Human Video β Robot Transfer Β· Tactile VLA Β· Cross-Embodiment.
flowchart TB
L6["L6 Β· Target-hand teleop / real demos<br/>scarce Β· gold-standard Β· on-hardware"]
L5["L5 Β· Simulation / synthetic grasp<br/>billion-scale Β· physics-labeled (force in sim)"]
L4["L4 Β· Retargeting (human β robot 5-finger)<br/>the load-bearing bridge"]
L3["L3 Β· Wearable glove / exoskeleton<br/>action-aligned finger kinematics + contact"]
L2["L2 Β· Egocentric instrumented dex datasets<br/>dense hand-pose (+ some pressure)"]
L1["L1 Β· Web / internet human-hand video<br/>unbounded Β· no action/force labels"]
L6 --- L5 --- L4 --- L3 --- L2 --- L1
T["cross-cut Β· TACTILE / FORCE (inject at L3/L5/L6)"] -.-> L4
Quantity grows downward; fidelity to the actual 5-finger hardware grows upward. The 2026 shift (see Β§5) is that the pyramid is inverting β human video (base) is becoming the primary data, with teleop reserved for the scarce tip.
| Tier | What it is | Pros | Cons (hand-specific) |
|---|---|---|---|
| L1 Web/internet human video | uninstrumented RGB human hands (web, YouTube, ego) | unbounded scale; rich contact & object-interaction priors; ~free | no action or force labels; morphology & retargeting gap; viewpoint/occlusion noise |
| L2 Egocentric instrumented dex | Vision-Pro / MoCap hand-pose, sometimes pressure | dense, curated hand-pose labels; some tactile | capture cost; still a human hand; task coverage bounded |
| L3 Wearable glove / exoskeleton | sensorized glove/exo captures finger kinematics + contact | action-aligned; natural grip diversity; captures contact; no robot needed | gloveβrobot-hand retarget loss; some contact cues (palm/compliance) lost; device-dependent |
| L4 Retargeting (bridge) | map human hand β robot 5-finger DoF | turns all human data (L1βL3) into robot-hand data; force-aware variants preserve contact | fidelity is the ceiling; kinematic-only retargeting drops force; calibration burden |
| L5 Simulation / synthetic grasp | physics-generated grasps/trajectories at billion scale | unbounded; physics-labeled (force available); safe; morphology-swappable | sim-to-real contact/force gap; grasp-centric (not full manipulation); asset/scene coverage |
| L6 Target-hand teleop / real | teleop or autonomous demos on the actual hand | exact embodiment; real dynamics & contact β gold standard | most expensive; lowest volume; operator burden |
| cross-cut: Tactile / force | fingertip/skin tactile + 6-axis F/T | the signal contact-rich dex actually needs; enables force control | sensor-specific; nearly absent in L1βL2; cross-sensor generalization gap |
Marks reflect training necessity, eval-only shown separately: β required for training Β· β auxiliary (optional in training) Β· β³ eval / deploy only β not training Β· blank = not used. Last column: is target-hand teleop/robot data used for fine-tuning? β Yes (+ volume) Β· eval-only Β· No. π = five-/multi-finger dexterous hand Β· π€ = gripper / 2-finger.
| Paper / method | L1 web-vid | L2 ego | L3 wearable | L4 retarget | L5 sim | L6 target-hand | Tac | Teleop β fine-tune? |
|---|---|---|---|---|---|---|---|---|
| π five-/multi-finger dexterous hand | ||||||||
| EgoScale (NVIDIA/Berkeley/UMD, 20,854 h) | β | β | β | β | Yes β human-robot play mid-training | |||
| UniDex (8 dex hands, 10M frames) | β | β | β | β | Yes | |||
| Being-H0/H0.7 (PKU, 200k h + 15k robot) | β | β | β | Yes β 15k h robot | ||||
| T-Rex (2606.17055, tactile-reactive) | β | β | Yes β 100 h teleop | |||||
| Shared-Autonomy VR+HandVLA (2511.00139) | β | β | β | Yes β teleop collects training data | ||||
| π Dexterous Point Policy (KAIST, 2606.10614) | β | β | β³ | eval-only β 75% vs VLA 1% | ||||
| DO AS I DO (Berkeley/NYU, 4D retarget) | β | β | β | β³ | eval-only | |||
| VideoManip / RGB-video (2602.09013) | β | β | β³ | eval-only | ||||
| DexUMI (sensorized glove) | β | β | β³ | β | No β glove-trained; robot = deploy | |||
| DexEXO (UCLA, wearable exo) | β | β | β³ | No β exo-trained; robot = deploy | ||||
| AnyDexRT (SJTU, calib-free retarget) | β | β³ | eval-only β downstream training untested | |||||
| Dex1B / SynGrasp-1B / GraspVLA (billion synth) | β | β | No β sim-trained | |||||
| DexGrasp-Zero (morphology-aligned) | β | β | No β sim-trained | |||||
| DexNDM / RobustDexGrasp (sim2real RL) | β | β³ | β | No β sim-trained; real = deploy | ||||
| One-Hand (cross-hand canonicalize) | β | β | No β retarget/sim; robot = eval | |||||
| π€ gripper / 2-finger (reference) | ||||||||
| YUBI (AIST Β· 2-finger bidigital, handheld) | β | β³ | No β handheld data = training (8,434 h) |
Reading it (papers): foundation-scale methods (EgoScale/UniDex/Being-H0/T-Rex) fine-tune on a target-hand teleop tip (β L6, "Yes"); device-free-video (Dexterous Point Policy, DO AS I DO) and capture-interface (DexUMI/DexEXO) methods keep the robot for eval/deploy only (β³).
| Company / program (hand) | L1 web-vid | L2 ego | L3 wearable | L4 retarget | L5 sim | L6 target-hand | Tac | Teleop β fine-tune? |
|---|---|---|---|---|---|---|---|---|
| π five-/multi-finger dexterous-hand companies | ||||||||
| Genesis GENE-26.5 (Genesis AI Β· 5-finger + glove) | β | β | β | β³ | β | β | Yes β < 1 h (< 200 ep); glove+video pretrain (>200k h), zero sim training (sim = eval) | |
| RLDX-1 (RLWRLD Β· ALLEX 5-finger) | β | β | β | β | β | β | Yes β small teleop set (+ synthetic training) | |
| GR-Dexter (ByteDance Seed Β· 21-DoF ByteDexter, 2512.24210) | β | β | β | Yes β ~20 h/task (trains VLM + action DiT) | ||||
| Figure (Figure 03 Β· Helix 02 Β· ~16-DoF) | β | β | β | Yes β 10 Gbps fleet offload β fleet-wide training | ||||
| Tesla (Optimus Gen 3 Β· 22-DoF, tendon) | β | β | β | Yes β teleop center + mocap + factory data | ||||
| Unitree (G1/H1 Β· Dex3-1 3-finger / Dex5, 2510.08807) | β | Yes β open teleop + open full-body dataset | ||||||
| Fourier (GR-1/GR-2 Β· dex hand + ActionNet) | β | Yes β ~140 h teleop dataset | ||||||
| Teleop tool vendors (MANUS Β· Open-TeleVision Β· TESOLLO Β· ByteDexter) | β | β | β | Yes β teleop is the data engine | ||||
| π€ gripper-based / generalist companies | ||||||||
| Physical Intelligence (Ο0.5 / Ο0.7 Β· parallel-jaw gripper) | β | β | Yes β multi-robot teleop | |||||
| DYNA-2 (Dyna Β· gripper-primary; claims 5-finger, unverified) | β | β | β | Yes β 13 minβhours | ||||
| AgiBot / Zhiyuan (GO-1 Β· modular gripper / opt. 6-DoF dex, 2503.06669) | β | β | Yes β > 1M free-form teleop (4,000 mΒ² facility) |
Fact-check: Genesis uses no simulation in training β glove+video only (>200k h); its Genesis-World sim is eval-only (L5 = β³). Gripper-based programs (Ο, Dyna, AgiBot) are included but marked π€ β grippers/2-finger, not 5-finger dexterity; Dyna's DYNA-2 claims a 5-finger extension (unverified).
- π Genesis AI β GENE-26.5 (May 2026, Paris + San Carlos; Zhou Xian; vendor claims). A glove-first data engine: a tactile-e-skin glove giving a 1:1:1 mapping glove β human hand β robot hand (claimed 100Γ cheaper, 5Γ more data-efficient than teleop). Pretrained on real human data only β glove demos + egocentric + third-person/internet video (>200,000 h); no simulation in training (Genesis-World sim is evaluation-only, "zero simulation training data"). Then most tasks need < 1 h of task-specific robot data (< 200 episodes) for fine-tuning.
- π GR-Dexter (ByteDance Seed, Jan 2026). Explicit three-source pyramid β web VL + cross-embodiment real-robot (Fourier ActionNet ~140 h, OpenLoong 100k+, RoboMIND 107k) + 800 h human trajectories. ~20 h/task bimanual teleop (Meta Quest + MANUS) is used to train both the VLM backbone and the action DiT (21-DoF ByteDexter V2). β teleop = training, not eval.
- EgoScale (NVIDIA/Berkeley/UMD). 20,854 h human-video pretrain + a mid-training stage on aligned human-robot play (robot data used for training; +54% over no-pretrain; 22-DoF hand).
- π Dexterous Point Policy (KAIST, Jun 2026). The counter-example: trains only on human-video keypoints, no robot data; the robot appears only at eval (75% vs a VLA baseline's 1%) β the "zero-teleop-training" frontier.
Verdict on "is teleop used for fine-tuning?" β Yes, in every SOTA foundation stack, but as a small fine-tuning tip: Genesis < 1 h Β· DYNA-2 minutesβhours Β· GR-Dexter ~20 h/task Β· RLDX-1 small Β· EgoScale human-robot play Β· T-Rex 100 h. Only device-free-video and pure capture-interface methods keep robot data out of training (eval/deploy only). The 2026 achievement is collapsing the teleop fine-tuning tip toward ~1 hour, not eliminating it.
Incumbents (Figure, Tesla, AgiBot, Unitree, Fourier) are teleop/fleet-first β they scale L6 by industrializing collection (Figure's fleet mmWave offload, AgiBot's 4,000 mΒ² facility, Tesla's factory floor); newcomers (Genesis, Dyna, RLWRLD) are video/glove-first β they shrink L6 to a sub-hour tip. Both converge on the same goal β make on-hand data cheaper β one by industrializing teleop, the other by replacing it with human video/glove. Note the geography: Chinese makers dominate the open large-scale teleop datasets (AgiBot World >1M, Unitree open full-body), while US labs (Figure, Tesla) keep fleet data proprietary. (Gripper/2-finger takeaway: UMI-style handheld grippers and parallel-jaw generalists β Ο, Dyna, AgiBot β win on data-collection speed but don't test in-hand five-finger dexterity.)
- L4 retargeting is the bottleneck. The abundant L1βL3 data only pays off if humanβrobot-hand mapping is faithful. 2026's move is dynamics/force-aware and calibration-free retargeting: DexMachina (decaying virtual-controller curriculum), Functional Force-Aware Retargeting, DO AS I DO (4D hand-object dynamics, 25%β71β81%), AnyDexRT (few-shot, calibration-free). Kinematic-only retargeting drops force and fails on contact-rich 5-finger tasks.
- Tactile/force is lost going down. L1βL2 have almost none, so contact-rich skills need force re-injected at L3 (glove contact), L5 (sim force), or L6 (real) β via DECO (plug-in tactile adapter), EgoTactile (pressure from video), T-Rex (tactile-reactive), and cross-sensor CTSRL.
The dominant production recipe for a humanoid 5-finger hand today is:
Teleoperation (MANUS glove / VR / exoskeleton) β retargeting β on-hand demos (L3 + L4 + L6), pretrained/augmented with simulation + synthetic grasps (L5).
- The industry splits two ways (Β§3c): hardware incumbents (Figure fleet-offload, Tesla factory data, AgiBot's 4,000 mΒ² facility + >1M-trajectory AgiBot World, Unitree/Fourier open teleop datasets) are teleop/fleet-first; software newcomers (Genesis glove-1:1:1, Dyna/RLWRLD human-video) shrink L6 to a sub-hour tip. Chinese makers lead the open large-scale teleop datasets; US labs keep fleet data proprietary.
- ICRA-2026 vendors (MANUS, TESOLLO, ByteDexter 20-DoF) center glove-based dexterous teleop + software retargeting.
- Shared autonomy is the emerging cost-cutter: a human drives the arm (macro) via VR while an autonomous hand-VLA handles fine finger control (micro) β 2511.00139 β reducing operator burden that pure teleop can't scale past.
- Sim remains the safe, force-labeled bulk (Dex1B, DexGrasp-Zero, sim-to-real RL like DexNDM/RobustDexGrasp).
- Teleop's role is now a fine-tuning tip, not the whole dataset (Β§3b). The SOTA stacks pretrain on video/glove/sim, then fine-tune on a small on-hand teleop set β Genesis GENE-26.5 < 1 h, GR-Dexter ~20 h/task, DYNA-2 minutesβhours, T-Rex 100 h. Glove-first engines (Genesis's 1:1:1 tactile glove) push the "teleop replacement" further while still keeping a sub-hour fine-tune.
In short: teleop is still the quality anchor β but shrunk to a β€1-hour fine-tuning tip; sim is the safe bulk; human video / glove is the new frontier base.
- The pyramid inverts β human-video scaling laws become the base. EgoScale shows dexterous performance scales log-linearly (RΒ²=0.9983) with human-video hours; DYNA-2 pushes to ~1M h; UniDex spans 8 hands. Teleop retreats to the scarce high-fidelity tip.
- Device-free RGB video drops the glove: VideoManip / DO AS I DO reconstruct 4D hand-object dynamics from ordinary video β the cheapest possible L1βL4 path.
- Calibration-free, operator-agnostic retargeting makes L4 universal: AnyDexRT (few-shot), DexEXO / YUBI (wearability-first / no-retarget interfaces). If L4 becomes cheap and faithful, the whole base transfers.
- Tactile/force as a first-class tier, plus cross-sensor foundations (T-Rex, CTSRL, EgoTactile pressure-from-video) β recovering the signal the human-video base can't carry.
- Cross-hand / morphology co-design β one pyramid, many hands: House of Dextra (co-optimize morphology + policy), One-Hand (canonical URDF), UniDex's 8-hand suite.
- Unified foundation + a shrinking teleop tip β full-stack engines consume every tier and drive the on-hand fine-tune toward ~1 hour: Genesis GENE-26.5 (glove-first + < 1 h robot FT), GR-Dexter (three-source pyramid, ~20 h/task), RLDX-1, plus the WAM/MoT convergence (DYNA-2, Motus, Being-H0) trained across L1βL6 with one checkpoint; see VLA Hybrid Architectures.
- The "zero-teleop-training" frontier β device-free-video methods that keep robot data out of training entirely (Dexterous Point Policy 2606.10614-style keypoints, DO AS I DO) β robot used only for eval. Whether these can match the sub-hour-fine-tune stacks is the open contest.
The through-line: whoever makes retargeting (L4) faithful-and-cheap and tactile (cross-cut) scalable will convert the abundant human-video/glove base into real 5-finger skill with a β€1-hour teleop tip β those two bridges (plus the shrinking tip), not raw data volume, are the 2026β27 battleground.
- Companion reviews: Dexterous Manipulation Β· Human Video β Robot Transfer Β· Tactile VLA Β· Cross-Embodiment Β· Single-Checkpoint Multi-Robot
- Key data papers: EgoScale Β· UniDex Β· EgoDex Β· DexUMI Β· DexMachina Β· DexGrasp-Zero Β· DexNDM Β· Being-H0
- Retargeting deep-dives: DO AS I DO (2606.19333) Β· AnyDexRT (2607.08341)
- Capture-interface deep-dives: YUBI (2606.10244) Β· DexEXO (2603.17323)
- Tactile / industry deep-dives: T-Rex (2606.17055) Β· RLDX-1 (RLWRLD)
- Industry deep-dive: Genesis GENE-26.5 (Genesis AI, vendor)
- Humanoid-company data engines (Β§3c): Figure 03 Β· Tesla Optimus Gen-3 hand Β· AgiBot World (2503.06669) Β· Unitree Humanoid Everyday (2510.08807)
- Recent additions (no page yet): GR-Dexter (ByteDance, 2512.24210) Β· Dexterous Point Policy (KAIST, 2606.10614) Β· Shared-Autonomy (2511.00139)