Review Dexterous Hand Data Pyramid - Heungwoo/research GitHub Wiki

In-Depth Review β€” The Data Pyramid for Humanoid 5-Finger Dexterous Hands

Goal: classify the data-acquisition/generation methods 2024–2026 papers & companies propose for training policies on a humanoid five-finger dexterous hand, as a data pyramid (broad-abundant-far-from-hardware base β†’ scarce-high-fidelity-on-hardware tip). Two hand-specific axes cut across every tier: the retargeting gap (human 5-finger β†’ robot 5-finger) and force/tactile (the signal lost first as you descend). Companions: Dexterous Manipulation Β· Human Video β†’ Robot Transfer Β· Tactile VLA Β· Cross-Embodiment.


1. The pyramid

flowchart TB
  L6["L6 Β· Target-hand teleop / real demos<br/>scarce Β· gold-standard Β· on-hardware"]
  L5["L5 Β· Simulation / synthetic grasp<br/>billion-scale Β· physics-labeled (force in sim)"]
  L4["L4 Β· Retargeting (human β†’ robot 5-finger)<br/>the load-bearing bridge"]
  L3["L3 Β· Wearable glove / exoskeleton<br/>action-aligned finger kinematics + contact"]
  L2["L2 Β· Egocentric instrumented dex datasets<br/>dense hand-pose (+ some pressure)"]
  L1["L1 Β· Web / internet human-hand video<br/>unbounded Β· no action/force labels"]
  L6 --- L5 --- L4 --- L3 --- L2 --- L1
  T["cross-cut Β· TACTILE / FORCE (inject at L3/L5/L6)"] -.-> L4
Loading

Quantity grows downward; fidelity to the actual 5-finger hardware grows upward. The 2026 shift (see Β§5) is that the pyramid is inverting β€” human video (base) is becoming the primary data, with teleop reserved for the scarce tip.


2. Data types β€” description, pros, cons

Tier What it is Pros Cons (hand-specific)
L1 Web/internet human video uninstrumented RGB human hands (web, YouTube, ego) unbounded scale; rich contact & object-interaction priors; ~free no action or force labels; morphology & retargeting gap; viewpoint/occlusion noise
L2 Egocentric instrumented dex Vision-Pro / MoCap hand-pose, sometimes pressure dense, curated hand-pose labels; some tactile capture cost; still a human hand; task coverage bounded
L3 Wearable glove / exoskeleton sensorized glove/exo captures finger kinematics + contact action-aligned; natural grip diversity; captures contact; no robot needed glove→robot-hand retarget loss; some contact cues (palm/compliance) lost; device-dependent
L4 Retargeting (bridge) map human hand β†’ robot 5-finger DoF turns all human data (L1–L3) into robot-hand data; force-aware variants preserve contact fidelity is the ceiling; kinematic-only retargeting drops force; calibration burden
L5 Simulation / synthetic grasp physics-generated grasps/trajectories at billion scale unbounded; physics-labeled (force available); safe; morphology-swappable sim-to-real contact/force gap; grasp-centric (not full manipulation); asset/scene coverage
L6 Target-hand teleop / real teleop or autonomous demos on the actual hand exact embodiment; real dynamics & contact β€” gold standard most expensive; lowest volume; operator burden
cross-cut: Tactile / force fingertip/skin tactile + 6-axis F/T the signal contact-rich dex actually needs; enables force control sensor-specific; nearly absent in L1–L2; cross-sensor generalization gap

3. ⭐ Which data each approach uses β€” papers & companies (separate tables)

Marks reflect training necessity, eval-only shown separately: ● required for training Β· β—‹ auxiliary (optional in training) Β· β–³ eval / deploy only β€” not training Β· blank = not used. Last column: is target-hand teleop/robot data used for fine-tuning? β€” Yes (+ volume) Β· eval-only Β· No. πŸ– = five-/multi-finger dexterous hand Β· 🀏 = gripper / 2-finger.

3a. Research papers / methods

Paper / method L1 web-vid L2 ego L3 wearable L4 retarget L5 sim L6 target-hand Tac Teleop β†’ fine-tune?
πŸ– five-/multi-finger dexterous hand
EgoScale (NVIDIA/Berkeley/UMD, 20,854 h) ● ● ● ● Yes β€” human-robot play mid-training
UniDex (8 dex hands, 10M frames) ● ● ● ● Yes
Being-H0/H0.7 (PKU, 200k h + 15k robot) ● ● ● Yes β€” 15k h robot
T-Rex (2606.17055, tactile-reactive) ● ● Yes β€” 100 h teleop
Shared-Autonomy VR+HandVLA (2511.00139) ● β—‹ ● Yes β€” teleop collects training data
πŸ†• Dexterous Point Policy (KAIST, 2606.10614) ● β—‹ β–³ eval-only β€” 75% vs VLA 1%
DO AS I DO (Berkeley/NYU, 4D retarget) ● ● β—‹ β–³ eval-only
VideoManip / RGB-video (2602.09013) ● ● β–³ eval-only
DexUMI (sensorized glove) ● ● β–³ β—‹ No β€” glove-trained; robot = deploy
DexEXO (UCLA, wearable exo) ● β—‹ β–³ No β€” exo-trained; robot = deploy
AnyDexRT (SJTU, calib-free retarget) ● β–³ eval-only β€” downstream training untested
Dex1B / SynGrasp-1B / GraspVLA (billion synth) β—‹ ● No β€” sim-trained
DexGrasp-Zero (morphology-aligned) β—‹ ● No β€” sim-trained
DexNDM / RobustDexGrasp (sim2real RL) ● β–³ β—‹ No β€” sim-trained; real = deploy
One-Hand (cross-hand canonicalize) ● β—‹ No β€” retarget/sim; robot = eval
🀏 gripper / 2-finger (reference)
YUBI (AIST Β· 2-finger bidigital, handheld) ● β–³ No β€” handheld data = training (8,434 h)

Reading it (papers): foundation-scale methods (EgoScale/UniDex/Being-H0/T-Rex) fine-tune on a target-hand teleop tip (● L6, "Yes"); device-free-video (Dexterous Point Policy, DO AS I DO) and capture-interface (DexUMI/DexEXO) methods keep the robot for eval/deploy only (β–³).

3b. Companies & industry programs (all vendor specs/claims)

Company / program (hand) L1 web-vid L2 ego L3 wearable L4 retarget L5 sim L6 target-hand Tac Teleop β†’ fine-tune?
πŸ– five-/multi-finger dexterous-hand companies
Genesis GENE-26.5 (Genesis AI Β· 5-finger + glove) ● ● ● β–³ ● ● Yes β€” < 1 h (< 200 ep); glove+video pretrain (>200k h), zero sim training (sim = eval)
RLDX-1 (RLWRLD Β· ALLEX 5-finger) β—‹ ● ● ● ● β—‹ Yes β€” small teleop set (+ synthetic training)
GR-Dexter (ByteDance Seed Β· 21-DoF ByteDexter, 2512.24210) ● ● ● Yes β€” ~20 h/task (trains VLM + action DiT)
Figure (Figure 03 Β· Helix 02 Β· ~16-DoF) β—‹ ● ● Yes β€” 10 Gbps fleet offload β†’ fleet-wide training
Tesla (Optimus Gen 3 Β· 22-DoF, tendon) β—‹ β—‹ ● Yes β€” teleop center + mocap + factory data
Unitree (G1/H1 Β· Dex3-1 3-finger / Dex5, 2510.08807) ● Yes β€” open teleop + open full-body dataset
Fourier (GR-1/GR-2 Β· dex hand + ActionNet) ● Yes β€” ~140 h teleop dataset
Teleop tool vendors (MANUS Β· Open-TeleVision Β· TESOLLO Β· ByteDexter) ● ● ● Yes β€” teleop is the data engine
🀏 gripper-based / generalist companies
Physical Intelligence (Ο€0.5 / Ο€0.7 Β· parallel-jaw gripper) β—‹ ● Yes β€” multi-robot teleop
DYNA-2 (Dyna Β· gripper-primary; claims 5-finger, unverified) ● ● ● Yes β€” 13 min–hours
AgiBot / Zhiyuan (GO-1 Β· modular gripper / opt. 6-DoF dex, 2503.06669) ● β—‹ Yes β€” > 1M free-form teleop (4,000 mΒ² facility)

Fact-check: Genesis uses no simulation in training β€” glove+video only (>200k h); its Genesis-World sim is eval-only (L5 = β–³). Gripper-based programs (Ο€, Dyna, AgiBot) are included but marked 🀏 β€” grippers/2-finger, not 5-finger dexterity; Dyna's DYNA-2 claims a 5-finger extension (unverified).

3c. Recent additions (last ~3 months) & the teleop-for-fine-tuning verdict

  • πŸ†• Genesis AI β€” GENE-26.5 (May 2026, Paris + San Carlos; Zhou Xian; vendor claims). A glove-first data engine: a tactile-e-skin glove giving a 1:1:1 mapping glove ↔ human hand ↔ robot hand (claimed 100Γ— cheaper, 5Γ— more data-efficient than teleop). Pretrained on real human data only β€” glove demos + egocentric + third-person/internet video (>200,000 h); no simulation in training (Genesis-World sim is evaluation-only, "zero simulation training data"). Then most tasks need < 1 h of task-specific robot data (< 200 episodes) for fine-tuning.
  • πŸ†• GR-Dexter (ByteDance Seed, Jan 2026). Explicit three-source pyramid β€” web VL + cross-embodiment real-robot (Fourier ActionNet ~140 h, OpenLoong 100k+, RoboMIND 107k) + 800 h human trajectories. ~20 h/task bimanual teleop (Meta Quest + MANUS) is used to train both the VLM backbone and the action DiT (21-DoF ByteDexter V2). β†’ teleop = training, not eval.
  • EgoScale (NVIDIA/Berkeley/UMD). 20,854 h human-video pretrain + a mid-training stage on aligned human-robot play (robot data used for training; +54% over no-pretrain; 22-DoF hand).
  • πŸ†• Dexterous Point Policy (KAIST, Jun 2026). The counter-example: trains only on human-video keypoints, no robot data; the robot appears only at eval (75% vs a VLA baseline's 1%) β€” the "zero-teleop-training" frontier.

Verdict on "is teleop used for fine-tuning?" β€” Yes, in every SOTA foundation stack, but as a small fine-tuning tip: Genesis < 1 h Β· DYNA-2 minutes–hours Β· GR-Dexter ~20 h/task Β· RLDX-1 small Β· EgoScale human-robot play Β· T-Rex 100 h. Only device-free-video and pure capture-interface methods keep robot data out of training (eval/deploy only). The 2026 achievement is collapsing the teleop fine-tuning tip toward ~1 hour, not eliminating it.

3d. The industry split (companies)

Incumbents (Figure, Tesla, AgiBot, Unitree, Fourier) are teleop/fleet-first β€” they scale L6 by industrializing collection (Figure's fleet mmWave offload, AgiBot's 4,000 mΒ² facility, Tesla's factory floor); newcomers (Genesis, Dyna, RLWRLD) are video/glove-first β€” they shrink L6 to a sub-hour tip. Both converge on the same goal β€” make on-hand data cheaper β€” one by industrializing teleop, the other by replacing it with human video/glove. Note the geography: Chinese makers dominate the open large-scale teleop datasets (AgiBot World >1M, Unitree open full-body), while US labs (Figure, Tesla) keep fleet data proprietary. (Gripper/2-finger takeaway: UMI-style handheld grippers and parallel-jaw generalists β€” Ο€, Dyna, AgiBot β€” win on data-collection speed but don't test in-hand five-finger dexterity.)


4. The two load-bearing bridges (why the base doesn't reach the hand for free)

  • L4 retargeting is the bottleneck. The abundant L1–L3 data only pays off if humanβ†’robot-hand mapping is faithful. 2026's move is dynamics/force-aware and calibration-free retargeting: DexMachina (decaying virtual-controller curriculum), Functional Force-Aware Retargeting, DO AS I DO (4D hand-object dynamics, 25%β†’71–81%), AnyDexRT (few-shot, calibration-free). Kinematic-only retargeting drops force and fails on contact-rich 5-finger tasks.
  • Tactile/force is lost going down. L1–L2 have almost none, so contact-rich skills need force re-injected at L3 (glove contact), L5 (sim force), or L6 (real) β€” via DECO (plug-in tactile adapter), EgoTactile (pressure from video), T-Rex (tactile-reactive), and cross-sensor CTSRL.

5. Current most-common method (2026)

The dominant production recipe for a humanoid 5-finger hand today is:

Teleoperation (MANUS glove / VR / exoskeleton) β†’ retargeting β†’ on-hand demos (L3 + L4 + L6), pretrained/augmented with simulation + synthetic grasps (L5).

  • The industry splits two ways (Β§3c): hardware incumbents (Figure fleet-offload, Tesla factory data, AgiBot's 4,000 mΒ² facility + >1M-trajectory AgiBot World, Unitree/Fourier open teleop datasets) are teleop/fleet-first; software newcomers (Genesis glove-1:1:1, Dyna/RLWRLD human-video) shrink L6 to a sub-hour tip. Chinese makers lead the open large-scale teleop datasets; US labs keep fleet data proprietary.
  • ICRA-2026 vendors (MANUS, TESOLLO, ByteDexter 20-DoF) center glove-based dexterous teleop + software retargeting.
  • Shared autonomy is the emerging cost-cutter: a human drives the arm (macro) via VR while an autonomous hand-VLA handles fine finger control (micro) β€” 2511.00139 β€” reducing operator burden that pure teleop can't scale past.
  • Sim remains the safe, force-labeled bulk (Dex1B, DexGrasp-Zero, sim-to-real RL like DexNDM/RobustDexGrasp).
  • Teleop's role is now a fine-tuning tip, not the whole dataset (Β§3b). The SOTA stacks pretrain on video/glove/sim, then fine-tune on a small on-hand teleop set β€” Genesis GENE-26.5 < 1 h, GR-Dexter ~20 h/task, DYNA-2 minutes–hours, T-Rex 100 h. Glove-first engines (Genesis's 1:1:1 tactile glove) push the "teleop replacement" further while still keeping a sub-hour fine-tune.

In short: teleop is still the quality anchor β€” but shrunk to a ≀1-hour fine-tuning tip; sim is the safe bulk; human video / glove is the new frontier base.


6. Future directions (insight)

  1. The pyramid inverts β€” human-video scaling laws become the base. EgoScale shows dexterous performance scales log-linearly (RΒ²=0.9983) with human-video hours; DYNA-2 pushes to ~1M h; UniDex spans 8 hands. Teleop retreats to the scarce high-fidelity tip.
  2. Device-free RGB video drops the glove: VideoManip / DO AS I DO reconstruct 4D hand-object dynamics from ordinary video β€” the cheapest possible L1β†’L4 path.
  3. Calibration-free, operator-agnostic retargeting makes L4 universal: AnyDexRT (few-shot), DexEXO / YUBI (wearability-first / no-retarget interfaces). If L4 becomes cheap and faithful, the whole base transfers.
  4. Tactile/force as a first-class tier, plus cross-sensor foundations (T-Rex, CTSRL, EgoTactile pressure-from-video) β€” recovering the signal the human-video base can't carry.
  5. Cross-hand / morphology co-design β€” one pyramid, many hands: House of Dextra (co-optimize morphology + policy), One-Hand (canonical URDF), UniDex's 8-hand suite.
  6. Unified foundation + a shrinking teleop tip β€” full-stack engines consume every tier and drive the on-hand fine-tune toward ~1 hour: Genesis GENE-26.5 (glove-first + < 1 h robot FT), GR-Dexter (three-source pyramid, ~20 h/task), RLDX-1, plus the WAM/MoT convergence (DYNA-2, Motus, Being-H0) trained across L1β†’L6 with one checkpoint; see VLA Hybrid Architectures.
  7. The "zero-teleop-training" frontier β€” device-free-video methods that keep robot data out of training entirely (Dexterous Point Policy 2606.10614-style keypoints, DO AS I DO) β€” robot used only for eval. Whether these can match the sub-hour-fine-tune stacks is the open contest.

The through-line: whoever makes retargeting (L4) faithful-and-cheap and tactile (cross-cut) scalable will convert the abundant human-video/glove base into real 5-finger skill with a ≀1-hour teleop tip β€” those two bridges (plus the shrinking tip), not raw data volume, are the 2026–27 battleground.


7. Links

← Back to Reviews Β· Home

⚠️ **GitHub.com Fallback** ⚠️