ICLR 2026 EgoDex - Heungwoo/research GitHub Wiki

EgoDex — Learning Dexterous Manipulation from Large-Scale Egocentric Video

Authors: Ryan Hoque, Peide Huang, David J. Yoon, Mouli Sivapurapu, Jian Zhang (Apple) · arXiv: 2505.11709 · ICLR 2026 Category: Data Trend tag: Trends 5 + 6

Approach diagram

flowchart LR
  V[Apple Vision Pro] --> E[Egocentric video stream<br/>1080p · 30 Hz]
  V --> H[Calibrated cameras + SLAM<br/>3D hand/finger + camera pose]
  E --> Pair[Paired data]
  H --> Pair
  Pair --> D[EgoDex dataset<br/>829 h · 338K episodes · 90M frames<br/>194 tabletop tasks · shoelaces → laundry]
  D --> Pre[Dexterous-policy pretraining]
Loading

Problem

Dexterous manipulation has historically been data-starved. Teleoperation with dexterous hands is slow, expensive, and produces unnatural trajectories. Human video is abundant but lacks the paired 3D hand and finger information that dexterous policies need.

Method

Use Apple Vision Pro to collect 829 hours of 1080p/30 Hz egocentric video paired with dense, continuous 3D hand and finger tracking (multiple calibrated cameras + on-device SLAM track the pose of every joint of each hand, plus camera pose). The dataset comprises 338,000 episodes / task demonstrations and 90 million frames across 194 tabletop tasks ranging from tying shoelaces to folding laundry.

Results

The benchmark task is hand-trajectory prediction (predict future 3D hand/keypoint trajectories from observations), not downstream robot deployment. The paper trains six imitation-learning variants — two architectures (encoder–decoder, decoder-only) × three policy heads (behavior cloning, DDPM diffusion, flow matching) — evaluated with a "best-of-K" Euclidean distance over 12 keypoints. Encoder–decoder flow matching is best, outperforming other models by up to ~34% at K=5/10. Performance degrades over longer horizons, improves with dataset size, and visual goal-conditioning cuts average distance 22% and final distance 53%.

Significance

Almost an order of magnitude more data than the next-largest dataset (90M frames vs. ~21M for Ego4D in Table 1) and the largest/most-diverse dexterous human-manipulation dataset to date. Captures natural variation (grip style, pose, speed) that teleop suppresses. A key enabler of the "dexterity enters the data + RL era" trend (Survey §4). Vision Pro collection methodology hints at a new scaling path: consumer AR/VR hardware as a dexterous-data pipeline. (Note: bridging the human→robot embodiment gap via co-training or fine-tuning on robot data is left as future work — the paper does not evaluate robot policies.)

Links

Related pages

  • UniHM (consumer of EgoDex)
  • DexNDM (sim-to-real for dexterity)
  • RFS (residual RL for dexterity)

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️