ICLR 2026 EgoDex - Heungwoo/research GitHub Wiki
Authors: Ryan Hoque, Peide Huang, David J. Yoon, Mouli Sivapurapu, Jian Zhang (Apple) · arXiv: 2505.11709 · ICLR 2026 Category: Data Trend tag: Trends 5 + 6
flowchart LR
V[Apple Vision Pro] --> E[Egocentric video stream<br/>1080p · 30 Hz]
V --> H[Calibrated cameras + SLAM<br/>3D hand/finger + camera pose]
E --> Pair[Paired data]
H --> Pair
Pair --> D[EgoDex dataset<br/>829 h · 338K episodes · 90M frames<br/>194 tabletop tasks · shoelaces → laundry]
D --> Pre[Dexterous-policy pretraining]
Dexterous manipulation has historically been data-starved. Teleoperation with dexterous hands is slow, expensive, and produces unnatural trajectories. Human video is abundant but lacks the paired 3D hand and finger information that dexterous policies need.
Use Apple Vision Pro to collect 829 hours of 1080p/30 Hz egocentric video paired with dense, continuous 3D hand and finger tracking (multiple calibrated cameras + on-device SLAM track the pose of every joint of each hand, plus camera pose). The dataset comprises 338,000 episodes / task demonstrations and 90 million frames across 194 tabletop tasks ranging from tying shoelaces to folding laundry.
The benchmark task is hand-trajectory prediction (predict future 3D hand/keypoint trajectories from observations), not downstream robot deployment. The paper trains six imitation-learning variants — two architectures (encoder–decoder, decoder-only) × three policy heads (behavior cloning, DDPM diffusion, flow matching) — evaluated with a "best-of-K" Euclidean distance over 12 keypoints. Encoder–decoder flow matching is best, outperforming other models by up to ~34% at K=5/10. Performance degrades over longer horizons, improves with dataset size, and visual goal-conditioning cuts average distance 22% and final distance 53%.
Almost an order of magnitude more data than the next-largest dataset (90M frames vs. ~21M for Ego4D in Table 1) and the largest/most-diverse dexterous human-manipulation dataset to date. Captures natural variation (grip style, pose, speed) that teleop suppresses. A key enabler of the "dexterity enters the data + RL era" trend (Survey §4). Vision Pro collection methodology hints at a new scaling path: consumer AR/VR hardware as a dexterous-data pipeline. (Note: bridging the human→robot embodiment gap via co-training or fine-tuning on robot data is left as future work — the paper does not evaluate robot policies.)
- arXiv: https://arxiv.org/abs/2505.11709
- Apple ML Research: https://machinelearning.apple.com/research/egodex
- OpenReview: https://openreview.net/forum?id=FFxkFMU89E
← Back to ICLR-2026