RSS 2026 EgoVerse - Heungwoo/research GitHub Wiki
EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World
Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Datasets and Benchmarks · paper #92 Authors: Ryan Punamiya, Simar Kareer, Zeyi Liu, Joshua Citron, Ri-Zhao Qiu, Xiongyi Cai, Alexey Gavryushin, Jiaqi Chen, Davide Liconti, Lawrence Y. Zhu, Patcharapong Aphiwetsa, Baoyu Li, Aniketh Cheluva, Pranav Kuppili, Yangcen Liu, Dhruv Patel, Aidan Gao, Ryan Co, Hye-Young Chung, Renee ...more> arXiv: 2604.07607 · program page
Summary compiled from the arXiv paper (v2); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Three-panel overview: data collection via Project Aria glasses (academic labs), custom industry hardware, and a community phone app feeding the EgoDB management system; the unified dataset (79,692 episodes, 1,362 h, 240 scenes, 1,965 tasks) with 21-keypoint hand pose, 6-DoF ego camera pose, and task-description annotations; and evaluation on three distinct robot embodiments across labs.
Problem
Robot demonstration data is expensive and slow to scale, while egocentric human data is abundant but fragmented: existing human datasets are one-off static releases with unresolved questions about the embodiment gap and scaling behavior of human-to-robot transfer.
Method
EgoVerse is a consortium platform (Georgia Tech, Stanford, UCSD, ETH Zürich, MIT, Meta Reality Labs, Mecka AI, Scale AI) unifying collection, processing, and access. It has two parts: EgoVerse-A (controlled protocols mirrored across labs, Project Aria glasses) and EgoVerse-I (in-the-wild industry data), aggregated by EgoDB, a "living dataset" ingestion system that also supports phone-based capture. The release totals 1,362 hours / 80k episodes over 1,965 tasks, 240 scenes, and 2,087 demonstrators, with 3D hand poses, camera motion, and subtask descriptions. The transfer study co-trains an encoder-decoder policy (ResNet-18 stems, shared transformer encoder, flow-matching action decoder, quantile-normalized camera-frame actions) on human+robot data, replicated on three platforms: dual ARX5 arms (two mounting variants) and a Unitree G1 with Inspire dexterous hands.
Results
Co-training with EgoVerse-A improves in-domain and out-of-domain robot performance by up to 30%, replicated across labs and embodiments (one exception: bag-grocery on Robot B, attributed to embodiment-forced strategy mismatch). Scaling diverse human data helps only when domain-aligned human data anchors training — e.g., 2 h of aligned data unlocks transfer from 2-8 h of diverse data. In controlled diversity studies (fold-clothes, offline Avg-MSE), demonstrator diversity improves robustness to unseen humans while scene diversity dominates generalization to novel environments under limited budgets.
Significance
The most comprehensive multi-lab, multi-embodiment study yet of human-data co-training, and infrastructure intended to outlive any single release; its "aligned data anchors scaling" finding is a concrete guide for anyone budgeting human-video collection. Related wiki thread: Review-VLA-Evaluation.
← Back to RSS 2026 survey · RSS-2026-Papers · Home