CVPR 2026 EgoScale - Heungwoo/research GitHub Wiki
Venue: CVPR 2026 Category: Egocentric data scaling / Dexterous Trend tag: Trend 5 Affiliations: NVIDIA (GEAR), UC Berkeley, University of Maryland
flowchart LR
EGO["20kh egocentric video"] --> LABEL["action-label via<br/>hand-pose extractor"]
LABEL --> EGOSCALE["EgoScale dataset"]
EGOSCALE --> MID["lightweight aligned<br/>human-robot mid-training"]
ROBOT["robot data"] --> MID
MID --> VLA["flow-based VLA<br/>(VLM backbone + DiT action expert,<br/>arch similar to GR00T N1)"]
VLA --> SCALE_LAW["log-linear scaling law<br/>(R²=0.9983)<br/>ego-data volume vs. val loss"]
How much ego video do you actually need before robot performance improves? Prior work (a few hundred to a few thousand hours) was inconclusive on the scaling curve. EgoScale builds the dataset and runs the experiment.
- Curate 20,854 hours of action-labeled egocentric human video (>20× prior human→robot transfer efforts), labeled with wrist motion and retargeted dexterous hand actions.
- Pretrain a flow-based VLA policy (VLM backbone + DiT action expert, architecture similar to GR00T N1, with embodiment-conditioned MLP adapters) using a flow-matching objective.
- Two-stage transfer recipe: large-scale human pretraining → lightweight aligned human-robot mid-training → lightweight robot fine-tuning.
- Fit a log-linear scaling law (R²=0.9983) relating ego-data volume to wrist/hand action validation loss, which in turn predicts real-robot success.
Near-perfect log-linear scaling (R²=0.9983) of wrist/hand action validation loss with ego-video volume — a quantitative scaling law for ego→robot pretraining at this scale, with validation loss predictive of real-robot success. The final policy improves average success rate by 54% over a no-pretraining baseline on a 22-DoF dexterous robotic hand, and transfers effectively to lower-DoF hands. Demonstrated long-horizon, one-shot-adaptable tasks include folding clothes, separating cards, and picking up fruit with tongs.
EgoScale's scaling law is the strongest empirical case that the next 10× of robot data is ego video. The ~20.9 kh dataset is the resource; the log-linear law (R²=0.9983) is the proof; the 54% real-robot success-rate gain on a 22-DoF hand is the validation. Alongside UniDex and EgoVLA (which EgoScale cites as concurrent but smaller-scale human-data work), CVPR 2026 hosts the trio that establishes ego-video pretraining as a dominant 2026 data strategy.
- arXiv: 2602.16710
- Project: NVIDIA GEAR EgoScale page
- GR00T series review — VLA architecture is similar to GR00T N1 (flow-based, VLM backbone + DiT action expert)
- Cross-Embodiment review · Dexterous Manipulation review
- EgoDex · Human-Video Pretraining · UniDex · EgoVLA
- CVPR 2026 survey
← Back to CVPR-2026