CVPR 2026 EgoScale - Heungwoo/research GitHub Wiki

EgoScale — Scaling Dexterous Manipulation from Egocentric Human Data

Venue: CVPR 2026 Category: Egocentric data scaling / Dexterous Trend tag: Trend 5 Affiliations: NVIDIA (GEAR), UC Berkeley, University of Maryland

Approach diagram

flowchart LR
  EGO["20kh egocentric video"] --> LABEL["action-label via<br/>hand-pose extractor"]
  LABEL --> EGOSCALE["EgoScale dataset"]
  EGOSCALE --> MID["lightweight aligned<br/>human-robot mid-training"]
  ROBOT["robot data"] --> MID
  MID --> VLA["flow-based VLA<br/>(VLM backbone + DiT action expert,<br/>arch similar to GR00T N1)"]
  VLA --> SCALE_LAW["log-linear scaling law<br/>(R²=0.9983)<br/>ego-data volume vs. val loss"]
Loading

Problem

How much ego video do you actually need before robot performance improves? Prior work (a few hundred to a few thousand hours) was inconclusive on the scaling curve. EgoScale builds the dataset and runs the experiment.

Method

  • Curate 20,854 hours of action-labeled egocentric human video (>20× prior human→robot transfer efforts), labeled with wrist motion and retargeted dexterous hand actions.
  • Pretrain a flow-based VLA policy (VLM backbone + DiT action expert, architecture similar to GR00T N1, with embodiment-conditioned MLP adapters) using a flow-matching objective.
  • Two-stage transfer recipe: large-scale human pretraining → lightweight aligned human-robot mid-training → lightweight robot fine-tuning.
  • Fit a log-linear scaling law (R²=0.9983) relating ego-data volume to wrist/hand action validation loss, which in turn predicts real-robot success.

Results

Near-perfect log-linear scaling (R²=0.9983) of wrist/hand action validation loss with ego-video volume — a quantitative scaling law for ego→robot pretraining at this scale, with validation loss predictive of real-robot success. The final policy improves average success rate by 54% over a no-pretraining baseline on a 22-DoF dexterous robotic hand, and transfers effectively to lower-DoF hands. Demonstrated long-horizon, one-shot-adaptable tasks include folding clothes, separating cards, and picking up fruit with tongs.

Significance

EgoScale's scaling law is the strongest empirical case that the next 10× of robot data is ego video. The ~20.9 kh dataset is the resource; the log-linear law (R²=0.9983) is the proof; the 54% real-robot success-rate gain on a 22-DoF hand is the validation. Alongside UniDex and EgoVLA (which EgoScale cites as concurrent but smaller-scale human-data work), CVPR 2026 hosts the trio that establishes ego-video pretraining as a dominant 2026 data strategy.

Links

  • arXiv: 2602.16710
  • Project: NVIDIA GEAR EgoScale page

Related pages

← Back to CVPR-2026

⚠️ **GitHub.com Fallback** ⚠️