CVPR 2026 EgoVLA - Heungwoo/research GitHub Wiki

EgoVLA โ€” Learning VLA Models from Egocentric Human Videos

Venue: CVPR 2026 Category: Egocentric VLA pretraining Trend tag: Trend 5 Affiliations: UCSD + UIUC + MIT + NVIDIA

Approach diagram

flowchart LR
  EGO["egocentric human video"] --> POSE["wrist pose + MANO<br/>hand parameters"]
  POSE --> PRE["pretrain NVILA-2B VLA"]
  PRE --> FT["robot fine-tune"]
  FT --> RETARGET["IK + hand retargeting<br/>to humanoid"]
  RETARGET --> POL["deployable VLA"]
Loading

Problem

Human ego video is abundant; robot teleop is not. But the gap between "what a human's wrist did" and "what the robot's end-effector should do" is non-trivial: different kinematics, different gripper.

Method

  • Backbone is NVILA-2B (compact VLM), chosen for vision-language understanding at small size.
  • Pretrain on egocentric human video to predict future wrist pose + MANO hand parameters โ€” a unified human action space โ€” from images, language, and proprioception. Pretraining uses ~500K image-action pairs from HOI4D, HOT3D, HoloAssist, and TACO.
  • Fine-tune on robot demonstrations; at deployment, human wrist+hand actions are mapped to the robot via inverse kinematics + hand retargeting rather than a single shared action space.

Results

Evaluated on the authors' Isaac Humanoid Manipulation Benchmark (NVIDIA Isaac Lab; Unitree H1 humanoid with two 12-DoF Inspire dexterous hands, 12 bimanual tasks: 7 short-horizon atomic + 5 long-horizon multi-stage). On seen backgrounds EgoVLA reaches 77.78% short-horizon vs 24.87% for the ACT baseline, and 45.93% long-horizon vs 2.22% for ACT. Human-video pretraining helps most on long-horizon and fine-grained tasks and on generalization to unseen backgrounds. Ablation: greater human-data diversity consistently improves generalization, but zero-shot deployment without robot fine-tuning yields 0% success.

Significance

EgoVLA is the cleanest demonstration that pretraining on ego video transfers to robot action prediction โ€” provided the action spaces are made compatible. Sister paper to EgoScale (which scales the data) and UniDex (which scales across hands). Combined, the three define the 2026 ego-video manipulation-pretraining recipe.

Links

  • arXiv: 2507.12440
  • Project: rchalyang.github.io/EgoVLA

Related pages

โ† Back to CVPR-2026

โš ๏ธ **GitHub.com Fallback** โš ๏ธ