RSS 2026 Human2Robot Emergence - Heungwoo/research GitHub Wiki
Emergence of Human to Robot Transfer in Vision-Language-Action Models
Venue: RSS 2026 (Imitation Learning session) Β· Authors: Simar Kareer, Karl Pertsch, James Darpinian, Judy Hoffman, Danfei Xu, Sergey Levine, Chelsea Finn, Suraj Nair β Physical Intelligence Γ Georgia Tech Β· arXiv: 2512.22414 Category: Human-video co-training for VLAs Trend tag: RSS 2026 thread 2 β human data & cross-embodiment transfer
Compiled from the verified RSS 2026 abstract.
Key figure

The paper's core result chart: absolute score improvement from adding human data, as a function of robot pre-training diversity (0% β 100% β 100%+cross-embodiment), across four evaluation tasks (bussing, spice, dresser, eggs). At 0β25% diversity human data contributes essentially nothing; gains appear at 50%, grow through 75β100%, and are largest with cross-embodiment data added (aggregate β+0.25, with individual tasks approaching +0.4). This is the emergence claim in one picture: the same human data is useless below a diversity threshold and transformative above it.
Problem
Human videos are the obvious cheap, diverse data source for VLAs β but training on them requires a humanβrobot mapping that has historically demanded manual engineering (retargeting, unified action spaces, domain adaptation). Does it have to?
Method
- A deliberately simple co-training recipe: mix human video into VLA pre-training without hand-crafted cross-embodiment machinery.
- The central question is framed as an emergence hypothesis, by analogy to LLMs: does the ability to absorb heterogeneous supervision appear once pre-training is scaled?
- Analysis probes why: diverse pre-training produces embodiment-agnostic representations under which human and robot data become mutually legible.
Results (as reported)
- Human-to-robot transfer emerges once the VLA is pre-trained on sufficient scenes, tasks, and embodiments β below that diversity threshold, the same recipe fails.
- With sufficiently diverse robot pre-training, the method nearly doubles performance on generalization settings seen only in human data.
Significance
The first-author lineage (Kareer: EgoMimic; Pertsch/Levine/Finn/Nair: the Ο series) makes this effectively Physical Intelligence's position paper on human data β and its answer is the opposite pole from Ξ¨β at the same conference: PI says co-train and let transfer emerge with diversity; Ξ¨β says decouple because the action distributions are incompatible. The synthesis question β whether emergence-at-scale eventually dominates staged decoupling, or only holds for gripper-class embodiments (PI's fleet) vs 36-DoF humanoids (Ξ¨β's) β is the sharpest open question RSS 2026 leaves behind. Relates to EgoScale and the LBM co-training study, which independently finds human video helps as a co-training modality.
β RSS 2026 survey Β· Home