RSS 2026 Human2Robot Emergence - Heungwoo/research GitHub Wiki

Emergence of Human to Robot Transfer in Vision-Language-Action Models

Venue: RSS 2026 (Imitation Learning session) Β· Authors: Simar Kareer, Karl Pertsch, James Darpinian, Judy Hoffman, Danfei Xu, Sergey Levine, Chelsea Finn, Suraj Nair β€” Physical Intelligence Γ— Georgia Tech Β· arXiv: 2512.22414 Category: Human-video co-training for VLAs Trend tag: RSS 2026 thread 2 β€” human data & cross-embodiment transfer

Compiled from the verified RSS 2026 abstract.

Key figure

Emergence result (from arXiv 2512.22414, Β© Physical Intelligence)

The paper's core result chart: absolute score improvement from adding human data, as a function of robot pre-training diversity (0% β†’ 100% β†’ 100%+cross-embodiment), across four evaluation tasks (bussing, spice, dresser, eggs). At 0–25% diversity human data contributes essentially nothing; gains appear at 50%, grow through 75–100%, and are largest with cross-embodiment data added (aggregate β‰ˆ+0.25, with individual tasks approaching +0.4). This is the emergence claim in one picture: the same human data is useless below a diversity threshold and transformative above it.

Problem

Human videos are the obvious cheap, diverse data source for VLAs β€” but training on them requires a human↔robot mapping that has historically demanded manual engineering (retargeting, unified action spaces, domain adaptation). Does it have to?

Method

  • A deliberately simple co-training recipe: mix human video into VLA pre-training without hand-crafted cross-embodiment machinery.
  • The central question is framed as an emergence hypothesis, by analogy to LLMs: does the ability to absorb heterogeneous supervision appear once pre-training is scaled?
  • Analysis probes why: diverse pre-training produces embodiment-agnostic representations under which human and robot data become mutually legible.

Results (as reported)

  • Human-to-robot transfer emerges once the VLA is pre-trained on sufficient scenes, tasks, and embodiments β€” below that diversity threshold, the same recipe fails.
  • With sufficiently diverse robot pre-training, the method nearly doubles performance on generalization settings seen only in human data.

Significance

The first-author lineage (Kareer: EgoMimic; Pertsch/Levine/Finn/Nair: the Ο€ series) makes this effectively Physical Intelligence's position paper on human data β€” and its answer is the opposite pole from Ξ¨β‚€ at the same conference: PI says co-train and let transfer emerge with diversity; Ξ¨β‚€ says decouple because the action distributions are incompatible. The synthesis question β€” whether emergence-at-scale eventually dominates staged decoupling, or only holds for gripper-class embodiments (PI's fleet) vs 36-DoF humanoids (Ξ¨β‚€'s) β€” is the sharpest open question RSS 2026 leaves behind. Relates to EgoScale and the LBM co-training study, which independently finds human video helps as a co-training modality.

← RSS 2026 survey Β· Home