ICLR 2026 Human Video Pretraining - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 ยท OpenReview: nmW77spR1I ยท arXiv: 2507.15597 Authors / affiliations: Hao Luo, Yicheng Feng, Wanpeng Zhang, Sipeng Zheng, Ye Wang, Haoqi Yuan, Jiazheng Liu, Chaoyi Xu, Qin Jin, Zongqing Lu (Peking University, Renmin University of China, BeingBeyond) Category: Data / Training Trend tag: Trend 5
flowchart LR
IW[Human hand videos<br/>mocap + VR + RGB] --> TOK[Part-level motion<br/>tokenizer GRQ]
TOK --> MT[Discrete motion tokens<br/>wrist + finger]
L[Language instructions] --> PT[Physical instruction tuning<br/>1. VLA pretraining]
MT --> PT
PT --> AL[2. Physical space<br/>alignment 3D reasoning]
AL --> POST[3. Post-training adaptation<br/>to robot]
POST --> Robot[Real-robot deployment]
Existing VLAs struggle with dexterous manipulation because they rely on synthetic data with sim-to-real gaps or on teleoperated demonstrations that lack scale and diversity. Large-scale human hand video is abundant and dexterous, but prior attempts to use it mostly yielded visual features rather than transferable, fine-grained manipulation behavior.
Being-H0 treats the human hand as a "manipulator template" and learns from web/mocap/VR human videos via physical instruction tuning, a three-part paradigm: (1) large-scale VLA pretraining from human videos, (2) physical space alignment for 3D reasoning, and (3) post-training adaptation to a robot. Crucially, hand motion is discretized into motion tokens by a part-level motion tokenizer using Grouped Residual Quantization (GRQ) โ separating wrist (global pose) from finger (fine manipulation) parameters and achieving millimeter-level reconstruction accuracy. The VLA (built on InternVL3 with an InternViT-300M encoder) then autoregressively predicts these motion tokens; it is not an inverse-dynamics objective. Training uses UniHand โ 165M+ motion-instruction pairs from 440K+ trajectories over 1,100+ hours โ sampled down to UniHand-2.5M for compute reasons.
The paper reports that Being-H0 scales favorably with model and data size and yields gains in real-world dexterous robotic manipulation after post-training adaptation. (Specific benchmark task names and per-task success rates were not extractable from the reachable source excerpts โ see report.)
Validates "pretrain from human hand video" as a credible foundation-policy pipeline, with the key technical bet being discrete part-level motion tokenization (GRQ) rather than continuous action regression or an inverse-dynamics loss. Combined with EgoDex and RoboCasa365, it reflects a shift in where pretraining data comes from: expensive teleop โ egocentric + synthetic + human hand video.
- OpenReview: https://openreview.net/forum?id=nmW77spR1I
- arXiv: https://arxiv.org/abs/2507.15597
โ Back to ICLR-2026