CoRL 2026 HiPHI - Heungwoo/research GitHub Wiki

CoRL 2026 — HiPHI: A Large-Scale Benchmark for High-Precision Human Motion & Object-Interaction

Venue: CoRL 2026 (Austin, TX, Nov 9–12) · Noitom Robotics. Paper: arXiv 2608.16222. Representative of: the humanoid data substrate — a 617.5 h high-precision mocap corpus that keeps improving whole-body policies with scale. Companions: Humanoid VLA · CoRL 2026 survey.

HiPHI overview: optical whole-body mocap with synchronized object trajectories and meshes, organized by FrameNet frames (figure from the authors, arXiv 2608.16222, © the authors)

1. Problem

Humanoid intelligence must learn over an extremely diverse space of whole-body motions and physically grounded interactions. Existing data sources force a trade-off: internet video is broad but lacks precise physical state, while lab motion-capture sets have accurate state but narrow behavioral coverage. HiPHI targets that gap with a large, high-fidelity corpus that systematically maximizes coverage of the human motion and interaction manifold.

2. Method

HiPHI is a 617.5-hour optical motion-capture dataset captured at 90 Hz (~200.1 M frames) from 132 performers, with sub-millimeter marker tracking. It splits into 371.8 h of whole-body movement and 245.7 h of human–object interaction, the latter covering 40 real physical objects across 12 categories with synchronized object trajectories and mesh-level 3D geometry. Coverage is organized using the FrameNet linguistic framework — 214 Frame-LU labels across 22 frames — to structure the behavioral space. A companion benchmark suite scores motion-space diversity, interaction grounding, object consistency, and downstream physical-AI applications.

3. Results

  • Coverage. Broader and more uniform kinematic coverage than prior sets (1620 vs 1438 occupied cells; 1443 vs 1114 effective occupancy) with a longer tail (14.1% vs 10.7%).
  • Quality. Low ground penetration (8 mm), 98.1% non-conflict frames, 95.7% near-surface grounding.
  • Matched-budget tracking. In physics-based imitation (DeepMimic on Unitree G1), HiPHI gives the highest success rates and fastest convergence at both 3 h and 20 h budgets; cross-dataset MPJPE keeps dropping as training data scales from 3 h to 300 h.
  • Real hardware. Policies transfer to a real Unitree G1, executing running, sitting, crawling, carrying, flipping, and pulling.

4. Why it matters

HiPHI argues that the bottleneck for whole-body humanoid policies is a data substrate — precise physical state plus systematically broad behavioral coverage — and shows the payoff is monotone with scale, so more of the same data keeps helping. The object-interaction split with meshes and trajectories makes it usable for contact-rich manipulation and loco-manipulation, not just locomotion.

Limitations (reviewer): optical mocap in a studio is expensive and constrains scene/lighting/object realism; tracking results center on the Unitree G1, so cross-embodiment generality is unverified; real-hardware behaviors are demonstrated qualitatively rather than with success-rate tables.

5. Links

← Back to CoRL 2026 survey · Home