CVPR 2026 Open Sim to Real - Heungwoo/research GitHub Wiki

Opening the Sim-to-Real Door for Humanoid Pixel-to-Action Policy Transfer

Venue: CVPR 2026 Category: Humanoid Loco-Manipulation / Sim-to-Real Trend tag: Trend 2 Affiliations: NVIDIA + UC Berkeley + CMU + CUHK

Approach diagram

flowchart LR
  SIM["sim doors<br/>(articulated)"] --> TEACH["privileged teacher<br/>PPO + staged-reset"]
  TEACH --> STUDENT["vision student<br/>DAgger distillation"]
  STUDENT --> RL["GRPO fine-tune<br/>(partial-obs)"]
  RL --> POL["pixel-to-action policy<br/>pure RGB"]
  POL --> HUMANOID["humanoid robot"]
  HUMANOID --> RES["31.7% faster than<br/>human teleop"]
Loading

Problem

Humanoid door opening โ€” an articulated-object loco-manipulation task used as a representative high-difficulty benchmark โ€” from pure RGB pixels is one of the hardest sim-to-real problems. RL exploration in sim is sparse; conventional resets destabilize long-horizon privileged-policy training, and partial observability degrades the vision student at deployment.

Method

A three-stage teacher-student-bootstrap framework:

  1. Privileged teacher (PPO) trained in sim with staged-reset exploration โ€” a curriculum-style reset distribution that stabilizes long-horizon privileged-policy training by gradually expanding initial-state difficulty so the agent sees rare failure-recovery situations.
  2. Vision student (DAgger distillation) โ€” distill the privileged teacher onto a pure-RGB student policy.
  3. GRPO fine-tune โ€” Group Relative Policy Optimization (the algorithm popularized in LLM RL contexts) applied to the student to mitigate partial observability and improve closed-loop consistency in sim-to-real RL.

The result is a pure RGB โ†’ action policy on a humanoid that achieves robust zero-shot transfer across diverse door types in the real world.

Results

Zero-shot real-world door opening, sim-trained only. The policy completes door opening up to 31.7% faster in task-completion time than human teleoperators (up to ~7.15 s faster), with robust zero-shot performance across diverse door types. The GRPO fine-tune contributes roughly a 20-30% success-rate improvement over the distilled student.

Significance

Pure-pixel humanoid manipulation that outpaces human teleop on completion time is a striking result. Suggests that for articulated tasks like door opening, a sim-trained visual policy can exceed human teleop in speed/efficiency when trained with enough diversity (note: the 31.7% is a task-time gain, not a success-rate gain over humans). Companion to VIRAL (same labs, infrastructure-focused). Closest sibling: Visual Imitation โ†’ Humanoid (CoRL 2025 best student paper).

Links

Related pages

โ† Back to CVPR-2026

โš ๏ธ **GitHub.com Fallback** โš ๏ธ