CVPR 2026 Open Sim to Real - Heungwoo/research GitHub Wiki
Venue: CVPR 2026 Category: Humanoid Loco-Manipulation / Sim-to-Real Trend tag: Trend 2 Affiliations: NVIDIA + UC Berkeley + CMU + CUHK
flowchart LR
SIM["sim doors<br/>(articulated)"] --> TEACH["privileged teacher<br/>PPO + staged-reset"]
TEACH --> STUDENT["vision student<br/>DAgger distillation"]
STUDENT --> RL["GRPO fine-tune<br/>(partial-obs)"]
RL --> POL["pixel-to-action policy<br/>pure RGB"]
POL --> HUMANOID["humanoid robot"]
HUMANOID --> RES["31.7% faster than<br/>human teleop"]
Humanoid door opening โ an articulated-object loco-manipulation task used as a representative high-difficulty benchmark โ from pure RGB pixels is one of the hardest sim-to-real problems. RL exploration in sim is sparse; conventional resets destabilize long-horizon privileged-policy training, and partial observability degrades the vision student at deployment.
A three-stage teacher-student-bootstrap framework:
- Privileged teacher (PPO) trained in sim with staged-reset exploration โ a curriculum-style reset distribution that stabilizes long-horizon privileged-policy training by gradually expanding initial-state difficulty so the agent sees rare failure-recovery situations.
- Vision student (DAgger distillation) โ distill the privileged teacher onto a pure-RGB student policy.
- GRPO fine-tune โ Group Relative Policy Optimization (the algorithm popularized in LLM RL contexts) applied to the student to mitigate partial observability and improve closed-loop consistency in sim-to-real RL.
The result is a pure RGB โ action policy on a humanoid that achieves robust zero-shot transfer across diverse door types in the real world.
Zero-shot real-world door opening, sim-trained only. The policy completes door opening up to 31.7% faster in task-completion time than human teleoperators (up to ~7.15 s faster), with robust zero-shot performance across diverse door types. The GRPO fine-tune contributes roughly a 20-30% success-rate improvement over the distilled student.
Pure-pixel humanoid manipulation that outpaces human teleop on completion time is a striking result. Suggests that for articulated tasks like door opening, a sim-trained visual policy can exceed human teleop in speed/efficiency when trained with enough diversity (note: the 31.7% is a task-time gain, not a success-rate gain over humans). Companion to VIRAL (same labs, infrastructure-focused). Closest sibling: Visual Imitation โ Humanoid (CoRL 2025 best student paper).
- arXiv: 2512.01061
- VIRAL (companion) ยท Visual Imitation โ Humanoid ยท WholeBodyVLA
- RL for VLA ยท Cross-Embodiment review
- CVPR 2026 survey
โ Back to CVPR-2026