CVPR 2026 Gallant - Heungwoo/research GitHub Wiki
Venue: CVPR 2026 (Poster #37444) Category: Humanoid Locomotion Trend tag: Trend 2 Affiliations: Shanghai AI Lab · CUHK · USTC · University of Tokyo · SJTU
flowchart LR
LIDAR["LiDAR / depth sensor"] --> VOX["voxelize to 3D grid"]
VOX --> ZG["z-grouped 2D CNN<br/>height-aware encoding"]
ZG --> POL["humanoid locomotion policy"]
POL --> ACT["walking action"]
TERRAIN["overhead, lateral 3D constraints"] -.-> POL
Existing humanoid locomotion policies mostly assume flat or moderately uneven ground. Real environments (warehouses, construction sites, kitchens) have overhead constraints (low pipes, shelving) and lateral constraints (narrow corridors, posts). A 2D height-map representation is insufficient.
- Sense the environment with LiDAR; voxelize into a 3D occupancy grid.
- Encode the voxel grid with a z-grouped 2D CNN — each height slice is processed by a shared 2D CNN, then the height channel is treated as an additional feature dimension. Captures the full 3D structure (overhead and lateral) without the cost of a full 3D CNN.
- End-to-end policy from the 3D representation to humanoid joint actions.
A single unified policy handles diverse terrain — lateral obstacles, overhead constraints, multi-level structures, and narrow passages — overcoming prior methods restricted to ground-level hazards. Reports near-100% success rates in challenging scenarios such as stair climbing and stepping onto elevated platforms (real-world deployment, Fig. 3), with ablations on terrain-traversal success and collision impulse (Fig. 2). Trained with a high-fidelity LiDAR simulation that generates realistic observations on the fly for sim-to-real transfer.
Gallant occupies the environment-perception corner of CVPR 2026's humanoid cluster (vs. VIRAL/Open-Sim-to-Real's task-perception corner and Humanoid-GPT's motion-prior corner). Together the four papers cover the full stack for humanoid deployment.
- arXiv: 2511.14625
- CVPR Poster: #37444
← Back to CVPR-2026