CVPR - Heungwoo/research GitHub Wiki

CVPR — Computer Vision and Pattern Recognition

Surveys and per-paper notes for CVPR, organized by year.

Years

Year Surveys Notes
[2026](/Heungwoo/research/wiki/CVPR-2026) 🆕 [VLA & Manipulation](/Heungwoo/research/wiki/CVPR-2026-VLA-Manipulation-Survey) 4,090 accepted of 16,092 (25.4 %). ~29 VLA/manipulation papers indexed across 10 buckets. Awards not yet announced. VLA papers are now the modal CVPR manipulation contribution for the first time; humanoid loco-manipulation has grown to a sub-cluster (VIRAL, Opening-the-Sim-to-Real-Door, Humanoid-GPT, Gallant); confirmed Highlight: Action-Sketcher.
[2025](/Heungwoo/research/wiki/CVPR-2025) [VLA & Manipulation](/Heungwoo/research/wiki/CVPR-2025-VLA-Manipulation-Survey) 2,878 accepted of 13,008 (22.1%). 19 VLA/manipulation papers indexed. Top awards went to 3D reconstruction, not VLA; closest VLA distinctions: RoboSpatial (Oral), OmniManip / DexGrasp Anything (Highlights).

Scope

CVPR is the vision-venue side of the VLA landscape. Its distinctive contribution vs. CoRL / NeurIPS / ICLR is a perception-first flavor — 3D VLAs, spatial-reasoning VLMs, visual chain-of-thought, grounding-mask policies, human-video→robot pipelines. Action-decoder innovation is lighter than at ML venues; scene understanding is heavier.

← Back to Home