CVPR 2025 - Heungwoo/research GitHub Wiki
CVPR 2025
Computer Vision and Pattern Recognition 2025 β Nashville, Tennessee, June 10β17, 2025.
Scale: 13,008 submissions β 2,878 accepted (22.1%). ~9,375 attendees from 75 countries. 118 workshops + 25 tutorials.
Awards (none directly VLA/manipulation):
- Best Paper: VGGT β Visual Geometry Grounded Transformer (Oxford/Meta AI) β 3D reconstruction
- Best Student Paper: Neural Inverse Rendering from Propagating Light (Toronto/Vector/CMU)
- Closest VLA-adjacent distinctions: RoboSpatial (Oral, ~0.74%), OmniManip (Highlight), DexGrasp Anything (Highlight)
Surveys hosted in this wiki
- VLA & Manipulation Survey (CVPR 2025) β 19 CVPR 2025 VLA/manipulation papers in 7 buckets with a CVPR 2025 β CoRL 2025 / NeurIPS 2025 / ICLR 2026 lineage map. Notable bridge: CoT-VLA (CVPR 2025) β dVLA (ICLR 2026) β visual chain-of-thought matures into unified discrete-diffusion decoding.
Quick paper index (by category)
VLA architecture & backbones
- Magma (Microsoft) β unified agent foundation (GUI + robot) via Set-of-Mark / Trace-of-Mark pretraining Β· 2502.13130
- UniAct (Tsinghua AIR/SenseTime) β VQ universal action codebook; 0.5B beats 14Γ larger models Β· 2501.10105
- MoManipVLA (BIT) β lifts fixed-base VLAs to mobile manipulation Β· 2503.13446
- RoboGround (Shanghai AI Lab) β grounding masks as VLMβpolicy interface Β· 2504.21530
- CrayonRobo (Tencent/THU) β sketch-style visual prompts drive SE(3) policies Β· 2505.02166
Reasoning / CoT / planning VLAs
- CoT-VLA (NVIDIA/Stanford/MIT) β Visual chain-of-thought: autoregressive subgoal-image generation before action chunks; +17% real / +6% sim. Direct ancestor of ICLR 2026's dVLA. Β· 2503.22020
- RoboBrain (BAAI) β unified MLLM emitting plans + affordances + trajectories Β· 2502.21257
- PhysVLM (CASIA) β Space-Physical Reachability Map (S-P Map) input; +14% over GPT-4o on EQA-phys Β· 2503.08481
- Code-as-Monitor β VLM failure detection via code constraints Β· 2412.04455
3D / point-cloud VLAs (CVPR's signature)
- RoboSpatial (Oral) β 1M images + 5k 3D scans + 3M spatial annotations Β· 2411.16537
- 3D-MVP (CMU/NVIDIA) β multi-view masked autoencoding on Objaverse Β· 2406.18158
- Lift3D Policy (PKU) β cheaply lift DINOv2/SigLIP to point-cloud policies Β· 2411.18623
- G3Flow (CUHK) β live 3D semantic flow conditions a diffusion policy Β· 2411.18369
Object-centric / affordance / grounded
- OmniManip (Highlight) (PKU/BAAI) β object-centric interaction primitives + 6D pose tracker; 68.3% rigid / 61.7% articulated zero-shot Β· 2501.03841
- AffordDP β affordance-guided diffusion policy Β· 2412.03142
Diffusion / flow-matching policies
- KStar Diffuser (CUHK/Tencent) β bimanual diffusion with kinematic-aware graph Β· 2503.10743
- FlowRAM (CAS) β region-aware Mamba flow-matching policy
- DexGrasp Anything (Highlight) β diffusion grasp generator with physics guidance; 3.4M grasps / 15k objects Β· 2503.08257
Dexterous / bimanual / humanoid
- ManipTrans (Tsinghua/BAAI) β two-stage MoCap β dex bimanual; releases 3.3k-episode DexManipNet Β· 2503.21860
- DexHandDiff β diffusion for dex hands Β· 2411.18562
- UniGraspTransformer Β· 2412.02699
- MobileH2R β human-to-robot for mobile manipulation Β· 2501.04595
Data / benchmarks / simulation
- GenManip (Shanghai AI Lab) β LLM-driven scene-graph sim; 10k rigid + 100 articulated assets Β· 2506.10966
- VidBot (ETH ZΓΌrich) β monocular human video β 3D hand trajectories β diffusion policy; +20% over SOTA zero-shot Β· 2503.07135
- RoboSpatial β also data (1M images, 3M annotations)
Relevant workshops
| Workshop | Focus |
|---|---|
| 6th Embodied AI Workshop | Real-world applications; Social Mobile Manipulation + Vision-Tactile Fusion challenges |
| Foundation Models Meet Embodied Agents | Foundations for embodied agents through the MDP lens |
| Generalization in Robotics Manipulation | Foundation models for generalizable manipulation |
| 3D Vision-Language Model for Robotics Manipulation | Robo-3DVLMs |
| MEIS β Multi-Agent Embodied Intelligence | RoboTwin Dual-Arm Challenge (64 teams / 17 bimanual tasks) |
| Humanoid Agents | Humanoid reasoning + control; best paper: Emergent Perception + Dexterity |
Trends β what's distinctive about CVPR 2025
- Perception-first VLAs. CVPR attacks what the VLM sees β spatial reasoning, grounding masks, visual prompts, semantic flow. Compare CoRL 2025 (action structure) and ICLR 2026 (decoders).
- 3D is CVPR's signature contribution. 3D-MVP, Lift3D, G3Flow, RoboSpatial, VidBot β the 3D VLA cluster is concentrated here.
- Visual chain-of-thought beats language CoT. CoT-VLA, Magma's Trace-of-Mark, OmniManip's primitives β CoT in pixel-space, not text.
- Human video + digital-twin sim. VidBot, ManipTrans, GenManip push sim-asset + human-video scaling rather than real-robot teleop.
- No flagship generalist. CVPR 2025 has no Ο0.5-class release β that lane belongs to robotics venues.