CVPR 2025 - Heungwoo/research GitHub Wiki

CVPR 2025

Computer Vision and Pattern Recognition 2025 β€” Nashville, Tennessee, June 10–17, 2025.

Scale: 13,008 submissions β†’ 2,878 accepted (22.1%). ~9,375 attendees from 75 countries. 118 workshops + 25 tutorials.

Awards (none directly VLA/manipulation):

  • Best Paper: VGGT β€” Visual Geometry Grounded Transformer (Oxford/Meta AI) β€” 3D reconstruction
  • Best Student Paper: Neural Inverse Rendering from Propagating Light (Toronto/Vector/CMU)
  • Closest VLA-adjacent distinctions: RoboSpatial (Oral, ~0.74%), OmniManip (Highlight), DexGrasp Anything (Highlight)

Surveys hosted in this wiki

  • VLA & Manipulation Survey (CVPR 2025) β€” 19 CVPR 2025 VLA/manipulation papers in 7 buckets with a CVPR 2025 β†’ CoRL 2025 / NeurIPS 2025 / ICLR 2026 lineage map. Notable bridge: CoT-VLA (CVPR 2025) β†’ dVLA (ICLR 2026) β€” visual chain-of-thought matures into unified discrete-diffusion decoding.

Quick paper index (by category)

VLA architecture & backbones

  • Magma (Microsoft) β€” unified agent foundation (GUI + robot) via Set-of-Mark / Trace-of-Mark pretraining Β· 2502.13130
  • UniAct (Tsinghua AIR/SenseTime) β€” VQ universal action codebook; 0.5B beats 14Γ— larger models Β· 2501.10105
  • MoManipVLA (BIT) β€” lifts fixed-base VLAs to mobile manipulation Β· 2503.13446
  • RoboGround (Shanghai AI Lab) β€” grounding masks as VLMβ†’policy interface Β· 2504.21530
  • CrayonRobo (Tencent/THU) β€” sketch-style visual prompts drive SE(3) policies Β· 2505.02166

Reasoning / CoT / planning VLAs

  • CoT-VLA (NVIDIA/Stanford/MIT) β€” Visual chain-of-thought: autoregressive subgoal-image generation before action chunks; +17% real / +6% sim. Direct ancestor of ICLR 2026's dVLA. Β· 2503.22020
  • RoboBrain (BAAI) β€” unified MLLM emitting plans + affordances + trajectories Β· 2502.21257
  • PhysVLM (CASIA) β€” Space-Physical Reachability Map (S-P Map) input; +14% over GPT-4o on EQA-phys Β· 2503.08481
  • Code-as-Monitor β€” VLM failure detection via code constraints Β· 2412.04455

3D / point-cloud VLAs (CVPR's signature)

  • RoboSpatial (Oral) β€” 1M images + 5k 3D scans + 3M spatial annotations Β· 2411.16537
  • 3D-MVP (CMU/NVIDIA) β€” multi-view masked autoencoding on Objaverse Β· 2406.18158
  • Lift3D Policy (PKU) β€” cheaply lift DINOv2/SigLIP to point-cloud policies Β· 2411.18623
  • G3Flow (CUHK) β€” live 3D semantic flow conditions a diffusion policy Β· 2411.18369

Object-centric / affordance / grounded

  • OmniManip (Highlight) (PKU/BAAI) β€” object-centric interaction primitives + 6D pose tracker; 68.3% rigid / 61.7% articulated zero-shot Β· 2501.03841
  • AffordDP β€” affordance-guided diffusion policy Β· 2412.03142

Diffusion / flow-matching policies

  • KStar Diffuser (CUHK/Tencent) β€” bimanual diffusion with kinematic-aware graph Β· 2503.10743
  • FlowRAM (CAS) β€” region-aware Mamba flow-matching policy
  • DexGrasp Anything (Highlight) β€” diffusion grasp generator with physics guidance; 3.4M grasps / 15k objects Β· 2503.08257

Dexterous / bimanual / humanoid

  • ManipTrans (Tsinghua/BAAI) β€” two-stage MoCap β†’ dex bimanual; releases 3.3k-episode DexManipNet Β· 2503.21860
  • DexHandDiff β€” diffusion for dex hands Β· 2411.18562
  • UniGraspTransformer Β· 2412.02699
  • MobileH2R β€” human-to-robot for mobile manipulation Β· 2501.04595

Data / benchmarks / simulation

  • GenManip (Shanghai AI Lab) β€” LLM-driven scene-graph sim; 10k rigid + 100 articulated assets Β· 2506.10966
  • VidBot (ETH ZΓΌrich) β€” monocular human video β†’ 3D hand trajectories β†’ diffusion policy; +20% over SOTA zero-shot Β· 2503.07135
  • RoboSpatial β€” also data (1M images, 3M annotations)

Relevant workshops

Workshop Focus
6th Embodied AI Workshop Real-world applications; Social Mobile Manipulation + Vision-Tactile Fusion challenges
Foundation Models Meet Embodied Agents Foundations for embodied agents through the MDP lens
Generalization in Robotics Manipulation Foundation models for generalizable manipulation
3D Vision-Language Model for Robotics Manipulation Robo-3DVLMs
MEIS β€” Multi-Agent Embodied Intelligence RoboTwin Dual-Arm Challenge (64 teams / 17 bimanual tasks)
Humanoid Agents Humanoid reasoning + control; best paper: Emergent Perception + Dexterity

Trends β€” what's distinctive about CVPR 2025

  1. Perception-first VLAs. CVPR attacks what the VLM sees β€” spatial reasoning, grounding masks, visual prompts, semantic flow. Compare CoRL 2025 (action structure) and ICLR 2026 (decoders).
  2. 3D is CVPR's signature contribution. 3D-MVP, Lift3D, G3Flow, RoboSpatial, VidBot β€” the 3D VLA cluster is concentrated here.
  3. Visual chain-of-thought beats language CoT. CoT-VLA, Magma's Trace-of-Mark, OmniManip's primitives β€” CoT in pixel-space, not text.
  4. Human video + digital-twin sim. VidBot, ManipTrans, GenManip push sim-asset + human-video scaling rather than real-robot teleop.
  5. No flagship generalist. CVPR 2025 has no Ο€0.5-class release β€” that lane belongs to robotics venues.

← Back to CVPR Β· Home