CVPR 2025 VLA Manipulation Survey - Heungwoo/research GitHub Wiki
CVPR 2025 β VLA & Manipulation Survey
Compiled April 2026 (in retrospect). Focus: CVPR 2025 VLA / manipulation papers, with a CVPR 2025 β CoRL 2025 / NeurIPS 2025 / ICLR 2026 lineage map.
TL;DR
CVPR 2025 (Nashville, June 10β17) accepted 2,878 of 13,008 submissions (22.1%). No VLA / manipulation paper won a top award β the Best Paper was VGGT (3D reconstruction), Best Student Paper was Neural Inverse Rendering. Closest VLA distinctions: RoboSpatial (Oral), OmniManip (Highlight), DexGrasp Anything (Highlight).
CVPR 2025's contribution to VLA is perception-first, distinct from CoRL 2025 (action structure) and ICLR 2026 (decoders). Five distinguishing trends:
3D is CVPR's signature contribution. 3D-MVP, Lift3D, G3Flow, RoboSpatial, VidBot all work in the 3D/point-cloud regime. CoRL 2025 and ICLR 2026 are more 2D-token dominated.
Visual chain-of-thought rather than language CoT. CoT-VLA's subgoal-image CoT, Magma's Trace-of-Mark, OmniManip's rendered primitives β CVPR's CoT is pixel-side, not text-side.
Human video as manipulation data + digital-twin sim. VidBot, ManipTrans, GenManip push sim-asset + human-video scaling rather than real-robot teleop.
No flagship generalist baseline. CVPR 2025 has no Ο0.5-class headline release. The closest are Magma (generalist multimodal agent) and UniAct (cross-embodiment universal action space). The production-baseline lane was claimed by CoRL / arXiv releases (Ο0.5, GR00T N1.x, Gemini Robotics).
flowchart LR
O[Obs + instruction] --> V[VILA-U 7B]
V -- AR --> SG[Subgoal image<br/>generated visually]
SG --> V2[VLA continues]
V2 --> A[Action chunk]
Loading
CoT-VLA (Zhao et al., NVIDIA + Stanford + MIT, CVPR 2025, 2503.22020) is the highest-impact CVPR 2025 VLA for this wiki's purposes. A 7B VLA built on VILA-U autoregressively generates a subgoal image before emitting an action chunk β the first visual CoT as an explicit generated frame inside a VLA. +17% real / +6% sim over prior SOTA. Its multimodal-CoT idea is directly generalized in ICLR 2026 dVLA (Diffusion VLA with Multimodal Chain-of-Thought) and adjacent to NeurIPS 2025 DreamVLA and ThinkAct.