CVPR 2025 VLA Manipulation Survey - Heungwoo/research GitHub Wiki

CVPR 2025 β€” VLA & Manipulation Survey

Compiled April 2026 (in retrospect). Focus: CVPR 2025 VLA / manipulation papers, with a CVPR 2025 β†’ CoRL 2025 / NeurIPS 2025 / ICLR 2026 lineage map.

TL;DR

CVPR 2025 (Nashville, June 10–17) accepted 2,878 of 13,008 submissions (22.1%). No VLA / manipulation paper won a top award β€” the Best Paper was VGGT (3D reconstruction), Best Student Paper was Neural Inverse Rendering. Closest VLA distinctions: RoboSpatial (Oral), OmniManip (Highlight), DexGrasp Anything (Highlight).

CVPR 2025's contribution to VLA is perception-first, distinct from CoRL 2025 (action structure) and ICLR 2026 (decoders). Five distinguishing trends:

  1. Perception-first VLAs β€” spatial reasoning (RoboSpatial, PhysVLM), grounding masks (RoboGround), object primitives (OmniManip), semantic flow (G3Flow), visual prompts (CrayonRobo).
  2. 3D is CVPR's signature contribution. 3D-MVP, Lift3D, G3Flow, RoboSpatial, VidBot all work in the 3D/point-cloud regime. CoRL 2025 and ICLR 2026 are more 2D-token dominated.
  3. Visual chain-of-thought rather than language CoT. CoT-VLA's subgoal-image CoT, Magma's Trace-of-Mark, OmniManip's rendered primitives β€” CVPR's CoT is pixel-side, not text-side.
  4. Human video as manipulation data + digital-twin sim. VidBot, ManipTrans, GenManip push sim-asset + human-video scaling rather than real-robot teleop.
  5. No flagship generalist baseline. CVPR 2025 has no Ο€0.5-class headline release. The closest are Magma (generalist multimodal agent) and UniAct (cross-embodiment universal action space). The production-baseline lane was claimed by CoRL / arXiv releases (Ο€0.5, GR00T N1.x, Gemini Robotics).

Strongest cross-venue seed: CoT-VLA β†’ ICLR 2026 dVLA (explicit multimodal-CoT β†’ discrete-diffusion action decoding).

Flagship: CoT-VLA

flowchart LR
  O[Obs + instruction] --> V[VILA-U 7B]
  V -- AR --> SG[Subgoal image<br/>generated visually]
  SG --> V2[VLA continues]
  V2 --> A[Action chunk]
Loading

CoT-VLA (Zhao et al., NVIDIA + Stanford + MIT, CVPR 2025, 2503.22020) is the highest-impact CVPR 2025 VLA for this wiki's purposes. A 7B VLA built on VILA-U autoregressively generates a subgoal image before emitting an action chunk β€” the first visual CoT as an explicit generated frame inside a VLA. +17% real / +6% sim over prior SOTA. Its multimodal-CoT idea is directly generalized in ICLR 2026 dVLA (Diffusion VLA with Multimodal Chain-of-Thought) and adjacent to NeurIPS 2025 DreamVLA and ThinkAct.

Section index

  1. VLA Architecture & Backbones
  2. Reasoning / CoT / Planning VLAs
  3. 3D / Point-Cloud VLAs
  4. Object-Centric / Affordance / Grounded
  5. Diffusion / Flow-Matching Policies
  6. Dexterous / Bimanual / Humanoid
  7. Data / Benchmarks / Simulation
  8. CVPR 2025 β†’ later venues lineage

1. VLA Architecture & Backbones

Paper Novelty
Magma (Microsoft) Unified agent foundation model across GUI + robot manipulation via Set-of-Mark / Trace-of-Mark pretraining on heterogeneous VL data
UniAct (Tsinghua AIR/SenseTime/PKU) VQ codebook of universal actions + per-embodiment decoder. 0.5B UniAct beats models 14Γ— larger on cross-embodiment evals
MoManipVLA (BIT) Adaptation framework to lift fixed-base pretrained VLAs to mobile-base settings; +4.2% vs prior mobile-manip SOTA
RoboGround (Shanghai AI Lab) Grounding masks as intermediate modality between VLM and low-level policy; auto-generated simulator corpus
CrayonRobo (Tencent/THU) 2D visual prompts (arrows, points overlaid on RGB) drive end-to-end SE(3) policy

2. Reasoning / CoT / Planning VLAs

Paper Novelty
CoT-VLA (NVIDIA/Stanford/MIT) Visual chain-of-thought: autoregressively generate a subgoal image before emitting an action chunk. +17% real / +6% sim
RoboBrain (BAAI) Unified MLLM producing plans + affordances + end-effector trajectories; new ShareRobot dataset
PhysVLM (CAS) Adds Space-Physical Reachability Map (S-P Map) to VLM inputs; +14% over GPT-4o on EQA-phys
Code-as-Monitor VLM failure detection via code constraints (monitoring, not policy)

3. 3D / Point-Cloud VLAs (CVPR's signature)

Paper Novelty
RoboSpatial (Oral) (OSU + NVIDIA) 1M images + 5k 3D scans + 3M spatial annotations β€” trains VLMs that lift spatial affordance + relation prediction + manipulation
3D-MVP (CMU/NVIDIA) Masked autoencoding on Objaverse pretrains a multi-view transformer for gripper-pose prediction
Lift3D Policy (PKU) Cheaply lift 2D foundation models (DINOv2, SigLIP) to point-cloud policies via task-aware MAE + positional lift
G3Flow (CUHK) Live 3D semantic flow (digital-twin + VFM + pose tracking) feeds a diffusion policy; robust to occlusion

4. Object-Centric / Affordance / Grounded

Paper Novelty
OmniManip (Highlight) (PKU/BAAI) Dual closed-loop: VLM reasons over object-centric interaction primitives; 6D pose tracker executes. 68.3% rigid / 61.7% articulated zero-shot
CrayonRobo (Tencent/THU) Sketch-style 2D visual prompts drive action generation
AffordDP Affordance-guided diffusion policy

5. Diffusion / Flow-Matching Policies

Paper Novelty
KStar Diffuser (CUHK/Tencent) Bimanual diffusion with kinematic-aware dynamic graph over both arms + differentiable kinematics regularizer
FlowRAM (CAS) Flow-matching policy with Mamba-based region-aware encoding + dynamic-radius perception schedule
DexGrasp Anything (Highlight) (ShanghaiTech) Diffusion grasp generator with physics-in-the-loop guidance (surface-pull + penetration-repulsion); 3.4M grasps / 15k objects
AffordDP Affordance-guided diffusion policy
G3Flow (also 3D) 3D semantic flow β†’ diffusion policy

6. Dexterous / Bimanual / Humanoid

Paper Novelty
ManipTrans (Tsinghua/BAAI) Two-stage (generalist trajectory imitator + residual FT) transfer from human MoCap to dex robots; releases DexManipNet (3.3k episodes)
DexGrasp Anything (Highlight) Physics-aware universal dex grasp synthesis
DexHandDiff Diffusion for dex hand manipulation
UniGraspTransformer Universal grasp transformer
MobileH2R Human-to-robot transfer for mobile manipulation
ManiVideo Dex hand-object video generation
Let Humanoid Robots Go Hiking Humanoid outdoor locomotion

7. Data / Benchmarks / Simulation

Paper Novelty
RoboSpatial (Oral) 1M images + 5k 3D scans + 3M spatial annotations
GenManip (Shanghai AI Lab) LLM-driven task-oriented scene graph sim; 10k rigid + 100 articulated assets + 200 curated scenarios
VidBot (ETH ZΓΌrich) Monocular human video β†’ depth + SfM β†’ metric 3D hand trajectories β†’ coarse-to-fine diffusion policy. +20% zero-shot across 13 tasks
RoboBrain's ShareRobot Multi-capability MLLM training dataset
ManipTrans's DexManipNet 3.3k-episode dex bimanual dataset
RoboTwin Dual-Arm Challenge (MEIS workshop) 64 teams, 17 bimanual tasks; tech report

8. Lineage

CVPR 2025 β†’ CoRL 2025 / NeurIPS 2025 / ICLR 2026

CVPR 2025 seed Descendant
CoT-VLA (subgoal-image CoT) dVLA (ICLR 2026 β€” multimodal CoT in discrete diffusion), ThinkAct (NeurIPS 2025 β€” RL reward-shaped MLLM plans), DreamVLA (NeurIPS 2025 β€” multi-modal world-knowledge forecasting)
Magma (Trace-of-Mark, unified agent) X-VLA, UniVLA, XR-1 (cross-embodiment / cross-domain generalists)
UniAct (VQ universal action space) Discrete Diffusion VLA, DIVA & Fast-dVLA, FASTER, OmniSAT (tokenized action families)
RoboBrain RoboBrain 2.0 technical report (arXiv 2507.02029, July 2025)
RoboSpatial / PhysVLM / RoboGround (spatial + grounding) VLM4VLA (VLM-benchmark analysis) + Embodied-R1 (keypoint-grounded RL)
3D-MVP / Lift3D / G3Flow (3D-pretrain) CoRL 2025 3DS-VLA; NeurIPS 2025 BridgeVLA (2D-heatmap unification)
VidBot (human-video β†’ 3D affordance) CoRL 2025 DexUMI, Visual Imitation β†’ Humanoid, X-Sim; ICLR 2026 EgoDex, Human-Video Pretraining
ManipTrans / DexGrasp Anything / DexHandDiff CoRL 2025 DexUMI, ClutterDexGrasp; ICLR 2026 DexNDM, UniHM, RFS
GenManip (LLM-driven scene graph sim) NeurIPS 2025 HumanoidGen, DexFlyWheel; ICLR 2026 RoboCasa365, WorldGym
FlowRAM / KStar Diffuser (flow/diffusion policies) CoRL 2025 DSRL, Streaming Flow Policy; NeurIPS 2025 ReinFlow, Real-Time Chunking

Validation notes (common CVPR 2025 miscitations)

  • CoT-VLA's arXiv is 2503.22020 β€” not 2403.09631 (that's 3D-VLA by Zhen et al., ICML 2024).
  • CoA-VLA (2412.20451) is ICCV 2025, not CVPR 2025.
  • SpatialVLA (2501.15830) is RSS 2025, not CVPR.
  • DexVLA (2502.05855) is CoRL 2025 β€” see DexVLA.
  • MolmoAct (2508.07917) is Ai2 arXiv-only, Aug 2025.

Sources

← Back to CVPR-2025 Β· CVPR Β· Home

⚠️ **GitHub.com Fallback** ⚠️