CVPR 2026 - Heungwoo/research GitHub Wiki

CVPR 2026

Computer Vision and Pattern Recognition 2026 β€” Denver Convention Center, June 3–7, 2026.

Scale: 16,092 submissions β†’ 4,090 accepted (25.4 %) β€” a +42 % absolute jump in accepted-paper volume over CVPR 2025. Embodied-AI / robotics share of accepted papers grew from ~2.9 % (CVPR 2025) to ~6.2 %.

Awards: not yet announced at the time of writing (revealed on-site). Oral / Highlight tiers not yet posted by the CVF. Among the manipulation papers in this wiki, the single confirmed Highlight is Action-Sketcher (PKU + BAAI).

Surveys hosted in this wiki

  • VLA & Manipulation Survey (CVPR 2026) β€” ~29 CVPR 2026 VLA / manipulation papers in 10 buckets, with a CVPR 2025 β†’ CoRL 2025 β†’ NeurIPS 2025 β†’ ICLR 2026 β†’ CVPR 2026 lineage map. Headline: VLA papers are now the modal CVPR manipulation contribution for the first time, and humanoid loco-manipulation has grown from a niche to a sub-cluster (VIRAL, Opening the Sim-to-Real Door, Humanoid-GPT, Gallant).

Quick paper index (by category)

VLA architecture & backbones

  • OptimusVLA (HIT-Shenzhen + Huawei) β€” dual-memory VLA, flow-matching head
  • Boost-FAN β€” Feasible Action Neighborhood prior bridges SFT/RFT
  • Fast-ThinkAct (NVIDIA Research Taiwan) β€” latent CoT distillation, βˆ’89.3 % inference latency
  • DM0 β€” Embodied-Native VLA with Embodied Spatial Scaffolding
  • RDT2 (THU-ML) β€” 7 B + RVQ on 10 k+ h UMI demos
  • DiT4DiT β€” cascaded video-DiT + action-DiT, 98.6 % LIBERO

Reasoning / CoT / self-correction

  • ACoT-VLA (AgiBot + BUAA) β€” action chain-of-thought (action-space CoT)
  • Action-Sketcher (Highlight) β€” See–Think–Sketch–Act, sketches as reasoning
  • CycleVLA (Oxford + Cambridge) β€” progress-aware + MBR decoding for self-correction

World-model / RL-augmented

  • GigaBrain-0.5M β€” RAMP world-model RL, +~30 % over RECAP-style baselines
  • Robo-Dopamine (FlagOpen + BAAI + PKU) β€” General Process Reward Model for VLA RL
  • CoWVLA β€” Chain of World, latent-motion subgoal thinking

Egocentric & human-video

  • UniDex β€” 50 k+ ego trajectories across 8 dex hands, FAAS + 3D VLA
  • EgoVLA (UCSD + UIUC + MIT + NVIDIA) β€” egocentric VLA, wrist + hand-pose unified action space
  • EgoScale (NVIDIA GEAR) β€” 20 kh action-labeled ego video + log-linear scaling law (the GR00T N1.7 data delta)
  • FunREC (ETH + MPI) β€” functional 3D digital twins from egocentric RGB-D

Humanoid / loco-manipulation

  • VIRAL (NVIDIA + Berkeley + CMU + UT Austin + CUHK) β€” visual sim-to-real for humanoid loco-manip
  • Open-Sim-to-Real β€” staged-reset + GRPO, pixel-RGB humanoid policy (+31.7 % over teleop)
  • Humanoid-GPT (Tsinghua + PKU EPIC + Galbot) β€” causal-attention GPT on 2 B-frame mocap
  • Gallant (Shanghai AI Lab) β€” voxel-LiDAR humanoid locomotion across 3D-constrained terrains

Affordance / grounding / active perception

  • RealVLG-R1 (Tongji) β€” RealVLG-11B real-world grounding benchmark (~165k images, ~1.3M annotations) + R1 RFT model
  • SaPaVe (PKU + BAAI) β€” active-perception VLA, ActiveViewPose-200K
  • PanoAffordanceNet (likely) β€” 360Β° affordance grounding

Diffusion / flow-matching / specialty

  • AnchorVLA (likely) β€” anchored diffusion for mobile manip

Tactile / haptic / bimanual

  • HapticVLA (likely) (Skoltech) β€” tactile knowledge distilled into a vision-only VLA
  • HandX (UIUC) β€” bimanual motion + interaction scaling

Mobile manipulation

  • EchoVLA (likely) β€” declarative memory VLA for mobile manip

Benchmarks / datasets / robustness

  • LIBERO-Plus (Fudan + Shanghai AI Lab) β€” 7-axis robustness perturbation; top VLAs 95 % β†’ < 30 %
  • RC-NF (Fudan ITEA + SMU) β€” robot-conditioned normalizing-flow anomaly detection, sub-100 ms

Trends β€” what's distinctive about CVPR 2026

  1. VLA papers are now the modal manipulation contribution at CVPR. Replaces the diffusion-policy + benchmark mix that dominated CVPR 2025.
  2. Humanoid loco-manipulation is a sub-cluster, not a niche. Four CVPR-2026 papers vs. one at CVPR 2025.
  3. World-model + RL for VLA crosses into CVPR for the first time. GigaBrain, Robo-Dopamine, CoWVLA all import ICLR-2026's world-model-as-policy thread.
  4. Reasoning VLAs split into action-CoT and image-sketch camps β€” both fundamentally interpretable, both abandon the language-CoT paradigm.
  5. Egocentric β†’ manipulation has matured. Four papers explicitly scaling on human video, anchored by NVIDIA's EgoScale β†’ GR00T N1.7 pipeline.
  6. Robustness backlash is starting. LIBERO-Plus shows top VLAs drop 65+ percentage points under modest perturbations β€” a credible challenge to benchmark inflation.

Relevant workshops & competitions

Workshop / Competition Focus
ManipArena Competition CVPR 2026 manipulation challenge; tech-report papers published as workshop
GigaBrain Challenge World-model + VLA challenge
ScaleBot Workshop Scaling laws for embodied AI
AI4Space Workshop (incl. GUIDE) Robotics in unstructured environments
Sense of Space Workshop (incl. PULSE) Privileged knowledge / sensor transfer

(Workshop papers are not indexed in this survey unless explicitly noted β€” see survey's "Validation notes".)

← Back to CVPR Β· Home