CVPR 2026 - Heungwoo/research GitHub Wiki
CVPR 2026
Computer Vision and Pattern Recognition 2026 β Denver Convention Center, June 3β7, 2026.
Scale: 16,092 submissions β 4,090 accepted (25.4 %) β a +42 % absolute jump in accepted-paper volume over CVPR 2025. Embodied-AI / robotics share of accepted papers grew from ~2.9 % (CVPR 2025) to ~6.2 %.
Awards: not yet announced at the time of writing (revealed on-site). Oral / Highlight tiers not yet posted by the CVF. Among the manipulation papers in this wiki, the single confirmed Highlight is Action-Sketcher (PKU + BAAI).
Surveys hosted in this wiki
- VLA & Manipulation Survey (CVPR 2026) β ~29 CVPR 2026 VLA / manipulation papers in 10 buckets, with a CVPR 2025 β CoRL 2025 β NeurIPS 2025 β ICLR 2026 β CVPR 2026 lineage map. Headline: VLA papers are now the modal CVPR manipulation contribution for the first time, and humanoid loco-manipulation has grown from a niche to a sub-cluster (VIRAL, Opening the Sim-to-Real Door, Humanoid-GPT, Gallant).
Quick paper index (by category)
VLA architecture & backbones
- OptimusVLA (HIT-Shenzhen + Huawei) β dual-memory VLA, flow-matching head
- Boost-FAN β Feasible Action Neighborhood prior bridges SFT/RFT
- Fast-ThinkAct (NVIDIA Research Taiwan) β latent CoT distillation, β89.3 % inference latency
- DM0 β Embodied-Native VLA with Embodied Spatial Scaffolding
- RDT2 (THU-ML) β 7 B + RVQ on 10 k+ h UMI demos
- DiT4DiT β cascaded video-DiT + action-DiT, 98.6 % LIBERO
Reasoning / CoT / self-correction
- ACoT-VLA (AgiBot + BUAA) β action chain-of-thought (action-space CoT)
- Action-Sketcher (Highlight) β SeeβThinkβSketchβAct, sketches as reasoning
- CycleVLA (Oxford + Cambridge) β progress-aware + MBR decoding for self-correction
World-model / RL-augmented
- GigaBrain-0.5M β RAMP world-model RL, +~30 % over RECAP-style baselines
- Robo-Dopamine (FlagOpen + BAAI + PKU) β General Process Reward Model for VLA RL
- CoWVLA β Chain of World, latent-motion subgoal thinking
Egocentric & human-video
- UniDex β 50 k+ ego trajectories across 8 dex hands, FAAS + 3D VLA
- EgoVLA (UCSD + UIUC + MIT + NVIDIA) β egocentric VLA, wrist + hand-pose unified action space
- EgoScale (NVIDIA GEAR) β 20 kh action-labeled ego video + log-linear scaling law (the GR00T N1.7 data delta)
- FunREC (ETH + MPI) β functional 3D digital twins from egocentric RGB-D
Humanoid / loco-manipulation
- VIRAL (NVIDIA + Berkeley + CMU + UT Austin + CUHK) β visual sim-to-real for humanoid loco-manip
- Open-Sim-to-Real β staged-reset + GRPO, pixel-RGB humanoid policy (+31.7 % over teleop)
- Humanoid-GPT (Tsinghua + PKU EPIC + Galbot) β causal-attention GPT on 2 B-frame mocap
- Gallant (Shanghai AI Lab) β voxel-LiDAR humanoid locomotion across 3D-constrained terrains
Affordance / grounding / active perception
- RealVLG-R1 (Tongji) β RealVLG-11B real-world grounding benchmark (~165k images, ~1.3M annotations) + R1 RFT model
- SaPaVe (PKU + BAAI) β active-perception VLA, ActiveViewPose-200K
- PanoAffordanceNet (likely) β 360Β° affordance grounding
Diffusion / flow-matching / specialty
- AnchorVLA (likely) β anchored diffusion for mobile manip
Tactile / haptic / bimanual
- HapticVLA (likely) (Skoltech) β tactile knowledge distilled into a vision-only VLA
- HandX (UIUC) β bimanual motion + interaction scaling
Mobile manipulation
- EchoVLA (likely) β declarative memory VLA for mobile manip
Benchmarks / datasets / robustness
- LIBERO-Plus (Fudan + Shanghai AI Lab) β 7-axis robustness perturbation; top VLAs 95 % β < 30 %
- RC-NF (Fudan ITEA + SMU) β robot-conditioned normalizing-flow anomaly detection, sub-100 ms
Trends β what's distinctive about CVPR 2026
- VLA papers are now the modal manipulation contribution at CVPR. Replaces the diffusion-policy + benchmark mix that dominated CVPR 2025.
- Humanoid loco-manipulation is a sub-cluster, not a niche. Four CVPR-2026 papers vs. one at CVPR 2025.
- World-model + RL for VLA crosses into CVPR for the first time. GigaBrain, Robo-Dopamine, CoWVLA all import ICLR-2026's world-model-as-policy thread.
- Reasoning VLAs split into action-CoT and image-sketch camps β both fundamentally interpretable, both abandon the language-CoT paradigm.
- Egocentric β manipulation has matured. Four papers explicitly scaling on human video, anchored by NVIDIA's EgoScale β GR00T N1.7 pipeline.
- Robustness backlash is starting. LIBERO-Plus shows top VLAs drop 65+ percentage points under modest perturbations β a credible challenge to benchmark inflation.
Relevant workshops & competitions
| Workshop / Competition | Focus |
|---|---|
| ManipArena Competition | CVPR 2026 manipulation challenge; tech-report papers published as workshop |
| GigaBrain Challenge | World-model + VLA challenge |
| ScaleBot Workshop | Scaling laws for embodied AI |
| AI4Space Workshop (incl. GUIDE) | Robotics in unstructured environments |
| Sense of Space Workshop (incl. PULSE) | Privileged knowledge / sensor transfer |
(Workshop papers are not indexed in this survey unless explicitly noted β see survey's "Validation notes".)