CVPR 2026 VLA Manipulation Survey - Heungwoo/research GitHub Wiki
CVPR 2026 β VLA & Manipulation Survey
Compiled May 2026. Focus: CVPR 2026 VLA / manipulation papers, with a CVPR 2025 β CoRL 2025 β NeurIPS 2025 β ICLR 2026 β CVPR 2026 lineage map. CVPR 2026 runs June 3β7 at the Denver Convention Center.
TL;DR
CVPR 2026 accepted 4,090 of 16,092 submissions (25.4 %) β a +42 % absolute jump in accepted-paper volume over CVPR 2025. Embodied-AI / robotics share of accepted papers grew from ~2.9 % (CVPR 2025) to ~6.2 %, with VLA papers now the modal manipulation contribution for the first time at CVPR.
Awards. No top-award announcements yet (Best Paper / Student Paper are revealed on-site). Oral / Highlight tiers are not yet posted at the time of writing. The single confirmed Highlight in this wiki's scope is Action-Sketcher (PKU + BAAI, sketch-as-reasoning VLA).
Five distinguishing trends relative to CVPR 2025:
- VLA papers replace the diffusion-policy + benchmark mix as CVPR's modal manipulation contribution. OptimusVLA, Fast-ThinkAct, CycleVLA, DM0, RDT2, GigaBrain, DiT4DiT, EgoVLA, UniDex, ACoT-VLA, CoWVLA, HapticVLA, AnchorVLA β full-VLA papers crowd out single-trick policy work that dominated CVPR 2025.
- Humanoid loco-manipulation graduates from a niche to a sub-cluster. VIRAL (NVIDIA + Berkeley + CMU + CUHK), Opening the Sim-to-Real Door (same lab cluster), Humanoid-GPT (Tsinghua + Galbot), Gallant (Shanghai AI Lab) β four CVPR-2026 humanoid papers, vs. one (Hike-Humanoid) at CVPR 2025.
- World-model-augmented VLA training (RAMP, latent-motion CoW) is new. GigaBrain-0.5M and CoWVLA bring the world-model thread from ICLR 2026 (Cosmos Policy, Ctrl-World, DreamGen) into CVPR's perception-first venue.
- Reasoning VLAs split: action-CoT vs. visual-sketch. ACoT-VLA reasons in action space (coarse-to-fine action chunks); Action-Sketcher reasons on the image (sketched points/arrows). Compared to CVPR 2025's CoT-VLA subgoal-image CoT, both are step-changes toward interpretable reasoning.
- Egocentric β manipulation has matured. UniDex (50 k+ egocentric trajs to 8 dex hands), EgoScale (20 kh action-labeled ego video β NVIDIA's GR00T N1.7 backbone), EgoVLA (UCSD's egocentric-pretrained VLA), FunREC (functional 3D scenes from ego video) β four papers all leveraging human ego data at scale.
Strongest cross-venue seed: EgoScale (CVPR 2026) β GR00T N1.7 (NVIDIA's Apr 2026 production humanoid VLA). The 20 kh EgoScale dataset is the data delta that distinguishes N1.7 from N1.6.
Section index
- VLA Architecture & Backbones
- Reasoning / CoT / Self-Correction VLAs
- World-Model / RL-augmented VLAs
- Egocentric & Human-Video VLAs
- Humanoid / Loco-Manipulation
- Affordance / Grounding / Active Perception
- Diffusion / Flow-Matching / Specialized Policies
- Tactile / Haptic / Bimanual
- Mobile Manipulation
- Benchmarks / Datasets / Robustness
- CVPR 2026 β other-venue lineage
1. VLA Architecture & Backbones
The dominant cluster: 7+ end-to-end VLA architecture papers, each proposing a non-trivial structural twist on the standard "VLM β flow / diffusion / AR action head" stack.
| Paper | Novelty |
|---|---|
| [OptimusVLA](/Heungwoo/research/wiki/CVPR-2026-OptimusVLA) (HIT-Shenzhen + Huawei) | Dual-memory VLA β global task prior + local action-consistency memory; flow-matching head on Galaxea R1 Lite |
| [Boost-FAN](/Heungwoo/research/wiki/CVPR-2026-Boost-FAN) (Boosting VLA Finetuning with Feasible Action Neighborhood) | Physical-tolerance prior bridges SFT and RFT β actions within a feasibility neighborhood are treated as positives, recovering RL gains from supervised data |
| [Fast-ThinkAct](/Heungwoo/research/wiki/CVPR-2026-Fast-ThinkAct) (NVIDIA Research Taiwan) | Distill explicit chain-of-thought into verbalizable latent CoT tokens β 89.3 % inference-latency reduction vs. ThinkAct |
| [DM0](/Heungwoo/research/wiki/CVPR-2026-DM0) | Embodied-Native VLA pretrained jointly on web text + driving + embodied data; Embodied Spatial Scaffolding CoT-constrained action prediction |
| [RDT2](/Heungwoo/research/wiki/CVPR-2026-RDT2) (Tsinghua THU-ML) | Scaling RDT to 7 B + RVQ-tokenized action head on 10 k+ hours UMI demos; zero-shot to unseen robots |
| [DiT4DiT](/Heungwoo/research/wiki/CVPR-2026-DiT4DiT) | Cascaded video-DiT + action-DiT with dual flow matching; 98.6 % LIBERO; transfers to Unitree G1 |
2. Reasoning / CoT / Self-Correction VLAs
A clear split between action-space reasoning (think in terms of action chunks) and image-space reasoning (sketch on the image). Plus one self-correction VLA (CycleVLA) that anticipates failure at subtask boundaries.
| Paper | Novelty |
|---|---|
| [ACoT-VLA](/Heungwoo/research/wiki/CVPR-2026-ACoT-VLA) (AgiBot + BUAA) | Action chain-of-thought β coarse trajectory intent + latent action patterns fused into final action; reasoning happens entirely in action space, not text |
| [Action-Sketcher](/Heungwoo/research/wiki/CVPR-2026-Action-Sketcher) (Highlight) (PKU + BAAI) | SeeβThinkβSketchβAct β foundation model emits human-readable sketches (points, arrows) on the image; supports human-in-the-loop sketch correction |
| [CycleVLA](/Heungwoo/research/wiki/CVPR-2026-CycleVLA) (Oxford + Cambridge) | Progress-aware VLA + VLM failure predictor + Minimum Bayes Risk decoding; anticipates subtask-boundary failures and backtracks proactively |
3. World-Model / RL-augmented VLAs
CVPR 2026 imports the world-model-as-policy thread from ICLR 2026 (Cosmos Policy, Ctrl-World) and bolts it on top of VLAs explicitly. RAMP and Robo-Dopamine are the first CVPR-2026 RL-for-VLA papers proper.
| Paper | Novelty |
|---|---|
| [GigaBrain-0.5M](/Heungwoo/research/wiki/CVPR-2026-GigaBrain-0.5M) | RAMP pipeline conditions a VLA on a learned world-model's value and future-state predictions; +~30 % over RECAP-style targets |
| [Robo-Dopamine](/Heungwoo/research/wiki/CVPR-2026-Robo-Dopamine) (FlagOpen + BAAI + PKU) | General Process Reward Model for VLA-RL β predicts task progress from multi-view (init / goal / intermediate) states; reward shaping is policy-invariant |
| [CoWVLA](/Heungwoo/research/wiki/CVPR-2026-CoWVLA) (HIT + Li Auto + BAAI) | Chain of World β VLA "thinks" by predicting latent motion chains and terminal frames in a video-VAE space (not pixel-space subgoals like Ο0.7 / CoT-VLA) |
4. Egocentric & Human-Video VLAs
Four papers all attacking the same bet: the next 10Γ of robot data is web-scale egocentric human video. Together they index ~50 k+ retargeted episodes, 20 kh of action-labeled human video, and a generic functional-3D pipeline.
| Paper | Novelty |
|---|---|
| [UniDex](/Heungwoo/research/wiki/CVPR-2026-UniDex) | 50 k+ trajectories retargeted across 8 dex hands; Function-Actuator-Aligned Space (FAAS) + 3D VLA; 81 % task progress, strong zero-shot cross-hand |
| [EgoVLA](/Heungwoo/research/wiki/CVPR-2026-EgoVLA) (UCSD + UIUC + MIT + NVIDIA) | Pretrains on egocentric human video predicting wrist + hand pose; finetunes on robot data through a unified action space |
| [EgoScale](/Heungwoo/research/wiki/CVPR-2026-EgoScale) (NVIDIA GEAR) | 20 kh action-labeled ego video + lightweight aligned human-robot mid-training; log-linear scaling law linking ego-data volume to validation loss. This is the GR00T N1.7 data delta. |
| [FunREC](/Heungwoo/research/wiki/CVPR-2026-FunREC) (ETH + MPI) | Builds simulation-ready functional 3D digital twins (articulated parts + kinematics) from egocentric RGB-D; transfers directly to a mobile manipulator |
5. Humanoid / Loco-Manipulation
The fastest-growing CVPR cluster β four papers, up from one at CVPR 2025. VIRAL and Opening the Sim-to-Real Door share lab affiliations (NVIDIA + Berkeley + CMU + CUHK) and are companion papers attacking visual sim-to-real for humanoids from teacher-student and pixel-RGB angles respectively.
| Paper | Novelty |
|---|---|
| [VIRAL](/Heungwoo/research/wiki/CVPR-2026-VIRAL) (NVIDIA + Berkeley + CMU + CUHK) | Visual sim-to-real at scale for humanoid loco-manip β privileged-state RL teacher + tiled-rendering student; 54 zero-shot cycles on Unitree G1 |
| [Open-Sim-to-Real](/Heungwoo/research/wiki/CVPR-2026-Open-Sim-to-Real) (same labs) | Staged-reset exploration + GRPO fine-tune for articulated-object loco-manipulation from pure RGB; exceeds human teleop (83 % vs 80 % expert success) and completes interactions 23.1β31.7 % faster |
| [Humanoid-GPT](/Heungwoo/research/wiki/CVPR-2026-Humanoid-GPT) (Tsinghua + PKU EPIC + Galbot) | Causal-attention GPT on 2 B-frame mocap corpus; closes agility-vs-generalization gap for zero-shot motion tracking |
| [Gallant](/Heungwoo/research/wiki/CVPR-2026-Gallant) (Shanghai AI Lab) | Voxelized LiDAR + z-grouped 2D CNN; end-to-end humanoid locomotion across overhead and lateral 3D constrained terrains |
6. Affordance / Grounding / Active Perception
CVPR's perception-first DNA reasserts itself β three of the strongest papers attack what the VLA sees, not how it acts.
| Paper | Novelty |
|---|---|
| [RealVLG-R1](/Heungwoo/research/wiki/CVPR-2026-RealVLG-R1) (Tongji) | RealVLG-11B dataset (~165 k images, ~1.3 M annotations) + R1-style RFT model jointly predicting masks, bboxes, grasp poses, contact points from language |
| [SaPaVe](/Heungwoo/research/wiki/CVPR-2026-SaPaVe) (PKU + BAAI) | Decouples camera motion from manipulation; introduces ActiveViewPose-200K dataset + ActiveManip-Bench |
| [PanoAffordanceNet](/Heungwoo/research/wiki/CVPR-2026-PanoAffordanceNet) (likely) (HNU) | 360-AGD dataset + Distortion-Aware Spectral Modulator + Omni-Spherical Densification Head for panoramic affordance grounding |
7. Diffusion / Flow-Matching / Specialized Policies
Smaller than at CVPR 2025 (KStar Diffuser, FlowRAM, DexGrasp Anything dominated), but with a sharper focus on deployment-grade efficiency.
| Paper | Novelty |
|---|---|
| [AnchorVLA](/Heungwoo/research/wiki/CVPR-2026-AnchorVLA) (likely) (UQ + CSIRO) | Lightweight VLA backbone + anchored diffusion β denoises locally around expert anchor trajectories with truncated schedule; residual self-correction module |
8. Tactile / Haptic / Bimanual
| Paper | Novelty |
|---|---|
| [HapticVLA](/Heungwoo/research/wiki/CVPR-2026-HapticVLA) (likely) (Skoltech) | Two-stage: train action model with safety-aware tactile rewards, then distill tactile knowledge into a vision-only VLA so it can act without sensors at deploy; ~87 % real-world success |
| [HandX](/Heungwoo/research/wiki/CVPR-2026-HandX) (UIUC + collaborators) | Bimanual motion + interaction generation at scale β consolidates mocap + LLM-generated descriptions; scaling-law study for bimanual synthesis |
9. Mobile Manipulation
| Paper | Novelty |
|---|---|
| [EchoVLA](/Heungwoo/research/wiki/CVPR-2026-EchoVLA) (likely) (SYSU + Huawei) | Brain-inspired declarative memory pairing a spatial-semantic scene map with episodic task memory, plugged into a VLA for mobile manipulation |
| [AnchorVLA](/Heungwoo/research/wiki/CVPR-2026-AnchorVLA) (likely) | (also fits Β§7) Lightweight anchored-diffusion VLA for end-to-end mobile manipulation |
10. Benchmarks / Datasets / Robustness
| Paper | Novelty |
|---|---|
| [LIBERO-Plus](/Heungwoo/research/wiki/CVPR-2026-LIBERO-Plus) (Fudan + Shanghai AI Lab) | Automated 7-axis perturbation framework; top VLAs drop 95 % β < 30 % under modest perturbations. The robustness wake-up call. |
| [RC-NF](/Heungwoo/research/wiki/CVPR-2026-RC-NF) (Fudan ITEA + SMU) | Task-aware robot-conditioned coupling normalizing flow with dual-branch point features for sub-100 ms anomaly detection atop Ο0 / OpenVLA-class policies |
11. Lineage
CVPR 2026 β prior-venue seeds
| Prior-venue seed | CVPR 2026 descendant(s) |
|---|---|
| CoT-VLA (CVPR 2025 visual subgoal CoT) | [Action-Sketcher](/Heungwoo/research/wiki/CVPR-2026-Action-Sketcher) (sketches replace generated frames) Β· [ACoT-VLA](/Heungwoo/research/wiki/CVPR-2026-ACoT-VLA) (action-space CoT) Β· [CoWVLA](/Heungwoo/research/wiki/CVPR-2026-CoWVLA) (latent-motion CoT) |
| [dVLA](/Heungwoo/research/wiki/ICLR-2026-dVLA) + [Discrete Diffusion VLA](/Heungwoo/research/wiki/Review-Discrete-Diffusion-VLA) (ICLR 2026 unified-token CoT) | [ACoT-VLA](/Heungwoo/research/wiki/CVPR-2026-ACoT-VLA) (carries the action-CoT thread to AR backbones) |
| [ThinkAct](/Heungwoo/research/wiki/NeurIPS-2025-ThinkAct) (NeurIPS 2025 reasoning-then-action) | [Fast-ThinkAct](/Heungwoo/research/wiki/CVPR-2026-Fast-ThinkAct) (explicit-CoT β latent-CoT distillation, β89 % latency) |
| [Knowledge Insulation / Ο0.6](/Heungwoo/research/wiki/Review-pi06) + [Ο0.7](/Heungwoo/research/wiki/Review-pi07) (PI flagship line) | [OptimusVLA](/Heungwoo/research/wiki/CVPR-2026-OptimusVLA) (dual-memory VLA) Β· [CoWVLA](/Heungwoo/research/wiki/CVPR-2026-CoWVLA) (subgoal-image alternative in latent space) |
| [GR00T N1.7](/Heungwoo/research/wiki/Review-GR00T-Series) (Apr 2026 humanoid VLA) | [EgoScale](/Heungwoo/research/wiki/CVPR-2026-EgoScale) (the 20 kh ego-video dataset that is the N1.7 data delta) |
| [RDT-1B](/Heungwoo/research/wiki/Review-VLA-Architecture) (ICLR 2025 1.2 B diffusion VLA) | [RDT2](/Heungwoo/research/wiki/CVPR-2026-RDT2) (7 B + RVQ; UMI scaling) |
| [DreamGen](/Heungwoo/research/wiki/CoRL-2025-DreamGen) + [Cosmos Policy](/Heungwoo/research/wiki/ICLR-2026-Cosmos-Policy) + [Ctrl-World](/Heungwoo/research/wiki/ICLR-2026-Ctrl-World) (world-model RFT) | [GigaBrain-0.5M](/Heungwoo/research/wiki/CVPR-2026-GigaBrain-0.5M) (RAMP world-model RL) Β· [DiT4DiT](/Heungwoo/research/wiki/CVPR-2026-DiT4DiT) (cascaded video+action DiT) |
| [DexUMI](/Heungwoo/research/wiki/CoRL-2025-DexUMI) + [Visual Imitation β Humanoid](/Heungwoo/research/wiki/CoRL-2025-Visual-Imitation-Humanoid) + ICLR 2026 [EgoDex](/Heungwoo/research/wiki/ICLR-2026-EgoDex) | [UniDex](/Heungwoo/research/wiki/CVPR-2026-UniDex) Β· [EgoVLA](/Heungwoo/research/wiki/CVPR-2026-EgoVLA) Β· [FunREC](/Heungwoo/research/wiki/CVPR-2026-FunREC) |
| [ManipBench](/Heungwoo/research/wiki/CoRL-2025-ManipBench) + [RoboArena β](/Heungwoo/research/wiki/ICLR-2026-RoboArena) | [LIBERO-Plus](/Heungwoo/research/wiki/CVPR-2026-LIBERO-Plus) (7-axis robustness perturbation) |
| CoRL 2025 humanoid cluster (Visual-Imit-Humanoid, et al.) | [VIRAL](/Heungwoo/research/wiki/CVPR-2026-VIRAL) Β· [Open-Sim-to-Real](/Heungwoo/research/wiki/CVPR-2026-Open-Sim-to-Real) Β· [Humanoid-GPT](/Heungwoo/research/wiki/CVPR-2026-Humanoid-GPT) Β· [Gallant](/Heungwoo/research/wiki/CVPR-2026-Gallant) |
CVPR 2026 β likely-future descendants
- EgoScale is already inside GR00T N1.7. The next downstream is whichever NeurIPS-2026 / ICLR-2027 paper redoes the 20 kh scaling-law study at 100 kh.
- Fast-ThinkAct's latent-CoT-distillation recipe is a generic acceleration technique β expect it to show up in production VLA releases (likely Ο-series or GR00T) within 6 months.
- LIBERO-Plus's 95 β 30 % drop is a red flag for benchmark inflation; expect Q3-2026 papers that re-report all 2024β2026 numbers on the perturbed version.
Validation notes (common CVPR 2026 miscitations)
- EgoVLA is CVPR 2026 despite the arXiv submission date of 2025; confirm against authors' project page rather than arXiv listing date.
- AnchorVLA / HapticVLA / EchoVLA / PanoAffordanceNet are tagged (likely) here β they appear in secondary aggregators citing CVPR 2026, but I could not verify the
cvpr.thecvf.com/virtual/2026/poster/*URL at write-time. Confirm before citing in publications. - Robo3R (2602.10101) is RSS 2026, not CVPR 2026.
- LaRA-VLA (2602.01166) is ICML 2026, not CVPR 2026.
- WholebodyVLA (OpenDriveLab) is ICLR 2026, not CVPR β see WholeBodyVLA.
- Symmetry-Aware Vision-Tactile (2602.13689) is ICRA 2026.
- GUIDE / PULSE are CVPR 2026 Workshop papers (AI4Space and Sense of Space respectively), not main-conference.
Sources
- CVPR 2026 Dates Β· Virtual Papers Index
- Paper Copilot β CVPR 2026 statistics
- Bohrium β CVPR 2026 Highlights
- Paper Digest CVPR 2026
- Encord β CVPR 2026 Trends
- Lab/group announcement pages: NVIDIA Research Taiwan, NVIDIA GEAR, AgiBot, FlagOpen + BAAI, Tsinghua THU-ML, JHU CCVL, A*STAR CFAR