CVPR 2026 VLA Manipulation Survey - Heungwoo/research GitHub Wiki

CVPR 2026 β€” VLA & Manipulation Survey

Compiled May 2026. Focus: CVPR 2026 VLA / manipulation papers, with a CVPR 2025 β†’ CoRL 2025 β†’ NeurIPS 2025 β†’ ICLR 2026 β†’ CVPR 2026 lineage map. CVPR 2026 runs June 3–7 at the Denver Convention Center.

TL;DR

CVPR 2026 accepted 4,090 of 16,092 submissions (25.4 %) β€” a +42 % absolute jump in accepted-paper volume over CVPR 2025. Embodied-AI / robotics share of accepted papers grew from ~2.9 % (CVPR 2025) to ~6.2 %, with VLA papers now the modal manipulation contribution for the first time at CVPR.

Awards. No top-award announcements yet (Best Paper / Student Paper are revealed on-site). Oral / Highlight tiers are not yet posted at the time of writing. The single confirmed Highlight in this wiki's scope is Action-Sketcher (PKU + BAAI, sketch-as-reasoning VLA).

Five distinguishing trends relative to CVPR 2025:

  1. VLA papers replace the diffusion-policy + benchmark mix as CVPR's modal manipulation contribution. OptimusVLA, Fast-ThinkAct, CycleVLA, DM0, RDT2, GigaBrain, DiT4DiT, EgoVLA, UniDex, ACoT-VLA, CoWVLA, HapticVLA, AnchorVLA β€” full-VLA papers crowd out single-trick policy work that dominated CVPR 2025.
  2. Humanoid loco-manipulation graduates from a niche to a sub-cluster. VIRAL (NVIDIA + Berkeley + CMU + CUHK), Opening the Sim-to-Real Door (same lab cluster), Humanoid-GPT (Tsinghua + Galbot), Gallant (Shanghai AI Lab) β€” four CVPR-2026 humanoid papers, vs. one (Hike-Humanoid) at CVPR 2025.
  3. World-model-augmented VLA training (RAMP, latent-motion CoW) is new. GigaBrain-0.5M and CoWVLA bring the world-model thread from ICLR 2026 (Cosmos Policy, Ctrl-World, DreamGen) into CVPR's perception-first venue.
  4. Reasoning VLAs split: action-CoT vs. visual-sketch. ACoT-VLA reasons in action space (coarse-to-fine action chunks); Action-Sketcher reasons on the image (sketched points/arrows). Compared to CVPR 2025's CoT-VLA subgoal-image CoT, both are step-changes toward interpretable reasoning.
  5. Egocentric β†’ manipulation has matured. UniDex (50 k+ egocentric trajs to 8 dex hands), EgoScale (20 kh action-labeled ego video β†’ NVIDIA's GR00T N1.7 backbone), EgoVLA (UCSD's egocentric-pretrained VLA), FunREC (functional 3D scenes from ego video) β€” four papers all leveraging human ego data at scale.

Strongest cross-venue seed: EgoScale (CVPR 2026) β†’ GR00T N1.7 (NVIDIA's Apr 2026 production humanoid VLA). The 20 kh EgoScale dataset is the data delta that distinguishes N1.7 from N1.6.


Section index

  1. VLA Architecture & Backbones
  2. Reasoning / CoT / Self-Correction VLAs
  3. World-Model / RL-augmented VLAs
  4. Egocentric & Human-Video VLAs
  5. Humanoid / Loco-Manipulation
  6. Affordance / Grounding / Active Perception
  7. Diffusion / Flow-Matching / Specialized Policies
  8. Tactile / Haptic / Bimanual
  9. Mobile Manipulation
  10. Benchmarks / Datasets / Robustness
  11. CVPR 2026 β†’ other-venue lineage

1. VLA Architecture & Backbones

The dominant cluster: 7+ end-to-end VLA architecture papers, each proposing a non-trivial structural twist on the standard "VLM β†’ flow / diffusion / AR action head" stack.

Paper Novelty
[OptimusVLA](/Heungwoo/research/wiki/CVPR-2026-OptimusVLA) (HIT-Shenzhen + Huawei) Dual-memory VLA β€” global task prior + local action-consistency memory; flow-matching head on Galaxea R1 Lite
[Boost-FAN](/Heungwoo/research/wiki/CVPR-2026-Boost-FAN) (Boosting VLA Finetuning with Feasible Action Neighborhood) Physical-tolerance prior bridges SFT and RFT β€” actions within a feasibility neighborhood are treated as positives, recovering RL gains from supervised data
[Fast-ThinkAct](/Heungwoo/research/wiki/CVPR-2026-Fast-ThinkAct) (NVIDIA Research Taiwan) Distill explicit chain-of-thought into verbalizable latent CoT tokens β€” 89.3 % inference-latency reduction vs. ThinkAct
[DM0](/Heungwoo/research/wiki/CVPR-2026-DM0) Embodied-Native VLA pretrained jointly on web text + driving + embodied data; Embodied Spatial Scaffolding CoT-constrained action prediction
[RDT2](/Heungwoo/research/wiki/CVPR-2026-RDT2) (Tsinghua THU-ML) Scaling RDT to 7 B + RVQ-tokenized action head on 10 k+ hours UMI demos; zero-shot to unseen robots
[DiT4DiT](/Heungwoo/research/wiki/CVPR-2026-DiT4DiT) Cascaded video-DiT + action-DiT with dual flow matching; 98.6 % LIBERO; transfers to Unitree G1

2. Reasoning / CoT / Self-Correction VLAs

A clear split between action-space reasoning (think in terms of action chunks) and image-space reasoning (sketch on the image). Plus one self-correction VLA (CycleVLA) that anticipates failure at subtask boundaries.

Paper Novelty
[ACoT-VLA](/Heungwoo/research/wiki/CVPR-2026-ACoT-VLA) (AgiBot + BUAA) Action chain-of-thought β€” coarse trajectory intent + latent action patterns fused into final action; reasoning happens entirely in action space, not text
[Action-Sketcher](/Heungwoo/research/wiki/CVPR-2026-Action-Sketcher) (Highlight) (PKU + BAAI) See–Think–Sketch–Act β€” foundation model emits human-readable sketches (points, arrows) on the image; supports human-in-the-loop sketch correction
[CycleVLA](/Heungwoo/research/wiki/CVPR-2026-CycleVLA) (Oxford + Cambridge) Progress-aware VLA + VLM failure predictor + Minimum Bayes Risk decoding; anticipates subtask-boundary failures and backtracks proactively

3. World-Model / RL-augmented VLAs

CVPR 2026 imports the world-model-as-policy thread from ICLR 2026 (Cosmos Policy, Ctrl-World) and bolts it on top of VLAs explicitly. RAMP and Robo-Dopamine are the first CVPR-2026 RL-for-VLA papers proper.

Paper Novelty
[GigaBrain-0.5M](/Heungwoo/research/wiki/CVPR-2026-GigaBrain-0.5M) RAMP pipeline conditions a VLA on a learned world-model's value and future-state predictions; +~30 % over RECAP-style targets
[Robo-Dopamine](/Heungwoo/research/wiki/CVPR-2026-Robo-Dopamine) (FlagOpen + BAAI + PKU) General Process Reward Model for VLA-RL β€” predicts task progress from multi-view (init / goal / intermediate) states; reward shaping is policy-invariant
[CoWVLA](/Heungwoo/research/wiki/CVPR-2026-CoWVLA) (HIT + Li Auto + BAAI) Chain of World β€” VLA "thinks" by predicting latent motion chains and terminal frames in a video-VAE space (not pixel-space subgoals like Ο€0.7 / CoT-VLA)

4. Egocentric & Human-Video VLAs

Four papers all attacking the same bet: the next 10Γ— of robot data is web-scale egocentric human video. Together they index ~50 k+ retargeted episodes, 20 kh of action-labeled human video, and a generic functional-3D pipeline.

Paper Novelty
[UniDex](/Heungwoo/research/wiki/CVPR-2026-UniDex) 50 k+ trajectories retargeted across 8 dex hands; Function-Actuator-Aligned Space (FAAS) + 3D VLA; 81 % task progress, strong zero-shot cross-hand
[EgoVLA](/Heungwoo/research/wiki/CVPR-2026-EgoVLA) (UCSD + UIUC + MIT + NVIDIA) Pretrains on egocentric human video predicting wrist + hand pose; finetunes on robot data through a unified action space
[EgoScale](/Heungwoo/research/wiki/CVPR-2026-EgoScale) (NVIDIA GEAR) 20 kh action-labeled ego video + lightweight aligned human-robot mid-training; log-linear scaling law linking ego-data volume to validation loss. This is the GR00T N1.7 data delta.
[FunREC](/Heungwoo/research/wiki/CVPR-2026-FunREC) (ETH + MPI) Builds simulation-ready functional 3D digital twins (articulated parts + kinematics) from egocentric RGB-D; transfers directly to a mobile manipulator

5. Humanoid / Loco-Manipulation

The fastest-growing CVPR cluster β€” four papers, up from one at CVPR 2025. VIRAL and Opening the Sim-to-Real Door share lab affiliations (NVIDIA + Berkeley + CMU + CUHK) and are companion papers attacking visual sim-to-real for humanoids from teacher-student and pixel-RGB angles respectively.

Paper Novelty
[VIRAL](/Heungwoo/research/wiki/CVPR-2026-VIRAL) (NVIDIA + Berkeley + CMU + CUHK) Visual sim-to-real at scale for humanoid loco-manip β€” privileged-state RL teacher + tiled-rendering student; 54 zero-shot cycles on Unitree G1
[Open-Sim-to-Real](/Heungwoo/research/wiki/CVPR-2026-Open-Sim-to-Real) (same labs) Staged-reset exploration + GRPO fine-tune for articulated-object loco-manipulation from pure RGB; exceeds human teleop (83 % vs 80 % expert success) and completes interactions 23.1–31.7 % faster
[Humanoid-GPT](/Heungwoo/research/wiki/CVPR-2026-Humanoid-GPT) (Tsinghua + PKU EPIC + Galbot) Causal-attention GPT on 2 B-frame mocap corpus; closes agility-vs-generalization gap for zero-shot motion tracking
[Gallant](/Heungwoo/research/wiki/CVPR-2026-Gallant) (Shanghai AI Lab) Voxelized LiDAR + z-grouped 2D CNN; end-to-end humanoid locomotion across overhead and lateral 3D constrained terrains

6. Affordance / Grounding / Active Perception

CVPR's perception-first DNA reasserts itself β€” three of the strongest papers attack what the VLA sees, not how it acts.

Paper Novelty
[RealVLG-R1](/Heungwoo/research/wiki/CVPR-2026-RealVLG-R1) (Tongji) RealVLG-11B dataset (~165 k images, ~1.3 M annotations) + R1-style RFT model jointly predicting masks, bboxes, grasp poses, contact points from language
[SaPaVe](/Heungwoo/research/wiki/CVPR-2026-SaPaVe) (PKU + BAAI) Decouples camera motion from manipulation; introduces ActiveViewPose-200K dataset + ActiveManip-Bench
[PanoAffordanceNet](/Heungwoo/research/wiki/CVPR-2026-PanoAffordanceNet) (likely) (HNU) 360-AGD dataset + Distortion-Aware Spectral Modulator + Omni-Spherical Densification Head for panoramic affordance grounding

7. Diffusion / Flow-Matching / Specialized Policies

Smaller than at CVPR 2025 (KStar Diffuser, FlowRAM, DexGrasp Anything dominated), but with a sharper focus on deployment-grade efficiency.

Paper Novelty
[AnchorVLA](/Heungwoo/research/wiki/CVPR-2026-AnchorVLA) (likely) (UQ + CSIRO) Lightweight VLA backbone + anchored diffusion β€” denoises locally around expert anchor trajectories with truncated schedule; residual self-correction module

8. Tactile / Haptic / Bimanual

Paper Novelty
[HapticVLA](/Heungwoo/research/wiki/CVPR-2026-HapticVLA) (likely) (Skoltech) Two-stage: train action model with safety-aware tactile rewards, then distill tactile knowledge into a vision-only VLA so it can act without sensors at deploy; ~87 % real-world success
[HandX](/Heungwoo/research/wiki/CVPR-2026-HandX) (UIUC + collaborators) Bimanual motion + interaction generation at scale β€” consolidates mocap + LLM-generated descriptions; scaling-law study for bimanual synthesis

9. Mobile Manipulation

Paper Novelty
[EchoVLA](/Heungwoo/research/wiki/CVPR-2026-EchoVLA) (likely) (SYSU + Huawei) Brain-inspired declarative memory pairing a spatial-semantic scene map with episodic task memory, plugged into a VLA for mobile manipulation
[AnchorVLA](/Heungwoo/research/wiki/CVPR-2026-AnchorVLA) (likely) (also fits Β§7) Lightweight anchored-diffusion VLA for end-to-end mobile manipulation

10. Benchmarks / Datasets / Robustness

Paper Novelty
[LIBERO-Plus](/Heungwoo/research/wiki/CVPR-2026-LIBERO-Plus) (Fudan + Shanghai AI Lab) Automated 7-axis perturbation framework; top VLAs drop 95 % β†’ < 30 % under modest perturbations. The robustness wake-up call.
[RC-NF](/Heungwoo/research/wiki/CVPR-2026-RC-NF) (Fudan ITEA + SMU) Task-aware robot-conditioned coupling normalizing flow with dual-branch point features for sub-100 ms anomaly detection atop Ο€0 / OpenVLA-class policies

11. Lineage

CVPR 2026 ← prior-venue seeds

Prior-venue seed CVPR 2026 descendant(s)
CoT-VLA (CVPR 2025 visual subgoal CoT) [Action-Sketcher](/Heungwoo/research/wiki/CVPR-2026-Action-Sketcher) (sketches replace generated frames) Β· [ACoT-VLA](/Heungwoo/research/wiki/CVPR-2026-ACoT-VLA) (action-space CoT) Β· [CoWVLA](/Heungwoo/research/wiki/CVPR-2026-CoWVLA) (latent-motion CoT)
[dVLA](/Heungwoo/research/wiki/ICLR-2026-dVLA) + [Discrete Diffusion VLA](/Heungwoo/research/wiki/Review-Discrete-Diffusion-VLA) (ICLR 2026 unified-token CoT) [ACoT-VLA](/Heungwoo/research/wiki/CVPR-2026-ACoT-VLA) (carries the action-CoT thread to AR backbones)
[ThinkAct](/Heungwoo/research/wiki/NeurIPS-2025-ThinkAct) (NeurIPS 2025 reasoning-then-action) [Fast-ThinkAct](/Heungwoo/research/wiki/CVPR-2026-Fast-ThinkAct) (explicit-CoT β†’ latent-CoT distillation, βˆ’89 % latency)
[Knowledge Insulation / Ο€0.6](/Heungwoo/research/wiki/Review-pi06) + [Ο€0.7](/Heungwoo/research/wiki/Review-pi07) (PI flagship line) [OptimusVLA](/Heungwoo/research/wiki/CVPR-2026-OptimusVLA) (dual-memory VLA) Β· [CoWVLA](/Heungwoo/research/wiki/CVPR-2026-CoWVLA) (subgoal-image alternative in latent space)
[GR00T N1.7](/Heungwoo/research/wiki/Review-GR00T-Series) (Apr 2026 humanoid VLA) [EgoScale](/Heungwoo/research/wiki/CVPR-2026-EgoScale) (the 20 kh ego-video dataset that is the N1.7 data delta)
[RDT-1B](/Heungwoo/research/wiki/Review-VLA-Architecture) (ICLR 2025 1.2 B diffusion VLA) [RDT2](/Heungwoo/research/wiki/CVPR-2026-RDT2) (7 B + RVQ; UMI scaling)
[DreamGen](/Heungwoo/research/wiki/CoRL-2025-DreamGen) + [Cosmos Policy](/Heungwoo/research/wiki/ICLR-2026-Cosmos-Policy) + [Ctrl-World](/Heungwoo/research/wiki/ICLR-2026-Ctrl-World) (world-model RFT) [GigaBrain-0.5M](/Heungwoo/research/wiki/CVPR-2026-GigaBrain-0.5M) (RAMP world-model RL) Β· [DiT4DiT](/Heungwoo/research/wiki/CVPR-2026-DiT4DiT) (cascaded video+action DiT)
[DexUMI](/Heungwoo/research/wiki/CoRL-2025-DexUMI) + [Visual Imitation β†’ Humanoid](/Heungwoo/research/wiki/CoRL-2025-Visual-Imitation-Humanoid) + ICLR 2026 [EgoDex](/Heungwoo/research/wiki/ICLR-2026-EgoDex) [UniDex](/Heungwoo/research/wiki/CVPR-2026-UniDex) Β· [EgoVLA](/Heungwoo/research/wiki/CVPR-2026-EgoVLA) Β· [FunREC](/Heungwoo/research/wiki/CVPR-2026-FunREC)
[ManipBench](/Heungwoo/research/wiki/CoRL-2025-ManipBench) + [RoboArena ∞](/Heungwoo/research/wiki/ICLR-2026-RoboArena) [LIBERO-Plus](/Heungwoo/research/wiki/CVPR-2026-LIBERO-Plus) (7-axis robustness perturbation)
CoRL 2025 humanoid cluster (Visual-Imit-Humanoid, et al.) [VIRAL](/Heungwoo/research/wiki/CVPR-2026-VIRAL) Β· [Open-Sim-to-Real](/Heungwoo/research/wiki/CVPR-2026-Open-Sim-to-Real) Β· [Humanoid-GPT](/Heungwoo/research/wiki/CVPR-2026-Humanoid-GPT) Β· [Gallant](/Heungwoo/research/wiki/CVPR-2026-Gallant)

CVPR 2026 β†’ likely-future descendants

  • EgoScale is already inside GR00T N1.7. The next downstream is whichever NeurIPS-2026 / ICLR-2027 paper redoes the 20 kh scaling-law study at 100 kh.
  • Fast-ThinkAct's latent-CoT-distillation recipe is a generic acceleration technique β€” expect it to show up in production VLA releases (likely Ο€-series or GR00T) within 6 months.
  • LIBERO-Plus's 95 β†’ 30 % drop is a red flag for benchmark inflation; expect Q3-2026 papers that re-report all 2024–2026 numbers on the perturbed version.

Validation notes (common CVPR 2026 miscitations)

  • EgoVLA is CVPR 2026 despite the arXiv submission date of 2025; confirm against authors' project page rather than arXiv listing date.
  • AnchorVLA / HapticVLA / EchoVLA / PanoAffordanceNet are tagged (likely) here β€” they appear in secondary aggregators citing CVPR 2026, but I could not verify the cvpr.thecvf.com/virtual/2026/poster/* URL at write-time. Confirm before citing in publications.
  • Robo3R (2602.10101) is RSS 2026, not CVPR 2026.
  • LaRA-VLA (2602.01166) is ICML 2026, not CVPR 2026.
  • WholebodyVLA (OpenDriveLab) is ICLR 2026, not CVPR β€” see WholeBodyVLA.
  • Symmetry-Aware Vision-Tactile (2602.13689) is ICRA 2026.
  • GUIDE / PULSE are CVPR 2026 Workshop papers (AI4Space and Sense of Space respectively), not main-conference.

Sources

← Back to CVPR-2026 Β· CVPR Β· Home