CVPR 2026 SaPaVe - Heungwoo/research GitHub Wiki
Venue: CVPR 2026 (Poster #37201) · arXiv 2603.12193 Category: Active Perception VLA Trend tag: Trend 1 Authors: Mengzhen Liu, Enshen Zhou, Cheng Chi, Yi Han, Shanyu Rong, Liming Chen, Pengwei Wang, Zhongyuan Wang, Shanghang Zhang Affiliations: Peking University (State Key Lab of Multimedia Information Processing) + Beihang University (School of Software) + BAAI
flowchart LR
TASK["manipulation task"] --> VLA["unified VLA<br/>(decoupled action heads)"]
VLA --> CAM_HEAD["camera-movement decoder<br/>where to look next"]
VLA --> MANIP_HEAD["manipulation decoder<br/>what to do"]
CAM_HEAD --> POSE["camera viewpoint"]
POSE --> OBS["new observation"]
OBS --> VLA
MANIP_HEAD --> ACT["robot action"]
Most VLAs assume a fixed camera viewpoint. Real manipulation often needs the robot to look around — peer behind objects, get closer for fine alignment, reposition to see occluded parts. Coupling camera motion and manipulation in one policy creates a credit-assignment mess.
A single end-to-end VLA that decouples camera movement from manipulation actions via two separate decoders (decoupled action heads on a Diffusion Transformer) — not two independent policies. Three components:
- Decoupled action heads + camera adapter — separate camera-movement and manipulation decoders reduce interference between the two action modalities.
- Universal spatial knowledge injection (3D geometry-aware module) — fuses RGB, geometric inputs, and language to keep manipulation robust under dynamic/changing viewpoints.
- Bottom-up, two-stage training — Stage 1 learns embodiment-agnostic semantic camera control on ActiveViewPose-200K; Stage 2 jointly optimizes camera movement and manipulation on a hybrid mix of manipulation datasets.
ActiveViewPose-200K: 200 k image–language–camera-movement pairs (semi-automatic pipeline: 3D asset curation, procedural indoor scenes, task-driven optimal-view annotation, LLM instruction augmentation) — the first large-scale semantic active-perception dataset.
ActiveManip-Bench (first active-manipulation benchmark): built on Isaac Sim with a humanoid + active head camera; 12 tasks, 100 objects, 20 scenes, with unoccluded / occluded / out-of-view conditions.
- Semantic active perception: 84.3% avg success on ActiveViewPose-200K, beating Gemini-2.5-Pro and other VLM baselines.
- Simulation (ActiveManip-Bench): 74.83% avg success.
- Real-world: 85.00% avg, vs GR00T-N1 53.75% and π0 45.00% (Table 3) — the headline "up to 31.25 pp higher" gain is SaPaVe (85.00) over GR00T-N1 (53.75). Deployed on a Unitree G1 with Inspire dexterous hands and a RealSense D455 RGB-D camera.
- Ablations validate both the two-stage training and the decoupled action-head design.
SaPaVe is the first principled "active perception VLA" — most prior VLAs assume passive cameras. As humanoid manipulation matures, head-camera control will be the standard. SaPaVe is the framework-and-dataset combo most likely to define how that's done.
- arXiv: 2603.12193 · Project page: https://lmzpai.github.io/SaPaVe/
- CVPR Poster: #37201
- ActiveViewPose-200K dataset + ActiveManip-Bench released alongside
← Back to CVPR-2026