CVPR 2026 RealVLG R1 - Heungwoo/research GitHub Wiki
RealVLG-R1 — Large-Scale Real-World Visual-Language Grounding Benchmark for Robotic Perception and Manipulation
Venue: CVPR 2026 Category: Affordance / Grounding Benchmark + Model Trend tag: Trend 1 Affiliations: Tongji University (School of Computer Science and Technology) Authors: Linfei Li, Lin Zhang, Ying Shen arXiv: 2603.14880v1 (16 Mar 2026)
flowchart LR
DATA["RealVLG-11B<br/>~165k images, ~1.3M annotations,<br/>~11B grasp examples"] --> RL["R1-style RFT<br/>(Qwen2.5-VL 3B/7B)"]
RL --> MODEL["RealVLG-R1 model"]
IMG["real-world image"] --> MODEL
LANG["language query"] --> MODEL
MODEL --> MASK["object mask"]
MODEL --> BBOX["bounding box"]
MODEL --> GRASP["grasp pose"]
MODEL --> CONTACT["contact point"]
Manipulation grounding tasks (masks, bboxes, grasps, contact points) have historically been siloed — separate datasets, separate models. There has been no large-scale real-world dataset that joints all four under a single language-grounding interface.
- Build RealVLG-11B — ~165 000 real-world images over ~800 object instances, with ~1.3 M annotations (segmentation, detection, language) spanning masks, bboxes, grasp poses, and contact points, all with human-verified fine-grained language descriptions. The "11B" denotes ~11 billion enumerable grasp examples. Split into ~120K training images plus Seen / Similar / Novel evaluation sets (~15K images each).
- Train an R1-style RFT model (Reasoning-then-Response, with rule-based reward shaping) on the unified dataset, reinforcement-fine-tuning a pretrained Qwen2.5-VL backbone (3B and 7B variants).
The released model jointly predicts all four output modalities from natural language and supports zero-shot perception/manipulation in unseen real-world environments. On the Seen evaluation set, the 7B RealVLG-R1 reports bbox gIoU 89.0%, segmentation Fβ 88.9%, grasp mIoU 33.6%, and contact-point gAcc 37.2%, establishing baselines for unified cross-task grounding.
RealVLG-R1 is the CVPR-2026 grounding equivalent of CVPR-2025's RoboSpatial — but with R1-style RFT training applied and four output modalities unified. Likely to become the standard grounding-benchmark dataset for VLA-adjacent perception, replacing per-task baselines.
- arXiv: 2603.14880
- Code/data: github.com/lif314/RealVLG-R1
- Embodied-R1 · VLM4VLA · RoboSpatial (predecessor)
- CVPR 2026 survey
← Back to CVPR-2026