CVPR 2026 RealVLG R1 - Heungwoo/research GitHub Wiki

RealVLG-R1 — Large-Scale Real-World Visual-Language Grounding Benchmark for Robotic Perception and Manipulation

Venue: CVPR 2026 Category: Affordance / Grounding Benchmark + Model Trend tag: Trend 1 Affiliations: Tongji University (School of Computer Science and Technology) Authors: Linfei Li, Lin Zhang, Ying Shen arXiv: 2603.14880v1 (16 Mar 2026)

Approach diagram

flowchart LR
  DATA["RealVLG-11B<br/>~165k images, ~1.3M annotations,<br/>~11B grasp examples"] --> RL["R1-style RFT<br/>(Qwen2.5-VL 3B/7B)"]
  RL --> MODEL["RealVLG-R1 model"]
  IMG["real-world image"] --> MODEL
  LANG["language query"] --> MODEL
  MODEL --> MASK["object mask"]
  MODEL --> BBOX["bounding box"]
  MODEL --> GRASP["grasp pose"]
  MODEL --> CONTACT["contact point"]
Loading

Problem

Manipulation grounding tasks (masks, bboxes, grasps, contact points) have historically been siloed — separate datasets, separate models. There has been no large-scale real-world dataset that joints all four under a single language-grounding interface.

Method

  • Build RealVLG-11B — ~165 000 real-world images over ~800 object instances, with ~1.3 M annotations (segmentation, detection, language) spanning masks, bboxes, grasp poses, and contact points, all with human-verified fine-grained language descriptions. The "11B" denotes ~11 billion enumerable grasp examples. Split into ~120K training images plus Seen / Similar / Novel evaluation sets (~15K images each).
  • Train an R1-style RFT model (Reasoning-then-Response, with rule-based reward shaping) on the unified dataset, reinforcement-fine-tuning a pretrained Qwen2.5-VL backbone (3B and 7B variants).

Results

The released model jointly predicts all four output modalities from natural language and supports zero-shot perception/manipulation in unseen real-world environments. On the Seen evaluation set, the 7B RealVLG-R1 reports bbox gIoU 89.0%, segmentation Fβ 88.9%, grasp mIoU 33.6%, and contact-point gAcc 37.2%, establishing baselines for unified cross-task grounding.

Significance

RealVLG-R1 is the CVPR-2026 grounding equivalent of CVPR-2025's RoboSpatial — but with R1-style RFT training applied and four output modalities unified. Likely to become the standard grounding-benchmark dataset for VLA-adjacent perception, replacing per-task baselines.

Links

Related pages

← Back to CVPR-2026

⚠️ **GitHub.com Fallback** ⚠️