IROS 2026 VCoT Grasp - Heungwoo/research GitHub Wiki

IROS 2026 — VCoT-Grasp: Grasp Foundation Model with Visual Chain-of-Thought

Venue: IROS 2026 (Pittsburgh) · paper #1923 · Xi'an Jiaotong University · Westlake University · BAAI (Zhang, Bai, Zhou, … Wang, Chen) · project/code. Representative of: language-driven grasp generation with reasoning — an end-to-end grasp foundation model that adds visual chain-of-thought and ships a large dataset. Companions: Dexterous Manipulation · VLA Architectures · IROS 2026 survey.

1. Problem

Language-driven grasp generation is promising but existing approaches either lack reasoning/generalization or rely on complex modular pipelines; current grasp foundation models overemphasize dialog/object semantics, giving inferior performance and restriction to single-object grasping.

2. Method

VCoT-Grasp is an end-to-end grasp foundation model with visual chain-of-thought reasoning:

  • A multi-turn processing paradigm that dynamically focuses on visual inputs while producing interpretable reasoning traces (the "visual CoT" — reasoning grounded in the image, cf. HALO's visual foresight and ICLR: Visual Reasoning).
  • Dataset — VCoT-GraspSet: 167K synthetic images / 1.36M+ grasps + 400+ real images / 1.2K+ grasps, annotated with intermediate bounding boxes (the CoT supervision).

3. Results

  • On VCoT-GraspSet and a real robot: significantly improves grasp success and generalizes to unseen objects, backgrounds, and distractors, including cluttered (multi-object) scenes — addressing the single-object limitation of prior grasp foundation models.

4. Why it matters (dexterous × reasoning lens)

VCoT-Grasp is the IROS 2026 representative of reasoning-augmented, end-to-end grasp foundation models: it fixes prior models' single-object / weak-reasoning limits by making the model reason visually (with interpretable bbox traces) before grasping. The visual-CoT idea is the same family as HALO (visual subgoal) and ICLR: Visual Reasoning (image-space reasoning traces) — evidence that visual reasoning is becoming a cross-cutting device in 2026 manipulation, not just a VLA add-on. Contrast DexKP-VLA (modular keypoint+solver, zero-shot) — VCoT-Grasp is the end-to-end + big-dataset route to the same goal.

Limitations (reviewer): grasp generation (not full manipulation policy); heavy on a new 167K-image dataset (mostly synthetic); multi-turn visual CoT adds inference cost.

5. Links

← Back to IROS 2026 survey · Home