CVPR 2026 Action Sketcher - Heungwoo/research GitHub Wiki

Action-Sketcher β€” From Reasoning to Action via Visual Sketches for Long-Horizon Robotic Manipulation

Venue: CVPR 2026 (Highlight) Category: Reasoning / CoT VLA (image-space reasoning) Trend tag: Trend 4 Affiliations: PKU + BAAI + U Sydney + CASIA

Approach diagram

flowchart LR
  OBS["current obs<br/>+ language goal"] --> FM["foundation model<br/>VLM"]
  FM --> SEE["See β€” perceive scene"]
  SEE --> THINK["Think β€” decide intent"]
  THINK --> SKETCH["Sketch β€” emit points,<br/>arrows, boxes on the image"]
  SKETCH --> HUMAN_OPT["optional: human edits<br/>the sketch"]
  HUMAN_OPT --> ACT["Act β€” low-level policy<br/>conditioned on sketch"]
Loading

Problem

Reasoning VLAs that emit text CoT are not human-readable in the same sense that the picture of what should happen is. For long-horizon manipulation, an interpretable visual trace would let a human (a) understand what the policy is about to do, and (b) correct it before the action.

Method

The See–Think–Sketch–Act loop. A VLM observes, reasons, and emits a human-readable Visual Sketch on the image β€” points, boxes, arrows, and typed relations annotating the next subtask. The framework is model-agnostic (it can wrap any VLA), and is instantiated with Ο€0 as the base low-level policy (SigLIP-SO400M vision + Gemma backbone + flow-matching action head). An event-driven loop summarizes the next subtask, emits the compact sketch, then synthesizes an action chunk conditioned on the sketch and robot state, with mode-switching between reasoning and acting handled by <BOR>/<BOA> gating tokens. A human-in-the-loop variant lets the operator edit the sketch before execution, correcting the policy's intent without retraining.

Results

Evaluated on LIBERO (96.9% avg success, vs. 97.1% OpenVLA-OFT and 96.8% Ο€0.5 β€” roughly on par on this short-horizon suite), RoboTwin 2.0 simulation, and real-world long-horizon tasks (Tidy Table 52.0%, Pour Tea 27.6%, Pick & Place 67.0%). The largest gains appear on long-horizon and spatially cluttered tasks. Ablations show the Visual Sketch is load-bearing: removing it collapses success to ~9.8%, and keypoints are the most critical sketch primitive.

The headline result is human-in-the-loop sketch correction: editing the sketch before execution lifts real-world success substantially (e.g., Tidy Table 52.0% β†’ ~75%, reported as +23.0%; Pour Tea +16.4%; Pick & Place +18.5%).

Significance

Action-Sketcher's sketches are the cleanest interpretable interface between a VLM's reasoning and a low-level policy published to date β€” strictly easier to inspect and edit than generated subgoal images (which require pixel-level scrutiny) or text plans (which require reading). Sits adjacent to CVPR 2025's CrayonRobo (which takes sketches as input, not output); Action-Sketcher closes the loop by generating the sketches.

Links

  • arXiv: 2601.01618
  • Project: action-sketcher.github.io

Related pages

← Back to CVPR-2026

⚠️ **GitHub.com Fallback** ⚠️