CVPR 2026 Action Sketcher - Heungwoo/research GitHub Wiki
Action-Sketcher β From Reasoning to Action via Visual Sketches for Long-Horizon Robotic Manipulation
Venue: CVPR 2026 (Highlight) Category: Reasoning / CoT VLA (image-space reasoning) Trend tag: Trend 4 Affiliations: PKU + BAAI + U Sydney + CASIA
flowchart LR
OBS["current obs<br/>+ language goal"] --> FM["foundation model<br/>VLM"]
FM --> SEE["See β perceive scene"]
SEE --> THINK["Think β decide intent"]
THINK --> SKETCH["Sketch β emit points,<br/>arrows, boxes on the image"]
SKETCH --> HUMAN_OPT["optional: human edits<br/>the sketch"]
HUMAN_OPT --> ACT["Act β low-level policy<br/>conditioned on sketch"]
Reasoning VLAs that emit text CoT are not human-readable in the same sense that the picture of what should happen is. For long-horizon manipulation, an interpretable visual trace would let a human (a) understand what the policy is about to do, and (b) correct it before the action.
The SeeβThinkβSketchβAct loop. A VLM observes, reasons, and emits a human-readable Visual Sketch on the image β points, boxes, arrows, and typed relations annotating the next subtask. The framework is model-agnostic (it can wrap any VLA), and is instantiated with Ο0 as the base low-level policy (SigLIP-SO400M vision + Gemma backbone + flow-matching action head). An event-driven loop summarizes the next subtask, emits the compact sketch, then synthesizes an action chunk conditioned on the sketch and robot state, with mode-switching between reasoning and acting handled by <BOR>/<BOA> gating tokens. A human-in-the-loop variant lets the operator edit the sketch before execution, correcting the policy's intent without retraining.
Evaluated on LIBERO (96.9% avg success, vs. 97.1% OpenVLA-OFT and 96.8% Ο0.5 β roughly on par on this short-horizon suite), RoboTwin 2.0 simulation, and real-world long-horizon tasks (Tidy Table 52.0%, Pour Tea 27.6%, Pick & Place 67.0%). The largest gains appear on long-horizon and spatially cluttered tasks. Ablations show the Visual Sketch is load-bearing: removing it collapses success to ~9.8%, and keypoints are the most critical sketch primitive.
The headline result is human-in-the-loop sketch correction: editing the sketch before execution lifts real-world success substantially (e.g., Tidy Table 52.0% β ~75%, reported as +23.0%; Pour Tea +16.4%; Pick & Place +18.5%).
Action-Sketcher's sketches are the cleanest interpretable interface between a VLM's reasoning and a low-level policy published to date β strictly easier to inspect and edit than generated subgoal images (which require pixel-level scrutiny) or text plans (which require reading). Sits adjacent to CVPR 2025's CrayonRobo (which takes sketches as input, not output); Action-Sketcher closes the loop by generating the sketches.
- arXiv: 2601.01618
- Project:
action-sketcher.github.io
- Goal-Image Conditioning review β the "visual reasoning interface" family
- CoT-VLA (subgoal-image CoT predecessor)
- CVPR 2026 survey
β Back to CVPR-2026