ICML 2026 CaP X - Heungwoo/research GitHub Wiki
CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation — Code-as-Policy gym, benchmark, and test-time scaling for embodied coding agents
Venue: ICML 2026 (Poster) Category: Benchmark Affiliations: Max Fu, Justin Yu, Karim El-Refai, Ethan Kou, Haoru Xue, Huang Huang, Wenli Xiao, Guanzhi Wang, Fei-Fei Li, Guanya Shi, Jiajun Wu, Shankar Sastry, Yuke Zhu, Ken Goldberg, Linxi "Jim" Fan (UC Berkeley / Stanford / CMU / NVIDIA) Traction (2026-06): 5 citations (arXiv)

Problem
"Code-as-Policy" (CaP) treats executable code as a complement to data-intensive VLA methods, but how effective frontier LLMs/VLMs are as autonomous controllers for embodied manipulation remains underexplored. Prior CaP work leans heavily on human-crafted high-level primitives, leaving open how much of the apparent competence comes from designer scaffolding versus the model itself, and how performance degrades as those priors are stripped away.
Method
CaP-X is an open-access framework with three pillars:
- CaP-Gym — an interactive environment in which agents control robots by synthesizing and executing programs that compose perception and control primitives, organized across tiers of abstraction (e.g., privileged state-based S1, noisy-perception S2, down to low-level primitives S4) and single- vs multi-turn interaction (M1–M4).
- CaP-Bench — evaluates frontier language and vision-language models across these varying levels of abstraction, interaction, and perceptual grounding (7 tasks, 12 models).
- CaP-Agent0 — a training-free agentic framework that scales test-time computation via multi-turn interaction, structured execution feedback, visual differencing (VDM), automatic skill-library synthesis, and ensembled/parallel reasoning across multiple models (e.g., Gemini-3-Pro, GPT-5.2, Claude Opus).
- CaP-RL — on-policy reinforcement learning with verifiable rewards (RLVR), applying GRPO to post-train a Qwen2.5-Coder-7B-Instruct base model directly in CaP-Gym.

Results
- CaP-Bench (12 models): A consistent trend — performance improves with human-crafted abstractions but degrades as priors are removed, exposing dependence on designer scaffolding. Models still trail human experts on generating manipulation programs even though they match humans in other coding benchmarks.
- CaP-Agent0 (100 trials/task): Despite operating solely on low-level primitives, it reaches success rates comparable to or exceeding human-written programs on 4 of 7 tasks, driven by visual differencing, a self-synthesized skill library, and parallel reasoning. On 30 LIBERO-PRO tasks it matches or exceeds training-based VLAs (OpenVLA, π0, π0.5) under position and instruction perturbations — while being entirely training-free.
- CaP-RL (Table 4): RL post-training a Qwen2.5-Coder-7B agent (50 iterations/task on S1) dramatically lifts success. In simulation (N=100): Cube Lift 25%→80%, Cube Stack 4%→44%, Spill Wipe 30%→93%. On a real Franka Emika robot (N=25): Cube Lift 24%→84%, Cube Stack 12%→76% — transferring from sim to real with minimal gap and approaching human-expert levels (92%/84%).
Significance
CaP-X provides a principled, open-access platform for studying embodied coding agents, disentangling genuine model capability from human-authored scaffolding. Its key findings — that frontier models depend on designer priors, but that this gap can be closed through agentic test-time computation (CaP-Agent0) and verifiable-reward RL (CaP-RL) — chart a practical path toward code-as-policy controllers that rival VLA post-training and human experts, with demonstrated sim-to-real transfer on a real robot.
Links
- arXiv: 2603.22435
- ICML 2026: https://icml.cc/virtual/2026/poster/66369
← Back to ICML-2026