ICLR 2026 VLMgineer - Heungwoo/research GitHub Wiki

VLMgineer β€” VLMs as Robotic Toolsmiths

Venue: ICLR 2026 Category: Embodied reasoning Β· Tool-action co-design Trend tag: Co-design / VLM creativity / Evolutionary search Affiliations: GRASP Lab, University of Pennsylvania (Gao, Li, Shi, Li, Zhang, Figueroa, Jayaraman)

Approach diagram

flowchart TB
  Inp[Inputs:<br/>1. Raw env source code E<br/>2. Overhead image I<br/>3. Task description d_task] --> Init[Initial sampling prompt to VLM<br/>Gemini-2.5-pro]
  Init --> Pop[Population<br/>n_agent parallel agents<br/>n_tool x n_action samples each]
  Pop --> Sim[PyBullet simulation<br/>k_sim steps per design]
  Sim --> Eval[Task fitness function F<br/>per-task reward R: S β†’ r ∈ 0..1]
  Eval --> Sel[Selection: top-k by reward<br/>+ reward_save threshold]
  Sel --> Cross[Inductive in-context<br/>crossover & mutation<br/>VLM evolution prompt]
  Cross --> Pop
  Sel -->|n_iteration done| Best[Best D* tool URDF + action waypoints]
  Best --> Real[3D print + Franka mount<br/>play action plan]
Loading

Problem

Most robot-learning research fixes the gripper / tool and improves the controller. The paper inverts that question: can a foundation model invent a tool so that control becomes easy? The authors recast tool geometry as a form of physical intelligence complementary to the policy. Doing this fully automatically β€” without human-specified parametric tool templates that prior morphology-optimization work relies on β€” is the real challenge.

Detailed Method

1. Inputs & VLM design agent

Three context streams:

  • Environment source code E β€” unmodified Python class for the simulation environment.
  • Workspace image I β€” overhead camera render.
  • Task description d_task β€” natural-language instruction.

The VLM (Gemini-2.5-pro is used in the main results) is prompted to write down (a) a free-form analysis of the scene, (b) a URDF tool design as modular blocks attached to the Franka end-effector, and (c) an NΓ—7 action sequence (6-DoF pose + gripper open/close per waypoint).

Tool representation. URDF is chosen over mesh / CAD / block representations because its modular text structure aligns with the VLM's strengths in code generation. The VLM output is dropped directly into a designated end-effector link of the robot URDF.

Action representation. Discrete waypoints β€” chosen intentionally to demonstrate that a smarter tool reduces the need for sophisticated policies. The framework supports force-vector or wrench actions if needed.

2. Joint tool-action candidate sampling

The VLM proposes n_tool tool designs and m = n_action action plans per tool in a single inference pass, totaling n_tool Γ— n_action tool-action pairs per iteration per agent. Authors emphasize this is "a kind of crude VLM-guided policy optimization" β€” far cheaper than RL β€” exploiting the insight that good tools simplify required action plans. They report empirically that visual feedback (passing the last frame back to the VLM) slightly hurts average reward (0.557 β†’ 0.523; Table 5), because VLMs struggle to ground fine-grained physical progress from raw frames.

3. Evolutionary loop (Algorithm 1)

for n_iteration iterations:
  D1...DK ~ VLM(E, I, d_task, PROMPT)        // sample K designs
  s1...sK = F(D1)...F(DK)                    // simulator evaluation
  Select top-k {Dj1, ..., Djk}               // selection
  PROMPT := EvolutionPrompt({Dj1, ..., Djk}) // in-context crossover + mutation
return D* = arg max F(D) across all iterations

The evolution prompt instructs the VLM:

"Your design decision is part of a genetic algorithm for tool creation, where each new design is produced either by mutationβ€”changing exactly one aspect (e.g., adjusting a component's dimension or adding/removing a component)β€”or by crossover, combining elements from two existing designs. All resulting mutations and crossovers should plausibly enhance task success while preserving design diversity."

4. Hyperparameters (Appendix)

Per-task hyperparameters live in an appendix table; the relevant knobs are:

  • n_agent β€” parallel evolutionary agents (multiple agents avoid local minima and improve runtime).
  • n_tool, n_action β€” per-iteration sample counts (total samples per iter per agent = n_tool Γ— n_action).
  • k_top, reward_save β€” top-k cutoff plus a minimum-reward threshold for retention.
  • n_iteration β€” number of evolution rounds.
  • k_sim β€” simulator steps per evaluation.
  • VLM: gemini-2.5-pro-preview-03-25 throughout main experiments.
  • Compute: AMD Ryzen 7 9800X3D 8-core, 64 GB RAM. One run on one task β‰ˆ 30 minutes.

The main experiment uses 4 evolution iterations Γ— 2000 samples each = 8000-sample budget (Table 1).

RoboToolBench (the contributed benchmark)

12 PyBullet manipulation tasks specifically chosen to be hard or impossible with a vanilla two-finger gripper. Franka Panda is the standard robot.

Task Inspiration
BringCube RLBench tool subset
CleanTable RLBench
DislodgeCube Caledonian-crow tool insertion (Jacobs et al., 2016)
ElevatePlate Co-design literature (Liu et al., 2023)
GatherSpheres RLBench
HighObject Co-design literature
LiftBox Everyday home
MoveBall Everyday home
OneBook Everyday home
ScoreGoal RLBench
SnatchCookie Everyday home
TurkeyLegs Everyday home

Reward R: S β†’ r ∈ [0, 1] is task-specific.

Comprehensive Results

Main result (Fig. 4 / Fig. 9)

Across all 12 tasks, VLMgineer's average and best-run reward exceed all baselines:

Baseline Description Result on ROBOTOOLBENCH
Franka Gripper Vanilla two-finger gripper, no tool Fails on majority of tasks
Human Prompts (Robotics expert) Graduate student prompts the VLM Lower mean & higher variance than VLMgineer
Human Prompts (LLM expert) Graduate student prompts the VLM Lower mean & higher variance
Human Prompts (Layperson) Undergrad prompts the VLM Lower mean & higher variance
RLBench Tools Original tool meshes from RLBench Lower than VLMgineer on all 4 RLBench tasks
VLMgineer Full pipeline Highest mean and best reward on all 12 tasks

Headline normalized improvements (abstract / Sec. 1):

  • +64.7% average normalized improvement over VLM-from-human-spec baselines.
  • +24.3% average normalized improvement over human-crafted RLBench tools.

(Per-task numerical reward bars are in Fig. 4 / Fig. 9 of the paper. Specific per-task scores beyond what's in the tables below are reported as bar charts and not as stated numerical values in the body β€” flagged here as not stated in numeric form for some tasks.)

Ablation: evolution vs brute-force sampling (Table 1)

Equal sample budget = 8000 samples. VLMgineer = 4 iterations Γ— 2000 samples; baseline = single-iteration VLM sampling Γ— 8000.

Method ElevatePlate RetrieveHigh CleanTable Average
VLMgineer (Evolution) 0.925 Β± 0.007 1.000 Β± 0.000 0.888 Β± 0.018 0.938 (+119.2%)
VLM Sampling (Baseline) 0.317 Β± 0.254 0.501 Β± 0.705 0.466 Β± 0.124 0.428

Evolution beats brute-force by +119.2% on average β€” the iterative refinement is the load-bearing piece, not just sampling many candidates. (Values verbatim from the paper's Table 1.)

Ablation: VLM choice (Tables 4 & 5)

Table 4 compares Gemini-2.5-pro vs GPT-o3; Table 5 compares the Gemini family:

Model Avg Reward Top Reward
Gemini-2.5-pro 0.6054 0.8222
GPT-o3 0.3775 0.5436
Gemini-2.5-flash 0.3393 0.4481
Gemini-2.0-flash 0.0686 0.0796

Gemini-2.5-pro is dramatically better than GPT-o3 and prior Gemini generations.

Ablation: image feedback (Table 6)

Tested on 3 tasks at 500 samples:

Method ElevatePlate GatherSpheres MoveBall Avg
VLMgineer w/ Img Feedback 0.30 0.45 0.80 0.52
VLMgineer w/o Img Feedback 0.28 0.56 0.82 0.55

Adding image feedback drops average reward by 5.4%; VLMs hallucinate progress from raw frames.

Sim-to-Real on a Franka Panda

3 tasks selected for ease of real-world replication. Best-of-simulation tool is 3D-printed (Bambu Lab P1S, supports added then removed), mounted via screws to a custom Franka end-effector mount, and the cached action waypoints are replayed by a position controller. 5 runs per task. Real-world average normalized rewards: MoveBall 0.959, ElevatePlate 0.761, GatherSpheres 0.713. Zero-shot transfer with no real-world fine-tuning.

Qualitative findings

  • GatherSpheres: evolution adds guardrails to an open scoop (after-evolution screenshot) β€” preventing spheres from bouncing out.
  • MoveBall: open-ended pusher gains a "hugging rim" through evolution.
  • ScoreGoal: VLMgineer produces long-and-bent shapes that simplify the action plan to one-axis motion, vs. straight tools from human prompting that need more careful control.
  • BringCube: cage-like structure that locks the cube reliably, vs. RLBench's simple stick which under-controls.

Limitations (as stated by authors, Sec. 7)

  1. Sim-to-real breadth. Successful on 3 tasks, but more dynamic real-world scenarios are not validated.
  2. Discrete waypoint actions. Chosen to demonstrate "smart tools reduce policy complexity", but limits handling of dynamic/contact-rich tasks. Future work: integrate force/wrench-conditioned closed-loop policies.
  3. Geometry constraints. URDF tool designs are simple geometries (modular blocks). No comprehensive evaluation of articulated or compliant tools.
  4. Single-task focus. No multitask optimization or generalization across tasks.
  5. Manufacturing constraints absent. Common failure: tools too thin for reliable 3D printing or too large for single-piece fabrication. Future work needs to bake manufacturability into the optimization.
  6. "Why design tools at all?" Honest framing: the paper's stance is that one cannot assume a convenient tool already exists. Integrating tool selection (choose-from-existing) with VLMgineer-style invention is future work.
  7. Requires raw env source code as input. Not a "pure photo-only" system. Digital-twin or sim-to-real construction modules are out of scope and left to existing pipelines.

Common pipeline failure modes (Appendix A.8): physically infeasible initial designs (penetrating environment / robot due to imperfect physics simulation); fabrication issues (thin / oversized parts).

Significance & Positioning

VLMgineer pushes VLA-adjacent research beyond "policy as the only learnable thing" β€” morphology itself becomes a search variable. Compared to neighbors:

  • Eureka (Ma et al., 2023) uses LLMs to evolve reward functions for RL. VLMgineer's analogue: evolve tools and actions directly. The two are complementary and could be stacked.
  • RoboMorph (Qiu et al., 2024), LASER (Song et al., 2025), Text2Robot (Ringel et al., 2025) apply LLM-aided evolution to locomotion robot morphology, where actions are still RL-trained. VLMgineer's contribution: extend to manipulation + joint tool-action sampling without RL.
  • Khan et al. (2025) "Evolution 6.0" uses VLMs for tool design but leans on retrieval from existing-object databases. VLMgineer's claim is that evolutionary search elicits genuine VLM physical creativity beyond single-shot prompting β€” supported by the 119.2% gap over brute-force sampling.
  • Liu et al. (2023) "Learning to design and use tools" does joint tool-policy design via differentiable simulation but requires manually-specified parametric design spaces. VLMgineer is fully open-vocabulary in tool space.
  • Articulate-Anything (Le et al., 2024) by overlapping authors validated the URDF-as-VLM-output pattern; VLMgineer is the manipulation/co-design instantiation.

The broader claim is that foundation models contain useful priors over physical morphology, not just over policies β€” and evolutionary search is the right amplifier. This is a different lens from the Ο€0/Ο€0.5/GR00T/OpenVLA family which all fix a hardware platform and improve the policy. If VLMgineer-style co-design becomes standard, future "VLAs" may co-evolve hardware and policy in deployment.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️