ICLR 2026 VLMgineer - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Category: Embodied reasoning Β· Tool-action co-design Trend tag: Co-design / VLM creativity / Evolutionary search Affiliations: GRASP Lab, University of Pennsylvania (Gao, Li, Shi, Li, Zhang, Figueroa, Jayaraman)
flowchart TB
Inp[Inputs:<br/>1. Raw env source code E<br/>2. Overhead image I<br/>3. Task description d_task] --> Init[Initial sampling prompt to VLM<br/>Gemini-2.5-pro]
Init --> Pop[Population<br/>n_agent parallel agents<br/>n_tool x n_action samples each]
Pop --> Sim[PyBullet simulation<br/>k_sim steps per design]
Sim --> Eval[Task fitness function F<br/>per-task reward R: S β r β 0..1]
Eval --> Sel[Selection: top-k by reward<br/>+ reward_save threshold]
Sel --> Cross[Inductive in-context<br/>crossover & mutation<br/>VLM evolution prompt]
Cross --> Pop
Sel -->|n_iteration done| Best[Best D* tool URDF + action waypoints]
Best --> Real[3D print + Franka mount<br/>play action plan]
Most robot-learning research fixes the gripper / tool and improves the controller. The paper inverts that question: can a foundation model invent a tool so that control becomes easy? The authors recast tool geometry as a form of physical intelligence complementary to the policy. Doing this fully automatically β without human-specified parametric tool templates that prior morphology-optimization work relies on β is the real challenge.
Three context streams:
-
Environment source code
Eβ unmodified Python class for the simulation environment. -
Workspace image
Iβ overhead camera render. -
Task description
d_taskβ natural-language instruction.
The VLM (Gemini-2.5-pro is used in the main results) is prompted to write down (a) a free-form analysis of the scene, (b) a URDF tool design as modular blocks attached to the Franka end-effector, and (c) an NΓ7 action sequence (6-DoF pose + gripper open/close per waypoint).
Tool representation. URDF is chosen over mesh / CAD / block representations because its modular text structure aligns with the VLM's strengths in code generation. The VLM output is dropped directly into a designated end-effector link of the robot URDF.
Action representation. Discrete waypoints β chosen intentionally to demonstrate that a smarter tool reduces the need for sophisticated policies. The framework supports force-vector or wrench actions if needed.
The VLM proposes n_tool tool designs and m = n_action action plans per tool in a single inference pass, totaling n_tool Γ n_action tool-action pairs per iteration per agent. Authors emphasize this is "a kind of crude VLM-guided policy optimization" β far cheaper than RL β exploiting the insight that good tools simplify required action plans. They report empirically that visual feedback (passing the last frame back to the VLM) slightly hurts average reward (0.557 β 0.523; Table 5), because VLMs struggle to ground fine-grained physical progress from raw frames.
for n_iteration iterations:
D1...DK ~ VLM(E, I, d_task, PROMPT) // sample K designs
s1...sK = F(D1)...F(DK) // simulator evaluation
Select top-k {Dj1, ..., Djk} // selection
PROMPT := EvolutionPrompt({Dj1, ..., Djk}) // in-context crossover + mutation
return D* = arg max F(D) across all iterations
The evolution prompt instructs the VLM:
"Your design decision is part of a genetic algorithm for tool creation, where each new design is produced either by mutationβchanging exactly one aspect (e.g., adjusting a component's dimension or adding/removing a component)βor by crossover, combining elements from two existing designs. All resulting mutations and crossovers should plausibly enhance task success while preserving design diversity."
Per-task hyperparameters live in an appendix table; the relevant knobs are:
-
n_agentβ parallel evolutionary agents (multiple agents avoid local minima and improve runtime). -
n_tool,n_actionβ per-iteration sample counts (total samples per iter per agent =n_tool Γ n_action). -
k_top,reward_saveβ top-k cutoff plus a minimum-reward threshold for retention. -
n_iterationβ number of evolution rounds. -
k_simβ simulator steps per evaluation. - VLM: gemini-2.5-pro-preview-03-25 throughout main experiments.
- Compute: AMD Ryzen 7 9800X3D 8-core, 64 GB RAM. One run on one task β 30 minutes.
The main experiment uses 4 evolution iterations Γ 2000 samples each = 8000-sample budget (Table 1).
12 PyBullet manipulation tasks specifically chosen to be hard or impossible with a vanilla two-finger gripper. Franka Panda is the standard robot.
| Task | Inspiration |
|---|---|
| BringCube | RLBench tool subset |
| CleanTable | RLBench |
| DislodgeCube | Caledonian-crow tool insertion (Jacobs et al., 2016) |
| ElevatePlate | Co-design literature (Liu et al., 2023) |
| GatherSpheres | RLBench |
| HighObject | Co-design literature |
| LiftBox | Everyday home |
| MoveBall | Everyday home |
| OneBook | Everyday home |
| ScoreGoal | RLBench |
| SnatchCookie | Everyday home |
| TurkeyLegs | Everyday home |
Reward R: S β r β [0, 1] is task-specific.
Across all 12 tasks, VLMgineer's average and best-run reward exceed all baselines:
| Baseline | Description | Result on ROBOTOOLBENCH |
|---|---|---|
| Franka Gripper | Vanilla two-finger gripper, no tool | Fails on majority of tasks |
| Human Prompts (Robotics expert) | Graduate student prompts the VLM | Lower mean & higher variance than VLMgineer |
| Human Prompts (LLM expert) | Graduate student prompts the VLM | Lower mean & higher variance |
| Human Prompts (Layperson) | Undergrad prompts the VLM | Lower mean & higher variance |
| RLBench Tools | Original tool meshes from RLBench | Lower than VLMgineer on all 4 RLBench tasks |
| VLMgineer | Full pipeline | Highest mean and best reward on all 12 tasks |
Headline normalized improvements (abstract / Sec. 1):
- +64.7% average normalized improvement over VLM-from-human-spec baselines.
- +24.3% average normalized improvement over human-crafted RLBench tools.
(Per-task numerical reward bars are in Fig. 4 / Fig. 9 of the paper. Specific per-task scores beyond what's in the tables below are reported as bar charts and not as stated numerical values in the body β flagged here as not stated in numeric form for some tasks.)
Equal sample budget = 8000 samples. VLMgineer = 4 iterations Γ 2000 samples; baseline = single-iteration VLM sampling Γ 8000.
| Method | ElevatePlate | RetrieveHigh | CleanTable | Average |
|---|---|---|---|---|
| VLMgineer (Evolution) | 0.925 Β± 0.007 | 1.000 Β± 0.000 | 0.888 Β± 0.018 | 0.938 (+119.2%) |
| VLM Sampling (Baseline) | 0.317 Β± 0.254 | 0.501 Β± 0.705 | 0.466 Β± 0.124 | 0.428 |
Evolution beats brute-force by +119.2% on average β the iterative refinement is the load-bearing piece, not just sampling many candidates. (Values verbatim from the paper's Table 1.)
Table 4 compares Gemini-2.5-pro vs GPT-o3; Table 5 compares the Gemini family:
| Model | Avg Reward | Top Reward |
|---|---|---|
| Gemini-2.5-pro | 0.6054 | 0.8222 |
| GPT-o3 | 0.3775 | 0.5436 |
| Gemini-2.5-flash | 0.3393 | 0.4481 |
| Gemini-2.0-flash | 0.0686 | 0.0796 |
Gemini-2.5-pro is dramatically better than GPT-o3 and prior Gemini generations.
Tested on 3 tasks at 500 samples:
| Method | ElevatePlate | GatherSpheres | MoveBall | Avg |
|---|---|---|---|---|
| VLMgineer w/ Img Feedback | 0.30 | 0.45 | 0.80 | 0.52 |
| VLMgineer w/o Img Feedback | 0.28 | 0.56 | 0.82 | 0.55 |
Adding image feedback drops average reward by 5.4%; VLMs hallucinate progress from raw frames.
3 tasks selected for ease of real-world replication. Best-of-simulation tool is 3D-printed (Bambu Lab P1S, supports added then removed), mounted via screws to a custom Franka end-effector mount, and the cached action waypoints are replayed by a position controller. 5 runs per task. Real-world average normalized rewards: MoveBall 0.959, ElevatePlate 0.761, GatherSpheres 0.713. Zero-shot transfer with no real-world fine-tuning.
- GatherSpheres: evolution adds guardrails to an open scoop (after-evolution screenshot) β preventing spheres from bouncing out.
- MoveBall: open-ended pusher gains a "hugging rim" through evolution.
- ScoreGoal: VLMgineer produces long-and-bent shapes that simplify the action plan to one-axis motion, vs. straight tools from human prompting that need more careful control.
- BringCube: cage-like structure that locks the cube reliably, vs. RLBench's simple stick which under-controls.
- Sim-to-real breadth. Successful on 3 tasks, but more dynamic real-world scenarios are not validated.
- Discrete waypoint actions. Chosen to demonstrate "smart tools reduce policy complexity", but limits handling of dynamic/contact-rich tasks. Future work: integrate force/wrench-conditioned closed-loop policies.
- Geometry constraints. URDF tool designs are simple geometries (modular blocks). No comprehensive evaluation of articulated or compliant tools.
- Single-task focus. No multitask optimization or generalization across tasks.
- Manufacturing constraints absent. Common failure: tools too thin for reliable 3D printing or too large for single-piece fabrication. Future work needs to bake manufacturability into the optimization.
- "Why design tools at all?" Honest framing: the paper's stance is that one cannot assume a convenient tool already exists. Integrating tool selection (choose-from-existing) with VLMgineer-style invention is future work.
- Requires raw env source code as input. Not a "pure photo-only" system. Digital-twin or sim-to-real construction modules are out of scope and left to existing pipelines.
Common pipeline failure modes (Appendix A.8): physically infeasible initial designs (penetrating environment / robot due to imperfect physics simulation); fabrication issues (thin / oversized parts).
VLMgineer pushes VLA-adjacent research beyond "policy as the only learnable thing" β morphology itself becomes a search variable. Compared to neighbors:
- Eureka (Ma et al., 2023) uses LLMs to evolve reward functions for RL. VLMgineer's analogue: evolve tools and actions directly. The two are complementary and could be stacked.
- RoboMorph (Qiu et al., 2024), LASER (Song et al., 2025), Text2Robot (Ringel et al., 2025) apply LLM-aided evolution to locomotion robot morphology, where actions are still RL-trained. VLMgineer's contribution: extend to manipulation + joint tool-action sampling without RL.
- Khan et al. (2025) "Evolution 6.0" uses VLMs for tool design but leans on retrieval from existing-object databases. VLMgineer's claim is that evolutionary search elicits genuine VLM physical creativity beyond single-shot prompting β supported by the 119.2% gap over brute-force sampling.
- Liu et al. (2023) "Learning to design and use tools" does joint tool-policy design via differentiable simulation but requires manually-specified parametric design spaces. VLMgineer is fully open-vocabulary in tool space.
- Articulate-Anything (Le et al., 2024) by overlapping authors validated the URDF-as-VLM-output pattern; VLMgineer is the manipulation/co-design instantiation.
The broader claim is that foundation models contain useful priors over physical morphology, not just over policies β and evolutionary search is the right amplifier. This is a different lens from the Ο0/Ο0.5/GR00T/OpenVLA family which all fix a hardware platform and improve the policy. If VLMgineer-style co-design becomes standard, future "VLAs" may co-evolve hardware and policy in deployment.
- OpenReview: https://openreview.net/forum?id=nESyz4PvJL
- PDF: https://openreview.net/pdf?id=nESyz4PvJL
- arXiv: https://arxiv.org/abs/2507.12644
- Project page: https://vlmgineer.github.io
β Back to ICLR-2026