ICRA 2026 Goal VLA - Heungwoo/research GitHub Wiki

Goal-VLA โ€” Image-Generative VLMs as Object-Centric World Models

Venue: ICRA 2026 ยท Authors: Haonan Chen, Jingxiang Guo, Bangjun Wang, Tianrui Zhang, Xuchuan Huang, Boren Zheng, Yiwen Hou, Chenrui Tie, Jiajun Deng, Lin Shao โ€” NUS LinS Lab (with HKU, Peking Univ., Tsinghua) ยท arXiv: 2506.23919 Category: Goal-image / world-model-conditioned VLA (zero-shot) Trend tag: Generated goal state as the interface

Approach diagram

flowchart LR
  INSTR["language instruction"] --> VLM["image-generative VLM<br/>(world model)"]
  OBS["initial observation"] --> VLM
  VLM --> GOAL["generated goal image"]
  GOAL --> RTS{"Reflection-through-<br/>Synthesis loop"}
  RTS -- "infeasible: refine" --> VLM
  RTS -- "validated goal<br/>(image + mask + depth)" --> GROUND["spatial grounding<br/>(feature match + point-cloud reg.)"]
  GROUND --> POSE["target object transform<br/>(object-pose interface)"]
  POSE --> LOW["training-free low-level policy<br/>(contact pose โ†’ motion plan)"]
  LOW --> ACT["robot execution"]
Loading

Problem

Recent Vision-Language-Action (VLA) models build policies on top of VLMs to inherit open-world semantics, but their zero-shot ability lags far behind the base VLM: instruction-vision-action data is too scarce to cover diverse scenes, tasks, and embodiments. Goal-VLA asks whether the generalizable knowledge of a VLM can drive manipulation without any action-labeled training, by choosing a representation that does not require action annotations at all.

Method

The core claim is that object state is the "golden interface" that cleanly splits a manipulation system into a generalizable high-level policy and a training-free low-level policy. The pipeline runs in three stages:

  1. Goal-state reasoning. An image-generative VLM acts as an object-centric world model, synthesizing a goal image from the instruction and initial observation. A Reflection-through-Synthesis loop iteratively validates and refines this image for task feasibility, emitting a validated goal as image + mask + depth.
  2. Spatial grounding. The object's rigid transformation between initial and goal states is recovered by feature matching and point-cloud registration โ€” no learned action head.
  3. Low-level policy. Applying that object transform to a contact pose yields the gripper goal pose, after which a motion planner produces the executed trajectory.

Because the interface is object pose rather than actions, the high-level VLM stays fully generalizable while control remains training-free.

Results

Verified numbers from the paper. Simulation (RLBench, 8 tasks, 100 runs each): Goal-VLA averages 59.9% success, versus MOKA 26.0%, MolmoAct 11.3%, VoxPoser 5.8%, OpenVLA 0.2%, ฯ€0 0.0%, SUSIE 0.0%. Real-world (4 tasks, 10 trials each): Goal-VLA 60% average (e.g., tomato placement 9/10, weighing duck 7/10), versus MolmoAct 27.5%, MOKA 22.5%, OpenVLA 0%. Ablation: a 40.0% baseline rises to 83.8% with input enhancement + reflector, and to 88.8% with three reflection iterations โ€” quantifying the Reflection-through-Synthesis contribution.

Significance

Goal-VLA belongs to the editing-diffusion / world-model-as-prompt clusters of Goal-Image-Conditioning review: rather than predicting actions, it generates a goal image and treats the derived object pose as the conditioning interface. Against the taxonomy of VLA Architectures review, it is a hierarchical, non-trained-policy design โ€” the VLM supplies semantics, classical geometry supplies control โ€” which is precisely why it can be zero-shot while learned VLAs (OpenVLA, ฯ€0) collapse to near-zero on these unseen tasks. The Reflection-through-Synthesis loop is the load-bearing idea: closing a validate-and-refine cycle over a generative world model converts an unreliable single-shot goal image into a usable plan, and the ablation shows most of the gain comes from it.

Links

Related pages

โ† Back to ICRA-2026

โš ๏ธ **GitHub.com Fallback** โš ๏ธ