ICRA 2026 Goal VLA - Heungwoo/research GitHub Wiki
Venue: ICRA 2026 ยท Authors: Haonan Chen, Jingxiang Guo, Bangjun Wang, Tianrui Zhang, Xuchuan Huang, Boren Zheng, Yiwen Hou, Chenrui Tie, Jiajun Deng, Lin Shao โ NUS LinS Lab (with HKU, Peking Univ., Tsinghua) ยท arXiv: 2506.23919 Category: Goal-image / world-model-conditioned VLA (zero-shot) Trend tag: Generated goal state as the interface
flowchart LR
INSTR["language instruction"] --> VLM["image-generative VLM<br/>(world model)"]
OBS["initial observation"] --> VLM
VLM --> GOAL["generated goal image"]
GOAL --> RTS{"Reflection-through-<br/>Synthesis loop"}
RTS -- "infeasible: refine" --> VLM
RTS -- "validated goal<br/>(image + mask + depth)" --> GROUND["spatial grounding<br/>(feature match + point-cloud reg.)"]
GROUND --> POSE["target object transform<br/>(object-pose interface)"]
POSE --> LOW["training-free low-level policy<br/>(contact pose โ motion plan)"]
LOW --> ACT["robot execution"]
Recent Vision-Language-Action (VLA) models build policies on top of VLMs to inherit open-world semantics, but their zero-shot ability lags far behind the base VLM: instruction-vision-action data is too scarce to cover diverse scenes, tasks, and embodiments. Goal-VLA asks whether the generalizable knowledge of a VLM can drive manipulation without any action-labeled training, by choosing a representation that does not require action annotations at all.
The core claim is that object state is the "golden interface" that cleanly splits a manipulation system into a generalizable high-level policy and a training-free low-level policy. The pipeline runs in three stages:
- Goal-state reasoning. An image-generative VLM acts as an object-centric world model, synthesizing a goal image from the instruction and initial observation. A Reflection-through-Synthesis loop iteratively validates and refines this image for task feasibility, emitting a validated goal as image + mask + depth.
- Spatial grounding. The object's rigid transformation between initial and goal states is recovered by feature matching and point-cloud registration โ no learned action head.
- Low-level policy. Applying that object transform to a contact pose yields the gripper goal pose, after which a motion planner produces the executed trajectory.
Because the interface is object pose rather than actions, the high-level VLM stays fully generalizable while control remains training-free.
Verified numbers from the paper. Simulation (RLBench, 8 tasks, 100 runs each): Goal-VLA averages 59.9% success, versus MOKA 26.0%, MolmoAct 11.3%, VoxPoser 5.8%, OpenVLA 0.2%, ฯ0 0.0%, SUSIE 0.0%. Real-world (4 tasks, 10 trials each): Goal-VLA 60% average (e.g., tomato placement 9/10, weighing duck 7/10), versus MolmoAct 27.5%, MOKA 22.5%, OpenVLA 0%. Ablation: a 40.0% baseline rises to 83.8% with input enhancement + reflector, and to 88.8% with three reflection iterations โ quantifying the Reflection-through-Synthesis contribution.
Goal-VLA belongs to the editing-diffusion / world-model-as-prompt clusters of Goal-Image-Conditioning review: rather than predicting actions, it generates a goal image and treats the derived object pose as the conditioning interface. Against the taxonomy of VLA Architectures review, it is a hierarchical, non-trained-policy design โ the VLM supplies semantics, classical geometry supplies control โ which is precisely why it can be zero-shot while learned VLAs (OpenVLA, ฯ0) collapse to near-zero on these unseen tasks. The Reflection-through-Synthesis loop is the load-bearing idea: closing a validate-and-refine cycle over a generative world model converts an unreliable single-shot goal image into a usable plan, and the ablation shows most of the gain comes from it.
- arXiv: 2506.23919 ยท HTML ยท PDF
- Project page: nus-lins-lab.github.io/goalvlaweb
- Goal-Image-Conditioning review (editing-diffusion / world-model-as-prompt clusters)
- VLA Architectures review (hierarchical, training-free low-level category)
- ICRA 2026 Survey
โ Back to ICRA-2026