Review Goal Image Conditioning - Heungwoo/research GitHub Wiki
Compiled April 2026 Β· Focus: every 2023β2026 VLA / foundation-model robot policy that feeds a goal or subgoal image into the low-level controller, compared axis-by-axis against the Ο0.7 recipe.
Companion reviews: Review-VLA-Architecture Β· Review-VLM-Action-Connection Β· Review-pi07 Β· ML-Attention.
Conditioning a policy on a picture of where things should end up is a much stronger signal than conditioning on a text description of where things should end up. The reason is concrete: a goal image pins down object identity, 3D layout, occlusion, and grasp geometry at a pixel level that no sentence can match. NeurIPS 2025's VLA-OS did the controlled ablation and found, unsurprisingly, that visual-grounded planning beats language-grounded planning across data efficiency and quality.
The question is no longer "should we use goal images?" β it's "where does the goal image come from, how does it enter the policy, and how often is it refreshed?" There are now at least 15 published answers, spanning from SuSIE (ICLR 2024, InstructPix2Pix editing) to Ο0.7 (Apr 2026, BAGEL-14B world model + dropout + CFG).
This page lays out that axis, clusters the papers by mechanism, and shows where Ο0.7's specific design decisions sit relative to every other option.
Every goal-image-conditioned policy picks a point in four largely orthogonal axes:
flowchart TB
AXES["Goal-image VLA design space"]
AXES --> A1["Axis 1 β Source<br/>Who produces the goal?"]
AXES --> A2["Axis 2 β Consumption<br/>How does the policy ingest it?"]
AXES --> A3["Axis 3 β Granularity<br/>Terminal vs. subgoal vs. video"]
AXES --> A4["Axis 4 β Cadence<br/>Once vs. per-step vs. async"]
A1 --> S1["Human / user"]
A1 --> S2["Image-editing diffusion<br/>InstructPix2Pix"]
A1 --> S3["Text-to-video diffusion<br/>UniPi, RoboDreamer"]
A1 --> S4["Foundation world model<br/>BAGEL-14B in Ο0.7"]
A1 --> S5["Inline AR from same backbone<br/>CoT-VLA"]
A1 --> S6["Co-generated in joint diffusion<br/>dVLA, Unified Diffusion VLA"]
A2 --> C1["Channel-wise concat<br/>SuSIE"]
A2 --> C2["Prepended tokens<br/>GR-MG, Ο0.7"]
A2 --> C3["Cross-attention<br/>Gen2Act Flamingo-style"]
A2 --> C4["Inverse dynamics<br/>UniPi, HiP, VLP"]
A2 --> C5["Unified diffusion stream<br/>dVLA"]
A3 --> G1["Terminal goal only<br/>Chain-of-Action"]
A3 --> G2["Single near-term subgoal<br/>SuSIE, CoT-VLA, GR-MG"]
A3 --> G3["Full H-frame video<br/>UniPi, Gen2Act, RoboDreamer"]
A3 --> G4["Multi-view waypoints<br/>Ο0.7"]
A4 --> F1["Once per episode<br/>UniPi"]
A4 --> F2["Per subtask<br/>Gen2Act, HiP"]
A4 --> F3["Fixed step interval<br/>SuSIE every 20 steps"]
A4 --> F4["Receding horizon / MPC<br/>VLP, RoboDreamer"]
A4 --> F5["Async background thread<br/>Ο0.7 every β€4 seconds"]
A4 --> F6["Inline every action<br/>CoT-VLA, dVLA"]
Β§5 walks each paper through these axes. Β§6 is the Ο0.7 comparison table.
Bold = dedicated wiki page.
| # | Paper | Venue / year | Cluster | Goal source | Consumption |
|---|---|---|---|---|---|
| 1 | Ο0.7 | Apr 2026 | Foundation world-model prompt | BAGEL-14B world model | Prompt tokens + dropout + CFG + block-causal attention |
| 2 | SuSIE | ICLR 2024 | Editing-diffusion subgoal | InstructPix2Pix fine-tuned on video+robot | Channel-wise concat to obs β diffusion policy |
| 3 | GR-MG | 2024 (Aug) | Editing-diffusion subgoal | InstructPix2Pix + progress-percentage text | MAE-tokenized, prepended to token stream |
| 4 | UniPi | NeurIPS 2023 | Text-to-video + inverse dynamics | Text-to-video diffusion (cascaded) | Inverse-dynamics MLP on frame pairs |
| 5 | Gen2Act | Sept 2024 (Google DeepMind) | Human-video diffusion | Frozen VideoPoet (web pretrain) | Perceiver-Resampler + gated cross-attention |
| 6 | HiP | NeurIPS 2023 | LLM + video diffusion + inv-dyn | LLM subgoal text β video diffusion | Inverse dynamics (VC-1 init) |
| 7 | VLP | 2023 (Oct) | Receding-horizon video planning | Text-to-video + VLM as value | Short-horizon goal-conditioned policy |
| 8 | RoboDreamer | ICML 2024 | Compositional text-to-video | Compositional decomposed video diffusion | Inverse dynamics, closed-loop replan |
| 9 | GR-1 / GR-2 | ICLR 2024 / Oct 2024 | Implicit co-prediction | Joint video+action GPT | Shared tokens (not goal-image input) |
| 10 | VLMPC | RSS 2024 | MPC cost only | Human-supplied goal image | Scored by VLM in MPC loop; no learned policy input |
| 11 | CoT-VLA | CVPR 2025 | Inline AR subgoal | VILA-U itself (AR before action chunk) | Re-ingested by same VLA |
| 12 | dVLA | ICLR 2026 | Unified parallel diffusion | Same transformer (co-denoised) | Actions attend to being-denoised frames |
| 13 | Unified Diffusion VLA | ICLR 2026 | Unified parallel diffusion | Same transformer (co-denoised) | Joint masked discrete diffusion |
| 14 | VideoVLA | NeurIPS 2025 | VAM backbone | Video-gen backbone itself | Co-generated with actions |
| 15 | Chain-of-Action | NeurIPS 2025 | Backward from terminal goal | Predicted goal keyframe (same VLA) | Backward trajectory unroll |
Not a goal-image method but cited here as the empirical anchor: VLA-OS (NeurIPS 2025) β controlled ablation showing visual-grounded planning > language-grounded planning > action-only.
flowchart LR
subgraph A["A. Editing-diffusion subgoal<br/>generate the next picture from the current one"]
A1["SuSIE"]
A2["GR-MG"]
end
subgraph B["B. Text-to-video + inverse dynamics<br/>generate a video, decode frames into actions"]
B1["UniPi"]
B2["HiP"]
B3["VLP"]
B4["RoboDreamer"]
end
subgraph C["C. Human-video prior + cross-attention<br/>web-pretrained video generator, fused into policy"]
C1["Gen2Act"]
end
subgraph D["D. Unified / co-generated frames and actions<br/>same network, same diffusion loop"]
D1["CoT-VLA"]
D2["dVLA"]
D3["Unified Diffusion VLA"]
D4["VideoVLA"]
D5["GR-1 / GR-2"]
D6["Chain-of-Action"]
end
subgraph E["E. Foundation world model as prompt<br/>separate generator, conditioned policy<br/>β
where Ο0.7 sits"]
E1["Ο0.7 + BAGEL-14B"]
end
Fine-tune InstructPix2Pix on robot demonstrations to produce a near-future subgoal frame from the current observation and language. The low-level policy then reaches that frame.
- Pro: Strong pixel grounding; simple architecture; works with modest compute.
- Con: The editor is trained on robot data only β it doesn't generalize to unseen appliances or object categories. Performance collapses at distribution boundary.
- Ο0.7 departure: BAGEL-14B is web-pretrained (billions of natural images), so it generates plausible subgoals for air fryers, bagel toasters, espresso machines that no robot dataset has ever seen. This is the single biggest capability delta vs. editing-diffusion methods.
A text-to-video diffusion model generates an H-frame sequence; an inverse-dynamics MLP decodes each consecutive pair of frames into an action. The policy, strictly speaking, never sees the goal image β only the inverse-dynamics net does.
- Pro: Decouples "what should happen" from "how to actuate."
- Con: Inverse-dynamics drift is brutal β any pixel error compounds into actuation error. Open-loop variants (UniPi) fail on contact-rich tasks. Inference cost of a full text-to-video rollout per replan is high.
- Ο0.7 departure: Ο0.7 does not use inverse dynamics. The flow-matching action expert consumes subgoal tokens alongside observation and state, and directly outputs actions β no decoding step, no drift compounding.
Use a frozen human-video generative model (VideoPoet) to hallucinate a clip of a human performing the task, then encode that clip with a ViT + Perceiver-Resampler and fuse into the policy via Flamingo-style gated cross-attention.
- Pro: Leverages human-video web-scale pretraining for out-of-distribution tasks.
- Con: Gated cross-attention adds parameters and latency; the human-hand-to-robot-arm gap must be bridged by the policy.
- Ο0.7 departure: Ο0.7 also uses web-pretrained priors, but the world model generates robot-view (not human-view) subgoals, and consumption is via prompt tokens in the existing attention path rather than an added Flamingo-style adapter.
4.D Unified / co-generated frames and actions (CoT-VLA, dVLA, Unified Diffusion VLA, VideoVLA, GR-1/GR-2, Chain-of-Action)
Same transformer produces both frames and actions. Inline autoregressive (CoT-VLA), parallel discrete diffusion (dVLA, Unified Diffusion VLA), or joint video+action GPT (GR-1, GR-2). Chain-of-Action is the backward variant β predict goal keyframe, unroll trajectory backward.
- Pro: One network, one loss, no alignment drift between frame generator and action head.
- Con: Cannot reuse a foundation video model without re-training everything together. Cannot dropout subgoals at inference because they're entangled with actions. Inference latency of frame generation is on the critical path per step (CoT-VLA, partially mitigated by dVLA's parallel diffusion).
-
Ο0.7 departure: Ο0.7 is the architectural opposite. BAGEL runs asynchronously in a separate thread, refreshed whenever the semantic intent changes or after Ξ = 4 s, and the policy accepts
subgoal = dropped_outas a valid training mode. This is what makes inference latency acceptable despite a 14B auxiliary model.
Ο0.7's home cluster. A separate, web-pretrained, generative world model produces subgoals; the policy ingests them as prompt tokens alongside subtask text and episode metadata; each of these context components has its own dropout schedule at training time so any subset is valid at test time; CFG is applied on the metadata component to steer style; attention is block-causal so the fixed prefix (obs + subgoal) is shared across the 5 denoising steps.
This cluster currently has exactly one member for the full Ο0.7 recipe. Β§6 spells out why each of its six design choices is a clean departure from every other cluster.
π ICRA 2026 cousin β Goal-VLA: an image-generative VLM produces a goal state from which the target object pose is derived as the interface, separating a highly-generalizable VLM from training-free low-level control (with a Reflection-through-Synthesis validation loop). Same "generated goal as the interface" spirit, but the consumed signal is an object pose rather than a subgoal-image prompt β and it is zero-shot (59.9% on 8 RLBench tasks vs OpenVLA 0.2% / Ο0 0.0%).
Two design choices make Goal-VLA a sharp contrast to the Ο0.7 recipe, and worth dwelling on for anyone reasoning about this cluster. (1) The interface is a derived object pose, not the generated pixels. Where Ο0.7 (and SuSIE/GR-MG) hand the image to a learned policy, Goal-VLA collapses the generated goal image into a rigid object transform via feature matching + point-cloud registration, then drives a classical motion planner β so the generator's pixel errors only have to be good enough to localize a pose, not good enough for a policy to imitate frame-by-frame. This is the same "decouple what-from-how" logic as the inverse-dynamics cluster (Β§4.B), but the decoupling happens at the object-pose level rather than the frame-pair level, which sidesteps the inverse-dynamics drift that Β§4.B suffers. (2) The validate-and-refine loop is explicit and load-bearing. Ο0.7 trusts BAGEL's single forward pass (with dropout/CFG as the only hedges against a bad subgoal); Goal-VLA instead closes a Reflection-through-Synthesis loop over the world model, and its ablation attributes most of the gain to it (40.0% β 83.8% with the reflector, β 88.8% with three iterations). The price is that Goal-VLA carries no learned low-level policy at all, so it cannot acquire skills that geometry-plus-planning cannot express (compliant, contact-rich, or dynamic motions) β exactly the regime where Ο0.7's flow-matching action expert lives. The two papers therefore bracket the cluster: a trained generalist policy fed an optional subgoal image (Ο0.7) versus a training-free planner fed a validated goal pose (Goal-VLA).
- arXiv: 2310.10639
- Mechanism: InstructPix2Pix fine-tuned on BridgeData V2 + Something-Something V2 to "edit" the current frame into a future subgoal; a ResNet-50 + diffusion policy reaches it.
- Consumption: subgoal and current obs stacked channel-wise β ResNet-50 β diffusion head (4-action chunks, temporal ensembling).
- Cadence: every 20 env steps at deployment; training uses future frames 11β14 steps ahead.
- Granularity: single near-term subgoal, continuously refreshed.
- Limitation: editor trained on robot+video data only β no broad object prior.
- arXiv: 2408.14368
- Mechanism: InstructPix2Pix fine-tune + GPT-style policy. Novel trick: append "X % completed" to the text prompt driving subgoal generation, giving explicit progress conditioning.
- Consumption: MAE-tokenizes generated goal image, prepends to input token stream, all obs/state/action tokens attend to it.
- Numbers: CALVIN 3.35 β 4.04 rollout length; real-world 68.7 β 78.1 % standard, 44.4 β 60.6 % generalization.
- Limitation: same distribution-boundary issue as SuSIE.
- arXiv: 2302.00111
- Mechanism: text-to-video diffusion (first-frame tiled for consistency, cascaded temporal super-resolution) β inverse dynamics (3Γ3 conv + residual blocks β 7-DoF action).
- Consumption: frames are targets, not policy inputs β policy lives in the inverse-dynamics net.
- Cadence: generated once per task, open-loop execution.
- Limitation: no closed-loop correction; fragile under contact.
- arXiv: 2409.16283
- Mechanism: frozen VideoPoet (270 M-clip pretraining, text+image conditioned) zero-shot hallucinates a human performing the task.
- Consumption: ViT encoding of generated video β Perceiver-Resampler β Flamingo-style gated cross-attention into policy. Auxiliary point-track prediction loss.
- Cadence: once per subtask; chain by reusing last rollout frame.
- Strength: human-video web pretraining generalizes to new scenarios without robot data.
- arXiv: 2309.08587
- Mechanism: LLM β language subgoals β video diffusion β frame sequence β VC-1-initialized inverse dynamics.
- Cadence: regenerate per-subtask boundary.
- Limitation: no frame-level feedback loop; subgoals frozen within a subtask.
- arXiv: 2310.10625
- Mechanism: tree search in the space of (language subgoal, generated video); VLM scores both policy and value; text-to-video is the dynamics model.
- Consumption: a short-horizon goal-conditioned policy consumes the first 10 synthesized frames.
- Cadence: receding horizon, replan after N executions.
- arXiv: 2404.12377
- Mechanism: compositional text-to-video diffusion (AVDC + Imagen 3-stage); text parser decomposes the instruction into verb and PP primitives, each conditioning a separate diffusion component.
- Consumption: inverse dynamics on consecutive frames.
- Cadence: closed-loop periodic regeneration.
- arXiv: GR-1 2312.13139 Β· GR-2 2410.06158
- Mechanism: joint video+action GPT. Future frames are auxiliary outputs, not inputs that condition a separate action head.
- Cluster note: strictly speaking implicit conditioning via shared tokens β included here because the literature groups them with "goal-predictive" methods. If the question is "does it consume a goal image?" the honest answer is "no β it predicts one alongside actions."
- arXiv: 2407.09829
- Mechanism: MPC with a human-provided goal image used as a cost-function reference; a VLM scores predicted frames against the goal.
- Cluster note: no learned policy consumes the goal image β it's a planner cost, not a conditioning input. Included because the term "goal image" is used, but it is structurally different.
See CVPR 2025 survey. VILA-U itself generates one subgoal frame autoregressively before emitting the action chunk, then conditions on its own generation. +17 % real / +6 % sim over prior SOTA.
See ICLR-2026-dVLA. Parallel discrete denoising of text reasoning tokens + future image frames + action tokens in a single transformer. Resolves CoT latency penalties by co-denoising.
See ICLR-2026-Unified-Diffusion-VLA. Joint discrete denoising of masked future image tokens and masked action tokens under block-wise causal masking. The claim is "world model as policy" in a single network.
See NeurIPS-2025-VideoVLA. Multimodal DiT with a video-generation backbone (not a VLM), fine-tuned with control tokens injected; jointly predicts action chunks and future video frames.
See NeurIPS-2025-Chain-of-Action. Predicts a terminal goal keyframe, then unrolls the trajectory backward to the current state. Each action token answers "what step produces the next keyframe on the backward path?"
All rows below are grounded in the Ο0.7 long-form review. The right column is the contrast against every other method in this review.
| Aspect | Ο0.7 (Apr 2026) | What everyone else does |
|---|---|---|
| Goal-image producer | BAGEL-14B, a web-pretrained mixture-of-transformers image-editing / generation model, flow-matching-finetuned on robot data | InstructPix2Pix fine-tune (SuSIE, GR-MG); text-to-video (UniPi, HiP, VLP, RoboDreamer); VideoPoet (Gen2Act); same-transformer co-gen (CoT-VLA, dVLA, VideoVLA, GR-1/2) |
| Pretraining scope | Web-scale natural images β generalizes to unseen appliances (air fryer, bagel toaster, French press) | Robot+video data only (SuSIE, GR-MG); web video (Gen2Act, UniPi); robot only (unified-diffusion cluster) |
| Train-time presence | Subgoal images in 25 % of training examples; of those, 30 % also drop the subtask text so subgoal alone substitutes | Always present (SuSIE, GR-MG, Gen2Act); never present during training (VLMPC); shared across all steps (CoT-VLA, dVLA) |
| Consumption path | Prompt tokens concatenated with subtask text, episode metadata, and control-mode tokens | Channel-wise concat (SuSIE); prepended MAE tokens (GR-MG); Flamingo cross-attention (Gen2Act); inverse dynamics, never consumed by policy (UniPi, HiP); same attention as actions (unified-diffusion cluster) |
| Attention mask | Block-causal: obs and subgoal tokens bidirectional within themselves; subgoal attends obs; text causal after; action tokens bidirectional. The fixed prefix (obs + subgoal) lets the policy reuse that context across the 5 denoising steps used to generate the 50-step action chunk | Standard causal (most); special per-paper (dVLA adapts ordering; Unified Diffusion uses its own block masks) |
| Inference-time optionality | Any subset of (subtask, subgoal, metadata, control-mode) can be dropped at test time because every component has its own training dropout | SuSIE/GR-MG require the subgoal (no dropout); unified-diffusion methods cannot drop the frame branch; CoT-VLA can skip CoT but at quality cost |
| Guidance | Classifier-Free Guidance with Ξ² β {1.3, 1.7, 2.2} applied on the metadata component | No CFG in any other goal-image-conditioned VLA in this review |
| Regeneration cadence | Async in a separate thread, refreshed when semantic intent changes or after Ξ = 4 s | Once per episode (UniPi); per-subtask (Gen2Act, HiP); every 20 steps (SuSIE); receding horizon (VLP, RoboDreamer); inline per-action (CoT-VLA, dVLA, VideoVLA) |
| Producer runs on critical path? | No β async refresh hides the 14B cost | Yes for inline methods (CoT-VLA, dVLA, co-gen cluster); yes-but-rarely for episode-level methods (UniPi) |
| Producer parameter budget | 14 B (BAGEL) β ~2.8Γ the 5 B policy β but off critical path | 200 Mβ1 B for InstructPix2Pix-based (SuSIE, GR-MG); 1β7 B for text-to-video cascades; same-network overhead for unified-diffusion cluster |
| Number of subgoal views | Multi-view (each camera that the policy sees at inference can get its own generated subgoal) | Single view (all others); implicit-only for GR-1/2, co-gen cluster |
| Empirical support | Reverse-fridge-microwave β subgoals are critical when language alone is not enough. Referential instructions ("the fruit on the largest plate") improve with subgoal conditioning. Performance β equal with language-only vs. generated-subgoal prompting on most short tasks β subgoals are a capability lever for hard cases, not a latency-free lunch | Per-paper numbers vary; no other 2026 VLA has published ablations on this many dimensions |
flowchart LR
subgraph P7["Ο0.7 Apr 2026"]
OBS1["current obs"] --> POL1["policy β Gemma3-4B + flow action expert"]
SUBT1["subtask text"] --> POL1
META1["metadata β quality, speed, mistake"] --> POL1
BAGEL["BAGEL-14B world model<br/>async thread, refresh β€4 s"] --> SG1["subgoal images<br/>multi-view"]
SG1 --> POL1
POL1 --> ACT1["action chunk"]
note1["all four context channels<br/>can be independently dropped<br/>CFG on metadata"]
end
subgraph TYP["Typical editing-diffusion VLA<br/>SuSIE / GR-MG"]
OBS2["current obs"] --> IP2P["InstructPix2Pix fine-tune"]
LANG2["language"] --> IP2P
IP2P --> SG2["subgoal image"]
SG2 --> POL2["low-level policy"]
OBS2 --> POL2
POL2 --> ACT2["action chunk"]
note2["subgoal is mandatory<br/>no dropout, no CFG<br/>regenerated every ~20 steps"]
end
flowchart LR
L1["Loose coupling"]
R1["Tight coupling"]
L1 --> P1["Ο0.7 + BAGEL<br/>separate model<br/>async, dropoutable"]
P1 --> P2["Gen2Act<br/>separate model<br/>cross-attention"]
P2 --> P3["SuSIE, GR-MG<br/>separate editor<br/>mandatory"]
P3 --> P4["HiP, UniPi, VLP<br/>separate video gen<br/>inv-dynamics gap"]
P4 --> P5["CoT-VLA<br/>inline AR<br/>same network"]
P5 --> P6["dVLA, Unified Diffusion VLA<br/>co-denoised<br/>same loop"]
P6 --> R1
Ο0.7 is the loosest-coupled goal-image VLA published to date β and does so without giving up the benefit. The dropout + CFG recipe lets the world model be optional at inference, which no other method in this review supports.
- Goal images should be train-time optional. Without per-component dropout you cannot ship a policy that gracefully degrades when the generator is unavailable or slow. This is an operational point, not a benchmark point β it matters in deployment.
- Web-pretrained generators generalize further than robot-pretrained editors. Reverse-fridge-microwave is the smoking gun: InstructPix2Pix-class methods cannot synthesize "a microwave with nothing inside" as a target state from robot data alone.
- Async refresh sidesteps the generator-on-critical-path problem that has plagued CoT-VLA-style inline methods. A 14 B model becomes tolerable when you run it in a separate async thread, refreshed only on semantic-intent change or after Ξ = 4 s.
- It does not imply the co-generated (dVLA / Unified Diffusion VLA) cluster is wrong. Those methods achieve frame/action consistency by construction and avoid the train/test distribution mismatch that plagues Ο0.7's sampling recipe (25 % end-of-segment + 75 % uniform 0β4 s). For tasks where subgoal-action alignment dominates, co-generation may still win.
- It does not imply BAGEL-scale is necessary. The Ο0.7 paper itself flags "what's the minimum BAGEL-scale world model that retains the benefit?" as an open question. A 1β2 B editing-diffusion model with web pretraining could plausibly hit the same capability envelope.
- It does not imply multi-view subgoals always help. The paper only reports the multi-view setup on bimanual tasks; single-view tasks may not need it.
ICRA 2026's cohort is deployment- and sensor-centric (see ICRA VLA topic), and it does not contribute a new "imitate-a-generated-video" or "world-model-as-prompt" method in the Ο0.7 mold β there is no ICRA analogue of UniPi-style text-to-video rollout or of BAGEL-style subgoal-image prompting. What ICRA contributes to this axis is a single, pointed data point: Goal-VLA, the cluster's only zero-shot entry, which keeps the generated-goal-state idea but moves the conditioning interface from pixels to a derived object pose consumed by a training-free planner (detailed in Β§4.E above). The broader ICRA framing in ICRA VLA topic Β§7 is "goal flexibility" β conditioning on coordinates, images, routes, or synthesized states rather than language alone β but most of that energy is in navigation (OmniVLA-Nav composing pose/image/language goals), not in manipulation subgoal images.
The reading against the Ο0.7 recipe is that ICRA validates the generated-goal premise from the opposite extreme of the coupling spectrum (Β§6.2): Ο0.7 spends a 14B web-pretrained world model and a 5B learned policy to make the subgoal image an optional, dropoutable lever for hard cases; Goal-VLA spends zero action-labeled training and leans entirely on the generated goal plus a reflection loop, paying for it with a control stack that only classical geometry can express. Relatedly, ICRA's world-model interest shows up elsewhere as a reasoning-time (not goal-image) tool β VLA-Reasoner rolls out a 600M action-aware world model inside MCTS to score imagined outcomes β which is a different use of "imagined futures" than conditioning a policy on a generated goal, and is out of scope for this axis. Net: ICRA 2026 widens the low-resource, zero-shot corner of the design space rather than challenging Ο0.7's foundation-world-model-as-prompt point.
- Can you make the world model smaller? BAGEL-14B is ~2.8Γ the 5 B policy size. Would a web-pretrained 2 B image-editing diffusion model capture the "unseen-appliance" generalization that is the single biggest delta vs. SuSIE / GR-MG?
- Is there a "SuSIE with web-pretrained editor" sweet spot? InstructPix2Pix is already web-pretrained on natural edits. Finetuning it more gently (LoRA on robot data, rather than full fine-tune) might recover much of Ο0.7's capability without a separate 14 B model.
- Can the unified-diffusion cluster achieve dropoutability? dVLA and Unified Diffusion VLA bake the frame branch into the same diffusion step as actions. A variant that trains both co-generation and goal-dropped modes via stochastic masking could match Ο0.7's inference flexibility while keeping tight coupling.
- Is terminal-only (Chain-of-Action) competitive on long horizons? Chain-of-Action predicts one goal keyframe and unrolls backward. Against Ο0.7's sequence of intermediate multi-view subgoals, which wins on 10+ minute tasks like laundry folding?
- Does CFG on the subgoal channel (not just metadata) help? Ο0.7 applies CFG only on metadata. Extending CFG to the subgoal channel β pushing the policy harder toward actions consistent with the subgoal β is unexplored.
| If you are shipping⦠| Probably pick |
|---|---|
| A generalist VLA that must work on unseen object categories with an optional goal-image path | Ο0.7-style foundation-world-model-as-prompt + per-component dropout + async refresh |
| A task-specific policy on a single embodiment, no novel objects, minimal inference budget | SuSIE / GR-MG-style editing-diffusion subgoal with channel-wise or prepended-token consumption |
| A zero-shot policy that must generalize to scenarios with no robot data | Gen2Act-style human-video prior + cross-attention |
| A tightly-integrated VLA where frame/action consistency dominates, and latency is less critical | dVLA / Unified Diffusion VLA-style unified parallel diffusion |
| A policy on an unfamiliar long-horizon task where only the terminal state matters | Chain-of-Action-style backward unroll from a predicted goal keyframe |
| Do not use | Text-to-video + inverse dynamics (UniPi, HiP) β the inverse-dynamics gap is a real practical problem on contact-rich tasks; the cluster has fallen behind |
- Review-pi07 β the full Ο0.7 long-form review.
- Review-VLA-Architecture Β§5.E β the world-model / VAM architectural family.
- Review-VLM-Action-Connection β where prompt-token consumption (Ο0.7's choice) sits relative to cross-attention, same-stack MoE, FiLM.
- NeurIPS-2025-VLA-OS β the visual-grounded > language-grounded controlled ablation.
- CVPR-2025-VLA-Manipulation-Survey β for CoT-VLA
- NeurIPS-2025-VLA-Manipulation-Survey β for VideoVLA, DreamVLA, Chain-of-Action, VLA-OS
- ICLR-2026-VLA-Manipulation-Survey β for dVLA, Unified Diffusion VLA, Cosmos Policy, Ctrl-World, Ο0.7
β Back to Home