Review VLM Action Connection - Heungwoo/research GitHub Wiki
Compiled April 2026 · Focus: the specific architectural question "how does the VLM condition the action generator?" — KV sharing, cross-attention, latent tokens, FiLM, shared parameters, async scheduling, and more.
This is the sibling review to Review-VLA-Architecture (which groups VLAs by action-decoder family) — this page slices the same papers along a different axis: where and how information flows from the VLM into the action generator. Companion reviews: RL for VLA · VLA Memory · Cross-Embodiment.
Everyone talks about "VLM + action expert," but the coupling mechanism varies wildly across 2024–2026 VLAs. This review isolates seven distinct interface patterns, with concrete mechanism details:
- Unified token stream (OpenVLA, VLA-0, discrete-diffusion VLAs) — no interface; actions are tokens in the VLM's own vocabulary
- Same-stack MoE + prefix-KV attention (π-series) — action expert is separate weights in the same transformer stack; action tokens attend into the VLM's cached prefix KV
- Cross-attention into VLM hidden states (GR00T N1, RDT-1B, ST4VLA) — separate action transformer cross-attends to VLM features, possibly from a specific intermediate layer
- Latent condition token (ThinkAct, RoboDual, WholeBodyVLA) — VLM emits a small summary; action head consumes only that
- FiLM / prefix conditioning (CogVLA, classic Diffusion Policy) — features modulate the action stream via γ/β affine, without cross-attention
- Parameter sharing at layer boundary (Fast-in-Slow) — S1 is the last 2 transformer blocks of the VLM, re-run at higher frequency
- Outside the VLM (RFS residual, VITA-VLA distillation, RTC async scheduling) — interface unchanged; a wrapper at the action level adapts behavior
Key finding: no one has published a matched-compute head-to-head of these mechanisms under the same backbone and data. The field is bifurcating — π-series on same-stack MoE, ICLR 2026 on unified streams, NeurIPS 2025 dual-system on latent/embedded — without a clean comparison. That ablation is the most valuable unpublished paper in the space.
Independent of the action generator's objective (flow matching vs. diffusion vs. cross-entropy), the interface choice affects:
| Axis | What the interface controls |
|---|---|
| Latency | KV caching + layer reuse (same-stack / embedded) → fastest; cross-attn → extra compute; unified streams → depends on decoding strategy (AR slow, masked diffusion fast) |
| Gradient hygiene | Does the action-head gradient corrupt VLM features? Knowledge Insulation says "stop the backflow" |
| Pretraining preservation | Unified streams (VLA-0 extreme) and embedded sharing (FiS) inherit VLM pretraining directly; cascaded latent (ThinkAct) bottlenecks it; distillation (VITA-VLA) explicitly tries to transfer it |
| Backbone compatibility | Same-stack MoE needs matched head-dim + layer count; embedded needs knowable block structure; cross-attn is most portable |
| Bandwidth | Unified stream = high; cross-attn = medium (features, not just a vector); latent = low (one bottleneck) |
| Which VLM layer? | Unified: every layer, inherently. Cross-attn: you must choose — GR00T picks layer 12 of Eagle-2 empirically, ST4VLA queries k intermediate layers. |
| Frequency decoupling | Orthogonal to most choices. Embedded (FiS 1:4), cascaded latent (GR00T ~10:100 Hz), and async scheduling (RTC) each achieve it differently. |
flowchart TB
Q{Where does VLM information enter the action path?}
Q --> I1[1. Unified token stream<br/>same transformer, same vocab]
Q --> I2[2. Same-stack MoE + prefix KV<br/>shared layers, separate weights]
Q --> I3[3. Cross-attention into VLM hidden states<br/>separate action transformer]
Q --> I4[4. Latent condition token<br/>bottlenecked summary vector]
Q --> I5[5. FiLM / prefix conditioning<br/>γ·β modulation]
Q --> I6[6. Parameter sharing at layer boundary<br/>embedded S1-in-S2]
Q --> I7[7. Outside the VLM<br/>residual / distilled / async]
Core idea. Actions are tokens in the VLM's own vocabulary or token space. One transformer, one loss, no "interface."
Exemplars and sub-patterns:
| Paper | Sub-pattern | How actions enter the stream |
|---|---|---|
| OpenVLA (CoRL 2024) | AR discrete tokens | 256 bins in extra vocab, AR next-token prediction via LLaMA-2 decoder |
| RT-2 / RT-X / π0-FAST | AR with custom tokenizer | Same stream, FAST/DCT tokenizer compresses actions to fewer tokens |
| VLA-0 (2510.13054) | Actions as plain text | No new tokens at all — actions encoded as decimal numerals; Qwen2.5-VL-3B backbone untouched (no action head/expert). 94.7% avg LIBERO — beats all same-data methods (π0.5-KI, OpenVLA-OFT, SmolVLA) and, without large-scale robot pretraining, also beats large-data methods (π0, GR00T-N1, MolmoAct) |
| Discrete Diffusion VLA (2508.20072) | Masked discrete diffusion | Action tokens in VLM vocab; parallel unmasking with adaptive order. 96.3% LIBERO |
| Unified Diffusion VLA | Joint frame+action diffusion | Block-wise causal mask; action tokens attend to still-being-denoised future-image tokens |
| dVLA | Multimodal CoT in one stream | Discrete diffusion over future frames + text CoT + actions, all in parallel |
| HybridVLA (2503.10631) | Interleaved AR + diffusion | Single LLM; diffusion denoising interleaved into next-token prediction |
- ✅ Zero architectural surgery · inherits LLM training + serving tooling · preserves VLM knowledge most directly · no inter-module interface to hand-design
- ❌ AR decoding is sequential (slow) unless you use parallel decoding (OFT) or masked diffusion · discretization caps precision on continuous control · mixing text CoT + action in one stream can compete for attention capacity
- Pick when: simplest possible recipe, need interpretable unified stream, or want VLA-0-style minimal modification
Core idea. The action expert is a separate set of weights, but lives in the same transformer stack as the VLM. Action tokens attend to the VLM's prefix using shared attention-head geometry, and the VLM's prefix KV is cached at inference — only action tokens get recomputed per step.
Key constraints (why π0.6 and π0.7 have "860M expert, same layer count as backbone"):
- Matched attention-head dim and layer count between VLM and expert → required so the attention operation is compatible
- KV caching of the VLM prefix → production latency (π0.6: 63 ms / chunk on H100 with 5 Euler steps)
- Bidirectional attention among action tokens, but action tokens attend the VLM prefix through the standard KV interface
Exemplars:
| Paper | Expert size | Details |
|---|---|---|
| π0 (2410.24164) | Smaller hidden/MLP, matched heads | Canonical π-style MoE action expert; the template every π-series release inherits |
| π0.5 | Same | + autoregressive subtask text as prompt prefix; VLM re-forwards after subtask emission |
| π0.6 | 860M, same layer count | Backbone upgrade to Gemma3-4B; interface identical |
| π0.7 | Same 860M | Interface unchanged; all deltas are prompt-side (subgoal images, metadata, CFG) |
| FLOWER (2509.04996) | ~950M | Efficient variant; prunes 50% LLM layers + Global-AdaLN; 200 H100-hr pretraining |
| Knowledge Insulation | Same interface | Gradient modifier, not a new interface. Blocks flow-matching gradient from corrupting VLM weights |
- ✅ Smooth continuous actions · 5-step inference · cached prefix KV → fast · clean gradient hygiene with KI · production-deployed at PI scale
- ❌ Requires backbone-compatible architecture (head-dim, layer count) · action-head gradient corrupts VLM without Knowledge Insulation · more parameters than pure unified stream
- Pick when: production latency matters, continuous control is primary, you can design expert + backbone together
3. Cross-attention into VLM hidden states
Core idea. Action transformer is a separate network. Cross-attention layers attend to VLM token embeddings, often from a specific intermediate VLM layer rather than the last.
Exemplars with the non-obvious details:
| Paper | Where cross-attn attends | Non-obvious detail |
|---|---|---|
| GR00T N1 / N1.5 / N1.6 (2503.14734) | Layer 12 of Eagle-2 (intermediate, not last) | Chosen empirically for speed + success; alternating self-attn + cross-attn blocks Flamingo/VIMA-style |
| RDT-1B (2410.07864) | Frozen SigLIP + frozen T5-XXL | "Alternating Condition Injection" — image tokens and text tokens are cross-attended in alternating layers, not both every layer, because image would drown text. Explicitly rejects AdaLN because conditions are "high-dim and variable length" |
| ST4VLA (ICLR 2026) | k intermediate VLM layers | Query transformer cross-attends to k intermediate layers (not just the last) to stabilize expert learning |
| RetoVLA | + register-token KV injection | Injects discarded register tokens as auxiliary KV pairs — cheap global spatial context |
| DexVLA (2502.05855) | Plug-in ~1B diffusion expert | Conditions on VLM features; exact pathway less crisply specified in the paper |
| Cosmos Policy | Cosmos video backbone + control tokens | VLM (video backbone) features feed control-token decoders |
- ✅ Flexible — VLM and expert can have different sizes · can freeze VLM, train expert · can pick which layer's features matter (ST4VLA ablation)
- ❌ Extra cross-attention compute · interface layer choice is ad-hoc (GR00T's layer-12 is empirical) · image tokens can drown text (RDT's alternating trick exists to fix this)
- Pick when: mixing vendor VLMs with custom experts, frozen-VLM setups, or you want to pluggably swap action transformers
Core idea. VLM emits a small fixed-size summary token/vector; the action head consumes only that.
Exemplars:
| Paper | What the latent is |
|---|---|
| ThinkAct (NVIDIA, 2507.16815) | RL-rewarded MLLM plan compressed into a visual latent that conditions a separate action head |
| RoboDual (2410.08001) | Generalist VLA emits latent; specialist DiT conditions on it · +26.7% real vs OpenVLA · specialist only 20M params |
| WholeBodyVLA | Unified latent decodes to coordinated base / arms / hands |
| Hi-Robot | Hierarchical planner latent feeds low-level controller |
- ✅ Sharp frequency decoupling (S2 ~10 Hz, S1 ~100 Hz) · minimal bandwidth between systems · interface is small and easily cached · training curricula can be separated
- ❌ Bottleneck loses information · hand-designed latent shape · poor fine-grained visual grounding for contact tasks · separate training cadence
- Pick when: long-horizon planning with slow S2 + fast S1; plan caching matters; modular development with separate teams
Core idea. Instruction / plan features modulate (γ, β affine) the action stream at multiple layers. No explicit cross-attention.
Exemplars:
| Paper | Where FiLM is applied |
|---|---|
| CogVLA (2508.21046) | Twice — EFA-Routing applies FiLM at the vision encoder for token aggregation; LFP-Routing applies FiLM at the LLM for token pruning. Plus V-L-A Coupled Attention (causal V-L + bidirectional action parallel decoding). 97.4% LIBERO, 2.5× training / 2.8× inference speedup over OpenVLA |
| Classic Diffusion Policy | FiLM conditions the U-Net at every block |
| RDT-1B | Considers and rejects AdaLN — "lossy for high-dim variable-length conditions" |
- ✅ Parameter-efficient · no explicit cross-attn layers · composes beautifully with token pruning (CogVLA's 2.8× speedup) · classical robustness
- ❌ Lossy for long / variable-length conditions · weaker than cross-attention on complex prompts (empirical: RDT chose cross-attn over AdaLN for this reason)
- Pick when: small efficient VLAs, instruction-conditioned vision pruning, or the classical diffusion-policy recipe
Core idea. Action head and VLM share some transformer blocks; S1 is a subset of S2's layers re-run at higher frequency.
Exemplar:
-
Fast-in-Slow (2506.01953) — the canonical example, and currently the only one.
- Last 2 of 32 LLM blocks = S1 (Prismatic-VLM backbone: SigLIP+DINOv2 vision + LLaMA-2-7B LLM; ablated optimum — performance saturates at 2 of 32 shared blocks)
- S2 = full 32 blocks at 1/4 the frequency of S1 (1:4 ratio, ablated optimum)
- S1 sees extra modalities S2 doesn't: 3D point clouds (lightweight tokenizer + shared encoder), robot state, noised actions
- 117.7 Hz control on NVIDIA 4090 with chunk=8
- See Review-Fast-in-Slow for the full deep-dive
-
✅ S1 inherits VLM pretraining for free (shared weights) · single set of weights to maintain · naturally handles frequency decoupling · production-rate control
-
❌ Block count is a hyperparameter (2 optimal for LLaVA; untested on other backbones) · still passes a latent forward at S2→S1 boundary · backbone-specific
-
Pick when: you want dual-system benefits without maintaining two networks; dense (non-MoE) backbones; your backbone has a consistent block structure
Core idea. Don't modify the VLM↔action link. Add a residual adapter, distill a teacher into the VLM, or change the scheduling around inference.
Exemplars (three different patterns):
| Paper | What it adds | Mechanism |
|---|---|---|
| RFS (Residual Flow Steering) | Residual adapter | Base flow-matching policy frozen; small residual steering policy trained with RL; outputs summed with base flow field at the action-vector level |
| VITA-VLA (2510.09607) | Teacher-student distillation (reverse direction) | Distills a small pretrained action model INTO a 7B VLM via hidden-state alignment. Two stages: (1) alignment — map VLM hiddens to teacher's action space; (2) fine-tune. 97.3% LIBERO, 82.0% real. Teacher's decoder is reused |
| Real-Time Chunking (2506.07339) | Async scheduling | Plug-and-play on any diffusion/flow VLA, no retraining. Async chunk inpainting: freeze committed actions, inpaint the rest while previous chunk executes |
| PLD | Residual RL + distill | Frozen base, residual policy trained with RL, then distilled back into the base |
- ✅ Plug-and-play on frozen base · preserves base VLA knowledge · sidesteps the interface question entirely · scales to new behaviors without touching the backbone
- ❌ Inherits the base's ceiling · doesn't fix a poor VLM↔action coupling · works only if the base already solves the core problem
- Pick when: fine-tuning to new embodiments / tasks; contact-rich residual correction; serving an existing production VLA without retraining
The ICRA 2026 VLA cohort is deployment- and sensor-centric, and that shifts where it pushes on the interface. It contributes no new coupling primitive, but it stress-tests three of the seven along axes the ML venues ignore — chiefly what extra modality enters the VLM, and how to upgrade a frozen interface from reward.
A "sensor-token into the VLM prefix" variant of mechanism 1/5. FD-VLA is the cleanest example: a Force Distillation Module fuses a learnable query token over vision + robot state into a predicted force token that is injected into the pretrained VLM, distilled at training time against the latent of a real force/torque signal so the sensor is needed only during training. The injected token rides the VLM's own stream (mechanism 1) rather than cross-attending from outside — but unlike VLA-0's text tokens it carries a non-linguistic sensor channel, and the design is explicitly engineered to preserve the VLM's vision-language semantics (the force token augments, does not disrupt, the prefix). The same theme appears as FiLM in the cohort — Enhancing VLA Precision via FiLM-based Force/Torque-Vision Integration modulates the visual stream with F/T features (mechanism 5) instead of adding a prefix token. Together they show the interface question now includes which sensor modality gets wired in, and via which of the seven slots.
Mechanism 4, in production dual-system form. Galaxea / G0 is a textbook latent/subtask cascade: a System-2 G0-VLM (Qwen2.5-VL) emits a high-level subtask goal that conditions a separate System-1 G0-VLA (PaliGemma-3B + SigLIP, FAST tokenizer + flow matching). The hand-off is a low-bandwidth subtask interface between two independently-trained networks — the canonical mechanism-4 frequency-decoupling trade-off — and G0 is a real-robot, 23-DoF mobile-bimanual instantiation of it rather than a benchmark study.
Mechanism 7, sharpened for flow heads. ICRA's RL cluster targets the "outside-the-VLM" slot. FPO is a drop-in online-RL recipe that needs no architectural change: it builds a likelihood-free policy ratio from per-sample changes in the conditional-flow-matching loss the policy is already trained with, leaving the VLM↔action coupling (here π0's same-stack MoE) untouched. Like RFS and RTC, it adapts behavior around a frozen interface — extending the mechanism-7 pattern to reward-driven fine-tuning of flow-matching experts specifically. See the ICRA VLA topic page for the full cohort.
These don't define new interface patterns — they optimize existing ones:
| Paper | Contribution |
|---|---|
| VLA-Cache (2502.02175) | Reuse KV cache for visual tokens that don't change step-to-step |
| KV-Efficient VLA | Compresses historical KV cache via recurrent gating |
| VLA-Adapter (2509.09372) | "Bridge Attention" — learned selector that autonomously picks which VLM conditions to inject. 0.5B backbone, trains in 8 hr on 1 consumer GPU. Closest paper to an interface-ablation |
| VLA-OS (2506.17561) | Controlled paradigm study: Hierarchical > Integrated > Action-Only · visual-grounded > language-grounded planning |
| ST4VLA (ICLR 2026) | Querying transformer over k intermediate VLM layers (layer-selection ablation) |
| Axis | Winner | Runner-up |
|---|---|---|
| Lowest latency at production scale | 2. Same-stack MoE (π0.6 @ 63 ms / chunk) | 6. Embedded (FiS @ 117.7 Hz control) |
| Best pretraining preservation | 1. Unified stream VLA-0 (zero modification, 94.7% LIBERO) | 2. Same-stack + KI |
| Highest bandwidth VLM→action | 3. Cross-attention | 1. Unified stream |
| Sharpest frequency decoupling | 4. Latent condition (10 Hz S2 / 100 Hz S1) | 6. Embedded (1:4 ratio) |
| Most portable across VLM sizes | 3. Cross-attention (VLM + expert can differ) | 7. Outside-the-VLM (plug-and-play) |
| Most parameter-efficient | 5. FiLM (CogVLA 2.8× inference speedup) | 6. Embedded |
| Best interpretability | 4. Latent condition (inspectable summary) | 1. Unified stream (text CoT visible) |
| Cheapest to adapt to new task | 7. Outside-the-VLM (RFS/PLD residual) | 5. FiLM adapters |
| Single unifying objective | 1. Unified stream | — |
| Empirical best on LIBERO | 1. Discrete Diffusion VLA (96.3%) | 1. VLA-0 (94.7%, actions as text) beats π0.5-KI/OpenVLA-OFT/SmolVLA (same data) and π0/GR00T-N1/MolmoAct (large data) |
-
The π-style same-stack MoE + prefix-KV has become the production default for continuous control (π0→π0.5→π0.6→π0.7, FLOWER). Knowledge Insulation formalized the gradient hygiene that makes it stable. But it requires backbone-compatible architecture (head-dim, layer count), which locks you into specific VLM sizes.
-
ICLR 2026 pulled the interface back into the VLM via discrete diffusion (Discrete Diffusion VLA, Unified Diffusion VLA, dVLA). This collapses mechanisms 2 + 3 back to mechanism 1 by making the VLM itself generate actions through masked parallel denoising.
-
VLA-0's counter-swing: the simplest mechanism (actions as text, zero architectural change) hits 94.7% avg LIBERO — beating same-data methods (π0.5-KI, OpenVLA-OFT, SmolVLA) and, with no large-scale robot pretraining, large-data methods (π0, GR00T-N1, MolmoAct). Suggests the "right" interface may be to not have one.
-
NeurIPS 2025 diversified dual-system interfaces into 3 distinct patterns at one conference: cascaded latent (ThinkAct), MoE-routed (ChatVLA-2), embedded-shared (Fast-in-Slow). No consensus.
-
Frequency decoupling is now orthogonal to the interface choice. Real-Time Chunking works on any diffusion/flow VLA, no retraining. Fast-in-Slow's 1:4 ratio is a separate axis. You can pick any interface and layer async scheduling on top.
-
The field is NOT converging — it's bifurcating:
- Production continuous control → mechanism 2 (π-series) + gradient insulation
- Research / unified objective / interpretability → mechanism 1 (discrete diffusion, VLA-0)
- Dual-system / long-horizon → mechanisms 4 or 6
-
By ICRA 2026 the mechanism set has stabilized — the frontier moved from inventing wirings to deploying them. The Vienna cohort adds no eighth coupling primitive: its contributions slot into the existing seven and stress-test them on deployment/sensor/RL axes the ML venues under-explore — a sensor token distilled into the VLM prefix (FD-VLA, a variant of mechanism 1/5), a production S2→S1 subtask hand-off (Galaxea G0, mechanism 4 at 23-DoF), and architecture-free RL around a frozen flow interface (FPO, mechanism 7). The interesting question is no longer "which wiring?" but "how cheaply can I adapt a frozen one at deploy time?"
-
Matched-compute head-to-head of KV-sharing vs. cross-attention vs. latent-token vs. FiLM on the same backbone, same data, same action decoder. VLA-OS compared paradigms (Hierarchical vs. Integrated vs. Action-Only) but not interfaces. VLA-Adapter ablates which conditions matter, not how to inject them. This is the single most useful unpublished study.
-
Which VLM layer should cross-attention target? GR00T's layer 12 of Eagle-2 was chosen empirically for speed + success. ST4VLA uses k intermediate layers. No published layer-sweep on a standard benchmark.
-
Does KV-prefix sharing (π-style) actually preserve VLM knowledge better than cross-attention? The motivating claim for KI is "action gradients corrupt VLM features" — but that's about gradient flow, not attention flow. Does the same failure mode occur in GR00T-style cross-attention? Unknown.
-
Is the embedded pattern (Fast-in-Slow) an artifact of LLaVA's 32-block depth? "2 blocks optimal" on LLaVA. Completely untested on Gemma3-4B (π-series), Qwen-VL, Eagle-2. If the optimum varies, the design isn't transferable.
-
Does teacher→student distillation (VITA-VLA) preserve VLM reasoning better than joint training + KI? No head-to-head. VITA-VLA reports 97.3% LIBERO but doesn't run open-world reasoning preservation metrics à la ChatVLA-2 / VLM4VLA.
-
Is the interface question obsolete under unified token streams? If Discrete Diffusion VLA / VLA-0 close the performance gap on continuous control, mechanisms 2–5 may be legacy. Conversely, if flow matching remains Pareto-optimal on latency, unified-stream approaches need to match 63 ms on H100.
-
How does the interface interact with cross-embodiment transfer? Soft prompts (X-VLA), MoE (HiMoE-VLA), shared codebooks (XR-1) place heterogeneity at different interface sites. No paper ablates interface × embodiment heterogeneity jointly.
If you're designing a VLA and wondering which interface to pick:
- Shipping continuous control at <100 ms/chunk? → Mechanism 2 (same-stack MoE + prefix KV) with Knowledge Insulation. Matched head-dim and layer count between VLM and expert. This is π0.6 / π0.7.
- Want maximum VLM-pretraining preservation + simplest recipe? → Mechanism 1 (unified stream). Try VLA-0 (actions as text) before anything else. If you need parallelism, Discrete Diffusion VLA.
- Mixing a frozen vendor VLM with a custom action network? → Mechanism 3 (cross-attention). Pick an intermediate layer to attend to (following GR00T / ST4VLA).
- Long-horizon + sharp frequency decoupling? → Mechanism 4 (latent cascade) or mechanism 6 (embedded sharing). Embedded saves parameters; cascaded caches plans.
- Small / efficient / CPU edge? → Mechanism 5 (FiLM, CogVLA-style). Or mechanism 1 with a small backbone (SmolVLA).
- Adapting an existing production VLA without retraining? → Mechanism 7. RFS for RL-residuals, RTC for latency, VITA-VLA for distillation, MAP-VLA-style prompt-library retrieval for new tasks.
- Don't stack mechanisms thoughtlessly. Async scheduling (RTC) composes with any interface. FiLM and cross-attention can coexist. But same-stack MoE + cross-attention in the same model doesn't make sense — you'd have two interfaces to the same VLM.
Papers with primary interface innovation:
- π0 — https://arxiv.org/abs/2410.24164
- π0.5 — https://arxiv.org/abs/2504.16054
- π0.6 model card — https://website.pi-asset.com/pi06star/PI06_model_card.pdf
- π0.7 — https://www.pi.website/blog/pi07 · https://www.pi.website/download/pi07.pdf
- Knowledge Insulation — https://arxiv.org/abs/2505.23705
- RDT-1B — https://arxiv.org/abs/2410.07864
- DexVLA — https://arxiv.org/abs/2502.05855
- GR00T N1 — https://arxiv.org/abs/2503.14734
- Fast-in-Slow — https://arxiv.org/abs/2506.01953 · Review-Fast-in-Slow
- ChatVLA-2 — https://arxiv.org/abs/2505.21906
- ThinkAct — https://arxiv.org/abs/2507.16815
- Real-Time Chunking — https://arxiv.org/abs/2506.07339
- CogVLA — https://arxiv.org/abs/2508.21046
- VITA-VLA — https://arxiv.org/abs/2510.09607
- VLA-0 — https://arxiv.org/abs/2510.13054
- OpenVLA — https://arxiv.org/abs/2406.09246
- Discrete Diffusion VLA — https://arxiv.org/abs/2508.20072
- Unified Diffusion VLA — https://openreview.net/forum?id=UvQOcw2oCD
- VLA-Adapter (Bridge Attention) — https://arxiv.org/abs/2509.09372
- VLA-OS — https://arxiv.org/abs/2506.17561
- RoboDual — https://arxiv.org/abs/2410.08001
- HybridVLA — https://arxiv.org/abs/2503.10631
- VLA-Cache — https://arxiv.org/abs/2502.02175
Companion reviews:
- Review: VLA Architectures — same papers sliced by action-decoder family
- Review: VLA Memory · Cross-Embodiment · RL for VLA
- Per-paper: Review-pi07 · Review-pi06 · Review-VLM4VLA · Review-Fast-in-Slow
Verdict: first direct ablations favor last-layer cross-attention, but margins are ~0.5 pp — the wiring axis is live, not settled.
2025 ended in a four-way standoff, one pattern per lab (π's same-stack MoE + prefix-KV · GR00T's cross-attention · LBM's adaLN token · Qwen-VLA's concatenation). H1 2026 produced the first direct evidence: Qwen-RobotManip Table 19 (cross-attention 87.5 > concatenation 87.0 > layer-wise fusion 86.4 on LIBERO-Plus, at lowest cost) and Ψ₀'s hardware ablation (SD3-style MM-DiT dual modulation + joint attention > naive DiT head).
| Pattern | Champion | Pros | Cons | 2026 evidence |
|---|---|---|---|---|
| Same-stack MoE + prefix-KV | π0.6/0.7 | KV reuse, tight coupling | Backbone surgery, hard to retrofit | Production-proven; never directly ablated vs others |
| Cross-attention (last layer) | GR00T, RobotManip | Decoupled; ~500M experts suffice | One-way information flow | Wins both 2026 ablations |
| Concatenation + joint self-attn | Qwen-VLA | Simplest | Re-attends full VLM state per denoise step | Loses narrowly in its sibling's ablation |
| adaLN single token | TRI LBM | Cheapest | Information bottleneck | No 2026 head-to-head |
| MM-DiT dual-stream | Ψ₀ | Timestep modulates VL & action branches separately | Newest, least replicated | Beats naive DiT on a real humanoid |
| Mixture-of-Transformers dual-system | HALO, LaST₀ (ICML 2026) | Separates low-frequency reasoning from high-frequency action experts in one model | Two-rate coherence | HALO +34.1% over π0 on RoboTwin |
| Dual-expert phase routing | Move-Then-Operate (ICML 2026) | Coarse "move" vs contact-critical "operate" experts, learned selector | Phase-label supervision needed | +24% over monolithic π0 |
- Margins ~0.5 pp on single benchmarks; no matched-scale cross-lab comparison; the two Qwen flagships disagree internally with no reconciliation.
- Latency implications (prefix-KV reuse vs full re-attention) are argued, never measured.
- Ablations are manipulation-only; the ranking for whole-body/humanoid conditioning rests on Ψ₀'s single comparison.
- A second axis opened at ICML 2026 — what flows through the connection: LangForce's Bayesian decomposition shows naive wiring lets policies shortcut past language entirely (+11.3% OOD when countered) — wiring and grounding are not independent choices.
← Back to Home