Review VLM Action Connection - Heungwoo/research GitHub Wiki

In-Depth Review — The VLM↔Action-Expert Interface in VLAs

Compiled April 2026 · Focus: the specific architectural question "how does the VLM condition the action generator?" — KV sharing, cross-attention, latent tokens, FiLM, shared parameters, async scheduling, and more.

This is the sibling review to Review-VLA-Architecture (which groups VLAs by action-decoder family) — this page slices the same papers along a different axis: where and how information flows from the VLM into the action generator. Companion reviews: RL for VLA · VLA Memory · Cross-Embodiment.


1. TL;DR

Everyone talks about "VLM + action expert," but the coupling mechanism varies wildly across 2024–2026 VLAs. This review isolates seven distinct interface patterns, with concrete mechanism details:

  • Unified token stream (OpenVLA, VLA-0, discrete-diffusion VLAs) — no interface; actions are tokens in the VLM's own vocabulary
  • Same-stack MoE + prefix-KV attention (π-series) — action expert is separate weights in the same transformer stack; action tokens attend into the VLM's cached prefix KV
  • Cross-attention into VLM hidden states (GR00T N1, RDT-1B, ST4VLA) — separate action transformer cross-attends to VLM features, possibly from a specific intermediate layer
  • Latent condition token (ThinkAct, RoboDual, WholeBodyVLA) — VLM emits a small summary; action head consumes only that
  • FiLM / prefix conditioning (CogVLA, classic Diffusion Policy) — features modulate the action stream via γ/β affine, without cross-attention
  • Parameter sharing at layer boundary (Fast-in-Slow) — S1 is the last 2 transformer blocks of the VLM, re-run at higher frequency
  • Outside the VLM (RFS residual, VITA-VLA distillation, RTC async scheduling) — interface unchanged; a wrapper at the action level adapts behavior

Key finding: no one has published a matched-compute head-to-head of these mechanisms under the same backbone and data. The field is bifurcating — π-series on same-stack MoE, ICLR 2026 on unified streams, NeurIPS 2025 dual-system on latent/embedded — without a clean comparison. That ablation is the most valuable unpublished paper in the space.


2. What does the interface choice actually affect?

Independent of the action generator's objective (flow matching vs. diffusion vs. cross-entropy), the interface choice affects:

Axis What the interface controls
Latency KV caching + layer reuse (same-stack / embedded) → fastest; cross-attn → extra compute; unified streams → depends on decoding strategy (AR slow, masked diffusion fast)
Gradient hygiene Does the action-head gradient corrupt VLM features? Knowledge Insulation says "stop the backflow"
Pretraining preservation Unified streams (VLA-0 extreme) and embedded sharing (FiS) inherit VLM pretraining directly; cascaded latent (ThinkAct) bottlenecks it; distillation (VITA-VLA) explicitly tries to transfer it
Backbone compatibility Same-stack MoE needs matched head-dim + layer count; embedded needs knowable block structure; cross-attn is most portable
Bandwidth Unified stream = high; cross-attn = medium (features, not just a vector); latent = low (one bottleneck)
Which VLM layer? Unified: every layer, inherently. Cross-attn: you must choose — GR00T picks layer 12 of Eagle-2 empirically, ST4VLA queries k intermediate layers.
Frequency decoupling Orthogonal to most choices. Embedded (FiS 1:4), cascaded latent (GR00T ~10:100 Hz), and async scheduling (RTC) each achieve it differently.

3. The seven mechanisms at a glance

flowchart TB
  Q{Where does VLM information enter the action path?}
  Q --> I1[1. Unified token stream<br/>same transformer, same vocab]
  Q --> I2[2. Same-stack MoE + prefix KV<br/>shared layers, separate weights]
  Q --> I3[3. Cross-attention into VLM hidden states<br/>separate action transformer]
  Q --> I4[4. Latent condition token<br/>bottlenecked summary vector]
  Q --> I5[5. FiLM / prefix conditioning<br/>γ·β modulation]
  Q --> I6[6. Parameter sharing at layer boundary<br/>embedded S1-in-S2]
  Q --> I7[7. Outside the VLM<br/>residual / distilled / async]
Loading

4. Per-mechanism deep-dive

1. Unified token stream — no interface

Core idea. Actions are tokens in the VLM's own vocabulary or token space. One transformer, one loss, no "interface."

Exemplars and sub-patterns:

Paper Sub-pattern How actions enter the stream
OpenVLA (CoRL 2024) AR discrete tokens 256 bins in extra vocab, AR next-token prediction via LLaMA-2 decoder
RT-2 / RT-X / π0-FAST AR with custom tokenizer Same stream, FAST/DCT tokenizer compresses actions to fewer tokens
VLA-0 (2510.13054) Actions as plain text No new tokens at all — actions encoded as decimal numerals; Qwen2.5-VL-3B backbone untouched (no action head/expert). 94.7% avg LIBERO — beats all same-data methods (π0.5-KI, OpenVLA-OFT, SmolVLA) and, without large-scale robot pretraining, also beats large-data methods (π0, GR00T-N1, MolmoAct)
Discrete Diffusion VLA (2508.20072) Masked discrete diffusion Action tokens in VLM vocab; parallel unmasking with adaptive order. 96.3% LIBERO
Unified Diffusion VLA Joint frame+action diffusion Block-wise causal mask; action tokens attend to still-being-denoised future-image tokens
dVLA Multimodal CoT in one stream Discrete diffusion over future frames + text CoT + actions, all in parallel
HybridVLA (2503.10631) Interleaved AR + diffusion Single LLM; diffusion denoising interleaved into next-token prediction
  • ✅ Zero architectural surgery · inherits LLM training + serving tooling · preserves VLM knowledge most directly · no inter-module interface to hand-design
  • ❌ AR decoding is sequential (slow) unless you use parallel decoding (OFT) or masked diffusion · discretization caps precision on continuous control · mixing text CoT + action in one stream can compete for attention capacity
  • Pick when: simplest possible recipe, need interpretable unified stream, or want VLA-0-style minimal modification

2. Same-stack MoE + prefix-KV attention — the π-series pattern

Core idea. The action expert is a separate set of weights, but lives in the same transformer stack as the VLM. Action tokens attend to the VLM's prefix using shared attention-head geometry, and the VLM's prefix KV is cached at inference — only action tokens get recomputed per step.

Key constraints (why π0.6 and π0.7 have "860M expert, same layer count as backbone"):

  • Matched attention-head dim and layer count between VLM and expert → required so the attention operation is compatible
  • KV caching of the VLM prefix → production latency (π0.6: 63 ms / chunk on H100 with 5 Euler steps)
  • Bidirectional attention among action tokens, but action tokens attend the VLM prefix through the standard KV interface

Exemplars:

Paper Expert size Details
π0 (2410.24164) Smaller hidden/MLP, matched heads Canonical π-style MoE action expert; the template every π-series release inherits
π0.5 Same + autoregressive subtask text as prompt prefix; VLM re-forwards after subtask emission
π0.6 860M, same layer count Backbone upgrade to Gemma3-4B; interface identical
π0.7 Same 860M Interface unchanged; all deltas are prompt-side (subgoal images, metadata, CFG)
FLOWER (2509.04996) ~950M Efficient variant; prunes 50% LLM layers + Global-AdaLN; 200 H100-hr pretraining
Knowledge Insulation Same interface Gradient modifier, not a new interface. Blocks flow-matching gradient from corrupting VLM weights
  • ✅ Smooth continuous actions · 5-step inference · cached prefix KV → fast · clean gradient hygiene with KI · production-deployed at PI scale
  • ❌ Requires backbone-compatible architecture (head-dim, layer count) · action-head gradient corrupts VLM without Knowledge Insulation · more parameters than pure unified stream
  • Pick when: production latency matters, continuous control is primary, you can design expert + backbone together

3. Cross-attention into VLM hidden states

Core idea. Action transformer is a separate network. Cross-attention layers attend to VLM token embeddings, often from a specific intermediate VLM layer rather than the last.

Exemplars with the non-obvious details:

Paper Where cross-attn attends Non-obvious detail
GR00T N1 / N1.5 / N1.6 (2503.14734) Layer 12 of Eagle-2 (intermediate, not last) Chosen empirically for speed + success; alternating self-attn + cross-attn blocks Flamingo/VIMA-style
RDT-1B (2410.07864) Frozen SigLIP + frozen T5-XXL "Alternating Condition Injection" — image tokens and text tokens are cross-attended in alternating layers, not both every layer, because image would drown text. Explicitly rejects AdaLN because conditions are "high-dim and variable length"
ST4VLA (ICLR 2026) k intermediate VLM layers Query transformer cross-attends to k intermediate layers (not just the last) to stabilize expert learning
RetoVLA + register-token KV injection Injects discarded register tokens as auxiliary KV pairs — cheap global spatial context
DexVLA (2502.05855) Plug-in ~1B diffusion expert Conditions on VLM features; exact pathway less crisply specified in the paper
Cosmos Policy Cosmos video backbone + control tokens VLM (video backbone) features feed control-token decoders
  • ✅ Flexible — VLM and expert can have different sizes · can freeze VLM, train expert · can pick which layer's features matter (ST4VLA ablation)
  • ❌ Extra cross-attention compute · interface layer choice is ad-hoc (GR00T's layer-12 is empirical) · image tokens can drown text (RDT's alternating trick exists to fix this)
  • Pick when: mixing vendor VLMs with custom experts, frozen-VLM setups, or you want to pluggably swap action transformers

4. Latent condition token (cascaded dual-system)

Core idea. VLM emits a small fixed-size summary token/vector; the action head consumes only that.

Exemplars:

Paper What the latent is
ThinkAct (NVIDIA, 2507.16815) RL-rewarded MLLM plan compressed into a visual latent that conditions a separate action head
RoboDual (2410.08001) Generalist VLA emits latent; specialist DiT conditions on it · +26.7% real vs OpenVLA · specialist only 20M params
WholeBodyVLA Unified latent decodes to coordinated base / arms / hands
Hi-Robot Hierarchical planner latent feeds low-level controller
  • ✅ Sharp frequency decoupling (S2 ~10 Hz, S1 ~100 Hz) · minimal bandwidth between systems · interface is small and easily cached · training curricula can be separated
  • ❌ Bottleneck loses information · hand-designed latent shape · poor fine-grained visual grounding for contact tasks · separate training cadence
  • Pick when: long-horizon planning with slow S2 + fast S1; plan caching matters; modular development with separate teams

5. FiLM / prefix conditioning

Core idea. Instruction / plan features modulate (γ, β affine) the action stream at multiple layers. No explicit cross-attention.

Exemplars:

Paper Where FiLM is applied
CogVLA (2508.21046) Twice — EFA-Routing applies FiLM at the vision encoder for token aggregation; LFP-Routing applies FiLM at the LLM for token pruning. Plus V-L-A Coupled Attention (causal V-L + bidirectional action parallel decoding). 97.4% LIBERO, 2.5× training / 2.8× inference speedup over OpenVLA
Classic Diffusion Policy FiLM conditions the U-Net at every block
RDT-1B Considers and rejects AdaLN — "lossy for high-dim variable-length conditions"
  • ✅ Parameter-efficient · no explicit cross-attn layers · composes beautifully with token pruning (CogVLA's 2.8× speedup) · classical robustness
  • ❌ Lossy for long / variable-length conditions · weaker than cross-attention on complex prompts (empirical: RDT chose cross-attn over AdaLN for this reason)
  • Pick when: small efficient VLAs, instruction-conditioned vision pruning, or the classical diffusion-policy recipe

6. Parameter sharing at layer boundary — embedded S1-in-S2

Core idea. Action head and VLM share some transformer blocks; S1 is a subset of S2's layers re-run at higher frequency.

Exemplar:

  • Fast-in-Slow (2506.01953) — the canonical example, and currently the only one.

    • Last 2 of 32 LLM blocks = S1 (Prismatic-VLM backbone: SigLIP+DINOv2 vision + LLaMA-2-7B LLM; ablated optimum — performance saturates at 2 of 32 shared blocks)
    • S2 = full 32 blocks at 1/4 the frequency of S1 (1:4 ratio, ablated optimum)
    • S1 sees extra modalities S2 doesn't: 3D point clouds (lightweight tokenizer + shared encoder), robot state, noised actions
    • 117.7 Hz control on NVIDIA 4090 with chunk=8
    • See Review-Fast-in-Slow for the full deep-dive
  • ✅ S1 inherits VLM pretraining for free (shared weights) · single set of weights to maintain · naturally handles frequency decoupling · production-rate control

  • ❌ Block count is a hyperparameter (2 optimal for LLaVA; untested on other backbones) · still passes a latent forward at S2→S1 boundary · backbone-specific

  • Pick when: you want dual-system benefits without maintaining two networks; dense (non-MoE) backbones; your backbone has a consistent block structure

7. Outside the VLM — residual / distilled / async

Core idea. Don't modify the VLM↔action link. Add a residual adapter, distill a teacher into the VLM, or change the scheduling around inference.

Exemplars (three different patterns):

Paper What it adds Mechanism
RFS (Residual Flow Steering) Residual adapter Base flow-matching policy frozen; small residual steering policy trained with RL; outputs summed with base flow field at the action-vector level
VITA-VLA (2510.09607) Teacher-student distillation (reverse direction) Distills a small pretrained action model INTO a 7B VLM via hidden-state alignment. Two stages: (1) alignment — map VLM hiddens to teacher's action space; (2) fine-tune. 97.3% LIBERO, 82.0% real. Teacher's decoder is reused
Real-Time Chunking (2506.07339) Async scheduling Plug-and-play on any diffusion/flow VLA, no retraining. Async chunk inpainting: freeze committed actions, inpaint the rest while previous chunk executes
PLD Residual RL + distill Frozen base, residual policy trained with RL, then distilled back into the base
  • ✅ Plug-and-play on frozen base · preserves base VLA knowledge · sidesteps the interface question entirely · scales to new behaviors without touching the backbone
  • ❌ Inherits the base's ceiling · doesn't fix a poor VLM↔action coupling · works only if the base already solves the core problem
  • Pick when: fine-tuning to new embodiments / tasks; contact-rich residual correction; serving an existing production VLA without retraining

ICRA 2026 developments

The ICRA 2026 VLA cohort is deployment- and sensor-centric, and that shifts where it pushes on the interface. It contributes no new coupling primitive, but it stress-tests three of the seven along axes the ML venues ignore — chiefly what extra modality enters the VLM, and how to upgrade a frozen interface from reward.

A "sensor-token into the VLM prefix" variant of mechanism 1/5. FD-VLA is the cleanest example: a Force Distillation Module fuses a learnable query token over vision + robot state into a predicted force token that is injected into the pretrained VLM, distilled at training time against the latent of a real force/torque signal so the sensor is needed only during training. The injected token rides the VLM's own stream (mechanism 1) rather than cross-attending from outside — but unlike VLA-0's text tokens it carries a non-linguistic sensor channel, and the design is explicitly engineered to preserve the VLM's vision-language semantics (the force token augments, does not disrupt, the prefix). The same theme appears as FiLM in the cohort — Enhancing VLA Precision via FiLM-based Force/Torque-Vision Integration modulates the visual stream with F/T features (mechanism 5) instead of adding a prefix token. Together they show the interface question now includes which sensor modality gets wired in, and via which of the seven slots.

Mechanism 4, in production dual-system form. Galaxea / G0 is a textbook latent/subtask cascade: a System-2 G0-VLM (Qwen2.5-VL) emits a high-level subtask goal that conditions a separate System-1 G0-VLA (PaliGemma-3B + SigLIP, FAST tokenizer + flow matching). The hand-off is a low-bandwidth subtask interface between two independently-trained networks — the canonical mechanism-4 frequency-decoupling trade-off — and G0 is a real-robot, 23-DoF mobile-bimanual instantiation of it rather than a benchmark study.

Mechanism 7, sharpened for flow heads. ICRA's RL cluster targets the "outside-the-VLM" slot. FPO is a drop-in online-RL recipe that needs no architectural change: it builds a likelihood-free policy ratio from per-sample changes in the conditional-flow-matching loss the policy is already trained with, leaving the VLM↔action coupling (here π0's same-stack MoE) untouched. Like RFS and RTC, it adapts behavior around a frozen interface — extending the mechanism-7 pattern to reward-driven fine-tuning of flow-matching experts specifically. See the ICRA VLA topic page for the full cohort.

5. Bonus: Interface-efficiency work

These don't define new interface patterns — they optimize existing ones:

Paper Contribution
VLA-Cache (2502.02175) Reuse KV cache for visual tokens that don't change step-to-step
KV-Efficient VLA Compresses historical KV cache via recurrent gating
VLA-Adapter (2509.09372) "Bridge Attention" — learned selector that autonomously picks which VLM conditions to inject. 0.5B backbone, trains in 8 hr on 1 consumer GPU. Closest paper to an interface-ablation
VLA-OS (2506.17561) Controlled paradigm study: Hierarchical > Integrated > Action-Only · visual-grounded > language-grounded planning
ST4VLA (ICLR 2026) Querying transformer over k intermediate VLM layers (layer-selection ablation)

6. Cross-axis comparison

Axis Winner Runner-up
Lowest latency at production scale 2. Same-stack MoE (π0.6 @ 63 ms / chunk) 6. Embedded (FiS @ 117.7 Hz control)
Best pretraining preservation 1. Unified stream VLA-0 (zero modification, 94.7% LIBERO) 2. Same-stack + KI
Highest bandwidth VLM→action 3. Cross-attention 1. Unified stream
Sharpest frequency decoupling 4. Latent condition (10 Hz S2 / 100 Hz S1) 6. Embedded (1:4 ratio)
Most portable across VLM sizes 3. Cross-attention (VLM + expert can differ) 7. Outside-the-VLM (plug-and-play)
Most parameter-efficient 5. FiLM (CogVLA 2.8× inference speedup) 6. Embedded
Best interpretability 4. Latent condition (inspectable summary) 1. Unified stream (text CoT visible)
Cheapest to adapt to new task 7. Outside-the-VLM (RFS/PLD residual) 5. FiLM adapters
Single unifying objective 1. Unified stream —
Empirical best on LIBERO 1. Discrete Diffusion VLA (96.3%) 1. VLA-0 (94.7%, actions as text) beats π0.5-KI/OpenVLA-OFT/SmolVLA (same data) and π0/GR00T-N1/MolmoAct (large data)

7. Trends 2024 → 2026

  1. The π-style same-stack MoE + prefix-KV has become the production default for continuous control (π0→π0.5→π0.6→π0.7, FLOWER). Knowledge Insulation formalized the gradient hygiene that makes it stable. But it requires backbone-compatible architecture (head-dim, layer count), which locks you into specific VLM sizes.

  2. ICLR 2026 pulled the interface back into the VLM via discrete diffusion (Discrete Diffusion VLA, Unified Diffusion VLA, dVLA). This collapses mechanisms 2 + 3 back to mechanism 1 by making the VLM itself generate actions through masked parallel denoising.

  3. VLA-0's counter-swing: the simplest mechanism (actions as text, zero architectural change) hits 94.7% avg LIBERO — beating same-data methods (π0.5-KI, OpenVLA-OFT, SmolVLA) and, with no large-scale robot pretraining, large-data methods (π0, GR00T-N1, MolmoAct). Suggests the "right" interface may be to not have one.

  4. NeurIPS 2025 diversified dual-system interfaces into 3 distinct patterns at one conference: cascaded latent (ThinkAct), MoE-routed (ChatVLA-2), embedded-shared (Fast-in-Slow). No consensus.

  5. Frequency decoupling is now orthogonal to the interface choice. Real-Time Chunking works on any diffusion/flow VLA, no retraining. Fast-in-Slow's 1:4 ratio is a separate axis. You can pick any interface and layer async scheduling on top.

  6. The field is NOT converging — it's bifurcating:

    • Production continuous control → mechanism 2 (π-series) + gradient insulation
    • Research / unified objective / interpretability → mechanism 1 (discrete diffusion, VLA-0)
    • Dual-system / long-horizon → mechanisms 4 or 6
  7. By ICRA 2026 the mechanism set has stabilized — the frontier moved from inventing wirings to deploying them. The Vienna cohort adds no eighth coupling primitive: its contributions slot into the existing seven and stress-test them on deployment/sensor/RL axes the ML venues under-explore — a sensor token distilled into the VLM prefix (FD-VLA, a variant of mechanism 1/5), a production S2→S1 subtask hand-off (Galaxea G0, mechanism 4 at 23-DoF), and architecture-free RL around a frozen flow interface (FPO, mechanism 7). The interesting question is no longer "which wiring?" but "how cheaply can I adapt a frozen one at deploy time?"


8. Open questions — the ablations no one has done

  1. Matched-compute head-to-head of KV-sharing vs. cross-attention vs. latent-token vs. FiLM on the same backbone, same data, same action decoder. VLA-OS compared paradigms (Hierarchical vs. Integrated vs. Action-Only) but not interfaces. VLA-Adapter ablates which conditions matter, not how to inject them. This is the single most useful unpublished study.

  2. Which VLM layer should cross-attention target? GR00T's layer 12 of Eagle-2 was chosen empirically for speed + success. ST4VLA uses k intermediate layers. No published layer-sweep on a standard benchmark.

  3. Does KV-prefix sharing (π-style) actually preserve VLM knowledge better than cross-attention? The motivating claim for KI is "action gradients corrupt VLM features" — but that's about gradient flow, not attention flow. Does the same failure mode occur in GR00T-style cross-attention? Unknown.

  4. Is the embedded pattern (Fast-in-Slow) an artifact of LLaVA's 32-block depth? "2 blocks optimal" on LLaVA. Completely untested on Gemma3-4B (π-series), Qwen-VL, Eagle-2. If the optimum varies, the design isn't transferable.

  5. Does teacher→student distillation (VITA-VLA) preserve VLM reasoning better than joint training + KI? No head-to-head. VITA-VLA reports 97.3% LIBERO but doesn't run open-world reasoning preservation metrics à la ChatVLA-2 / VLM4VLA.

  6. Is the interface question obsolete under unified token streams? If Discrete Diffusion VLA / VLA-0 close the performance gap on continuous control, mechanisms 2–5 may be legacy. Conversely, if flow matching remains Pareto-optimal on latency, unified-stream approaches need to match 63 ms on H100.

  7. How does the interface interact with cross-embodiment transfer? Soft prompts (X-VLA), MoE (HiMoE-VLA), shared codebooks (XR-1) place heterogeneity at different interface sites. No paper ablates interface × embodiment heterogeneity jointly.


9. Practical decision guide

If you're designing a VLA and wondering which interface to pick:

  1. Shipping continuous control at <100 ms/chunk? → Mechanism 2 (same-stack MoE + prefix KV) with Knowledge Insulation. Matched head-dim and layer count between VLM and expert. This is π0.6 / π0.7.
  2. Want maximum VLM-pretraining preservation + simplest recipe? → Mechanism 1 (unified stream). Try VLA-0 (actions as text) before anything else. If you need parallelism, Discrete Diffusion VLA.
  3. Mixing a frozen vendor VLM with a custom action network? → Mechanism 3 (cross-attention). Pick an intermediate layer to attend to (following GR00T / ST4VLA).
  4. Long-horizon + sharp frequency decoupling? → Mechanism 4 (latent cascade) or mechanism 6 (embedded sharing). Embedded saves parameters; cascaded caches plans.
  5. Small / efficient / CPU edge? → Mechanism 5 (FiLM, CogVLA-style). Or mechanism 1 with a small backbone (SmolVLA).
  6. Adapting an existing production VLA without retraining? → Mechanism 7. RFS for RL-residuals, RTC for latency, VITA-VLA for distillation, MAP-VLA-style prompt-library retrieval for new tasks.
  7. Don't stack mechanisms thoughtlessly. Async scheduling (RTC) composes with any interface. FiLM and cross-attention can coexist. But same-stack MoE + cross-attention in the same model doesn't make sense — you'd have two interfaces to the same VLM.

10. Links

Papers with primary interface innovation:

Companion reviews:


🗓 State of the Field (updated Aug 2026)

Verdict: first direct ablations favor last-layer cross-attention, but margins are ~0.5 pp — the wiring axis is live, not settled.

📈 Trend

2025 ended in a four-way standoff, one pattern per lab (π's same-stack MoE + prefix-KV · GR00T's cross-attention · LBM's adaLN token · Qwen-VLA's concatenation). H1 2026 produced the first direct evidence: Qwen-RobotManip Table 19 (cross-attention 87.5 > concatenation 87.0 > layer-wise fusion 86.4 on LIBERO-Plus, at lowest cost) and Ψ₀'s hardware ablation (SD3-style MM-DiT dual modulation + joint attention > naive DiT head).

⚖️ Approaches & trade-offs

Pattern Champion Pros Cons 2026 evidence
Same-stack MoE + prefix-KV π0.6/0.7 KV reuse, tight coupling Backbone surgery, hard to retrofit Production-proven; never directly ablated vs others
Cross-attention (last layer) GR00T, RobotManip Decoupled; ~500M experts suffice One-way information flow Wins both 2026 ablations
Concatenation + joint self-attn Qwen-VLA Simplest Re-attends full VLM state per denoise step Loses narrowly in its sibling's ablation
adaLN single token TRI LBM Cheapest Information bottleneck No 2026 head-to-head
MM-DiT dual-stream Ψ₀ Timestep modulates VL & action branches separately Newest, least replicated Beats naive DiT on a real humanoid
Mixture-of-Transformers dual-system HALO, LaST₀ (ICML 2026) Separates low-frequency reasoning from high-frequency action experts in one model Two-rate coherence HALO +34.1% over π0 on RoboTwin
Dual-expert phase routing Move-Then-Operate (ICML 2026) Coarse "move" vs contact-critical "operate" experts, learned selector Phase-label supervision needed +24% over monolithic π0

⚠️ Limitations & open problems

  • Margins ~0.5 pp on single benchmarks; no matched-scale cross-lab comparison; the two Qwen flagships disagree internally with no reconciliation.
  • Latency implications (prefix-KV reuse vs full re-attention) are argued, never measured.
  • Ablations are manipulation-only; the ranking for whole-body/humanoid conditioning rests on Ψ₀'s single comparison.
  • A second axis opened at ICML 2026 — what flows through the connection: LangForce's Bayesian decomposition shows naive wiring lets policies shortcut past language entirely (+11.3% OOD when countered) — wiring and grounding are not independent choices.

← Back to Home

⚠️ **GitHub.com Fallback** ⚠️