Review VLA Attention - Heungwoo/research GitHub Wiki

In-Depth Review β€” VLA Attention Architectures

Compiled May 2026 Β· Focus: per-family attention structure (mask pattern, cross vs self, block-causal details, softmax vs linear) for the four dominant VLA families of 2025–2026 β€” Ο€ series, GR00T series, StarVLA, and Qwen-VL-based VLAs.

This page is the attention-centric companion to the broader VLA Architectures review (which catalogs categories by action-head paradigm) and VLM↔Action Connection review (which catalogs interface mechanisms β€” KV-share / cross-attn / FiLM / latent / shared-params). Where those pages classify, this page diagrams β€” for each major family, what attention pattern is actually used, with block-causal masks drawn out and source citations to the paper section.

For the foundational primer on attention variants (causal, bidirectional, GQA-MQA-MHA, sliding-window, linear, gated), see ML-Attention. This page assumes you've read it.


1. TL;DR

The 2025–2026 VLA landscape has converged on softmax attention β€” no production VLA uses linear or gated-linear attention in the action path (RoboMamba SSM is the only mainstream exception, and it's small/efficient-class). The interesting variation is in mask patterns and VLM↔action coupling, not the attention kernel.

The 4 families separate cleanly:

  1. Ο€ series β€” same-stack MoE; conditional block-causal mask that switches between global bidirectional (no image goals) and block-causal with causal text (when subgoal images present).
  2. GR00T series β€” separate DiT with softmax cross-attention into truncated VLM features; DiT internally bidirectional. N1.6+ adds RDT-1B-style alternating image/text cross-attn.
  3. StarVLA β€” Lego-like; 4 head variants from pure causal AR (FAST) to layer-wise multi-layer cross-DiT (Ο€-style on Qwen).
  4. NORA / Qwen-VL-based β€” inherits Qwen2.5-VL's windowed ViT + full-causal LLM + M-RoPE; AR head uses the native causal mask, optional flow expert adds bidir on action chunk.

The single most important correction over earlier wiki notes: Ο€0.5 used global bidirectional attention (image + text), and Ο€0.7's "block-causal" mask is conditional on subgoal images / metadata being in the prompt β€” without them it falls back to Ο€0.5-style global bidir.


2. Quick comparison table

Model Action head location Mask pattern Bidir region Cross-attn? Linear? VLM backbone attn
Ο€0 Same stack (MoE) causal obs only βœ— βœ— PaliGemma full causal
Ο€0.5 Same stack (MoE) global bidirectional (image + text) image + text + actions βœ— βœ— full causal LLM, bidir mask added
Ο€0.6 Same stack (MoE 860M) (per Ο€0.7 App. B implication) global bidirectional image + text + actions βœ— βœ— Gemma3-4B (sliding window + global hybrid)
Ο€0.7 Same stack (MoE) conditional block-causal (see Β§3.1) obs + subgoal + actions; text causal if image goals present βœ— βœ— Gemma3-4B
GR00T N1 Separate DiT DiT bidir, VLM causal DiT internal (state + action) βœ“ softmax into truncated Eagle-2 (layer 12) βœ— Eagle-2 causal
GR00T N1.5 Separate DiT same same βœ“ into Eagle-2.5 last layer + adapter βœ— Eagle-2.5 causal (frozen)
GR00T N1.6 Separate DiT (32 layer) same + AlternateVLDiT (image-only / image+text every 2N blocks) DiT internal βœ“ + alternation βœ— Cosmos-Reason-2B causal
GR00T N1.7 Separate DiT same + vlln LN + vl_self_attention pre-DiT DiT internal βœ“ βœ— Qwen3-VL-2B (windowed ViT + causal LLM + M-RoPE)
NORA v1 Same stack AR Qwen2.5-VL native causal none βœ— βœ— windowed ViT + causal LLM + M-RoPE
NORA-1.5 + flow expert causal + bidir on action chunk action expert internal partial layer-wise βœ— Qwen2.5-VL
StarVLA-FAST Same stack AR Qwen native causal none βœ— βœ— Qwen2.5/3-VL
StarVLA-OFT MLP on hidden Qwen native causal none βœ— βœ— Qwen2.5/3-VL
StarVLA-Ο€ Layer-wise cross-DiT DiT bidir, VLM causal DiT internal βœ“ multi-layer (every Qwen layer β†’ DiT) βœ— Qwen3-VL
StarVLA-GR00T Separate DiT DiT bidir, VLM causal DiT internal βœ“ last-layer βœ— Qwen3-VL

3. Per-family deep dives

3.1 Ο€ series β€” same-stack MoE with conditional block-causal

The Ο€ series is architecturally one transformer: the VLM backbone (PaliGemma-3B β†’ Gemma3-4B) and the action expert (~300M β†’ 860M) sit in the same attention pool as a Mixture-of-Experts pair, with separate weights but a shared sequence and shared KV cache (prefix-KV reuse). The cross-version evolution is in the mask, not the topology.

3.1.1 The two attention modes of Ο€0.7 (per Appendix B)

Mode A β€” no image goals in prompt (Ο€0.7 falls back to Ο€0.5):

flowchart LR
  subgraph Seq[Token sequence]
    direction LR
    I[Image tokens]
    T[Text tokens<br/>task / subtask / metadata]
    A[Action tokens Γ— 50]
  end
  I <-->|bidir| T
  T <-->|bidir| A
  I <-->|bidir| A
  classDef img fill:#bbdefb,stroke:#1565c0,color:#000
  classDef txt fill:#fff9c4,stroke:#f57f17,color:#000
  classDef act fill:#c8e6c9,stroke:#2e7d32,color:#000
  class I img
  class T txt
  class A act
Loading

Mask matrix (token-by-block, all green = attend):

image text action
image βœ“ βœ“ βœ“
text βœ“ βœ“ βœ“
action βœ“ βœ“ βœ“

This is the pattern Ο€0.5 used. Paper Appendix B: "in absence of image goals we use the same attention patterns as in Ο€0.5, with global bidirectional attention between embeddings for all".

Mode B β€” with subgoal images / metadata (the new Ο€0.7 mask):

flowchart LR
  subgraph Seq[Token sequence β€” 4 blocks]
    direction LR
    O[Block 1: obs<br/>image + proprio]
    SG[Block 2: subgoal images]
    T[Block 3: text<br/>task / subtask / metadata]
    A[Block 4: action Γ— 50]
  end
  O <-->|bidir within| O
  SG <-->|bidir within| SG
  SG -->|attend| O
  T -->|causal within| T
  T -->|attend| O
  T -->|attend| SG
  A <-->|bidir within| A
  A -->|attend| O
  A -->|attend| SG
  A -->|attend| T
  classDef obs fill:#bbdefb,stroke:#1565c0,color:#000
  classDef sg fill:#ffcc80,stroke:#e65100,color:#000
  classDef txt fill:#fff9c4,stroke:#f57f17,color:#000
  classDef act fill:#c8e6c9,stroke:#2e7d32,color:#000
  class O obs
  class SG sg
  class T txt
  class A act
Loading

Block-level mask matrix (Q rows ↓, K cols β†’):

obs subgoal text action
obs BIDIR βœ— βœ— βœ—
subgoal βœ“ BIDIR βœ— βœ—
text βœ“ βœ“ CAUSAL βœ—
action βœ“ βœ“ βœ“ BIDIR

Legend: green BIDIR = bidirectional within block Β· green βœ“ = attend to prior block Β· orange CAUSAL = causal within block Β· red βœ— = blocked.

Paper Β§VI.B (Model architecture): "We employ a block-causal masking scheme, such that the observation tokens and the subgoal image tokens use bidirectional attention within themselves, and goal-image tokens can additionally attend the observations. The following text tokens use causal attention. [...] The 50 [action] tokens attend bidirectionally to each other and can also attend to the VLM backbone activations."

Why the switch? When the prompt grew from {task} (Ο€0/Ο€0.5) to {task, subtask, subgoal images, metadata, control mode} (Ο€0.7) with independent dropout per component + CFG on metadata, full bidirectional would let metadata leak into subtask, and CFG would have no defined "before/after". Causal text imposes an order that makes per-component dropout and CFG well-defined.

3.1.2 Ο€ series same-stack MoE topology

flowchart TB
  subgraph Prompt[Prompt tokens]
    direction TB
    O[Obs: 4cam Γ— 6hist 448Β²<br/>+ proprio]
    SG[Subgoal images<br/>BAGEL-14B generated]
    T[Text: task + subtask + metadata + ctrl]
  end

  subgraph SameStack[Single transformer stack β€” shared attention pool]
    direction TB
    L1[Layer 1: Gemma3-4B weights βŠ• Action-expert 860M weights<br/>shared KV cache, MoE routing per token type]
    L2[Layer 2: same]
    LN[Layer N: same]
    L1 --> L2 --> LN
  end

  Prompt --> L1
  N[Noise z] --> L1
  LN --> AT[50 action tokens<br/>flow matching, 5 Euler steps]
  AT --> Robot[Real robot 50Hz]

  KI[Knowledge Insulation:<br/>action gradient stops at MoE boundary,<br/>VLM weights protected]
  KI -.governs.-> L1

  classDef shared fill:#e3f2fd,stroke:#1565c0,color:#000
  class L1,L2,LN shared
Loading

Key point: there is no cross-attention in the Ο€ series. Action tokens and VLM tokens share the same attention layers (different MoE weights), so action queries reach VLM features via the same KV cache, not a separate cross-attention pass. This is what enables prefix-KV reuse β€” the VLM portion of the sequence is computed once per chunk, then 5 flow-matching denoising steps reuse the same KV.

3.1.3 Gemma3-4B's own attention (Ο€0.6/Ο€0.7 backbone)

Gemma3-4B uses a 5:1 sliding-window : global layer ratio internally. So the Ο€0.7 backbone has two levels of attention masking stacked:

  1. Gemma3 layer-internal: 5 of every 6 layers use sliding window, 1 in 6 is full global (Gemma3 design)
  2. Ο€ policy-level: the block-causal mask described above is applied on top of whichever Gemma3 layer is computing

Both use softmax. No linear attention anywhere.


3.2 GR00T series β€” separate DiT with cross-attention

GR00T's structural choice is the opposite of Ο€'s: two transformers. The VLM backbone (Eagle-2 β†’ Eagle-2.5 β†’ Cosmos-Reason β†’ Qwen3-VL) produces hidden states; a separate DiT (Diffusion Transformer) is the action head. The DiT cross-attends into the VLM hidden states at every block.

3.2.1 GR00T DiT block β€” internal attention

flowchart TB
  IN[Input: state tokens + action tokens<br/>concat as one sequence inside DiT]

  subgraph Block[One DiT block β€” BasicTransformerBlock]
    direction TB
    SA[1. Self-attention<br/>Q, K, V from state + action<br/>mask = BIDIRECTIONAL within chunk<br/>softmax]
    CA[2. Cross-attention<br/>Q from state + action<br/>K, V from VLM backbone features vl_embeds<br/>softmax]
    FF[3. FFN + AdaLN timestep t<br/>timestep injected via adaptive LayerNorm]
    SA --> CA --> FF
  end

  IN --> Block
  VLE[vl_embeds<br/>from truncated VLM] -.KV source.-> CA
  Block --> NEXT[to next DiT block]

  classDef sa fill:#c8e6c9,stroke:#2e7d32,color:#000
  classDef ca fill:#ffe0b2,stroke:#e65100,color:#000
  classDef ff fill:#e1bee7,stroke:#6a1b9a,color:#000
  class SA sa
  class CA ca
  class FF ff
Loading

So inside DiT: bidirectional self-attention on the action chunk (no causal mask β€” chunks are denoised in parallel) + softmax cross-attention into VLM features.

3.2.2 GR00T VLM-side preprocessing (the vlln + vl_self_attention evolution)

flowchart LR
  V[Multi-view images] --> ViT[Vision encoder]
  L[Language] --> Tok[Tokenizer]
  ViT --> LLM
  Tok --> LLM
  LLM[VLM LLM layers β€” physically truncated at select_layer]
  LLM --> BF[backbone_features]

  BF --> N1[N1: identity<br/>raw features]
  BF --> N15[N1.5: adapter MLP + LN]
  BF --> N16[N1.6: vlln LayerNorm]
  BF --> N17[N1.7: vlln + vl_self_attention<br/>SelfAttentionTransformer<br/>Q-former-lite]

  N1 -.cross-attn KV.-> DiT
  N15 -.cross-attn KV.-> DiT
  N16 -.cross-attn KV.-> DiT
  N17 -.cross-attn KV.-> DiT

  DiT[DiT action head]

  classDef new fill:#ffe8c2,stroke:#b47820,color:#000
  class N16,N17 new
Loading

The VLM is physically truncated with layers.pop() rather than tapping a mid-layer β€” N1's select_layer=12 actually drops layers 13+ and uses the new last layer. N1.5+ uses -1 (full stack last layer).

3.2.3 AlternateVLDiT (N1.6+) β€” RDT-1B-style alternating image/text

Problem: standard DiT has every block cross-attend to ALL VLM tokens (image + text). Image tokens (~hundreds) drown out text tokens (~tens) in the softmax.

Solution: every 2N-th block (default N=2 β†’ text every 4 blocks) attends text + image; the other blocks attend image only.

DiT block index 0 1 2 3 4 5 6 7 8 …
Cross-attn to TEXT βœ“ βœ— βœ— βœ— βœ“ βœ— βœ— βœ— βœ“ …
Cross-attn to IMAGE βœ“ βœ“ βœ“ βœ“ βœ“ βœ“ βœ“ βœ“ βœ“ …

Code reference (gr00t/model/modules/dit.py):

if idx % (2 * self.attend_text_every_n_blocks) == 0:
    # cross-attend to TEXT (+ image)
else:
    # cross-attend to IMAGE only

This is the direct analogue of RDT-1B's Alternating Condition Injection, adapted for VL inputs. Without alternation, image tokens overwhelm the text gradient signal.

3.2.4 GR00T full attention pipeline diagram

flowchart TB
  subgraph S2[System 2 β€” VLM]
    direction LR
    V[Images] --> ViT[ViT β€” windowed for Qwen3-VL N1.7,<br/>full softmax for Eagle-2/2.5]
    L[Language] --> Tk[Tokenizer]
    ViT --> LM[LLM causal softmax + M-RoPE for Qwen3-VL]
    Tk --> LM
  end

  LM --> BF[backbone_features]
  BF --> Pre[vlln + vl_self_attention<br/>N1.6/N1.7 only]
  Pre --> VLE[vl_embeds β€” cross-attn KV source]

  subgraph S1[System 1 β€” DiT 32 layers in N1.6/N1.7]
    direction TB
    SA[Self-attn:<br/>state + action<br/>BIDIRECTIONAL softmax]
    CA[Cross-attn:<br/>Q from state+action, KV from vl_embeds<br/>softmax + AlternateVLDiT alternation]
    FFN[FFN + AdaLN timestep injection]
    SA --> CA --> FFN
  end

  VLE -.KV source.-> CA
  ST[State history] --> SA
  AN[Noised action chunk Γ— 16] --> SA
  FFN --> OUT[Action chunk H=16<br/>flow matching K=4–5 steps]
Loading

Compared to Ο€ series: GR00T pays a separate cross-attn pass (extra latency, but VLM features stay clean), while Ο€ pays MoE branching cost (cheaper per pass but needs Knowledge Insulation to keep VLM gradients safe).


3.3 StarVLA β€” Lego-like, 4 attention variants on the same Qwen-VL backbone

StarVLA (arXiv 2604.05014) by Jiaya Jia's group is a modular codebase, not a single architecture. The same Qwen2.5-VL / Qwen3-VL / Qwen3.5 backbone can host 4 different attention/head configurations:

flowchart TB
  subgraph Backbone[Qwen-VL backbone β€” common to all 4 variants]
    ViT[ViT: windowed attn 28 local + 4 global, 2D RoPE per window]
    LLM[LLM: full softmax causal + M-RoPE multimodal RoPE]
    ViT --> LLM
  end

  LLM --> H1[Head: FAST AR tokens<br/>causal next-token]
  LLM --> H2[Head: OFT MLP<br/>parallel regression on placeholder hidden states]
  LLM --> H3[Head: Ο€ flow expert<br/>layer-wise cross-DiT]
  LLM --> H4[Head: GR00T-style DiT<br/>last-layer cross-attn]

  H1 --> A1[Action]
  H2 --> A2[Action]
  H3 --> A3[Action]
  H4 --> A4[Action]

  classDef variant fill:#f3e5f5,stroke:#6a1b9a,color:#000
  class H1,H2,H3,H4 variant
Loading

Per-variant attention pattern:

Variant Action paradigm Attention added on top of Qwen Mask
StarVLA-FAST AR FAST tokens (cross-entropy) none β€” uses Qwen native causal causal
StarVLA-OFT Parallel MLP regression on placeholder action token hidden states none β€” placeholder tokens enter Qwen causal stream causal
StarVLA-Ο€ Flow matching layer-wise cross-DiT: every Qwen layer β†’ LayerNorm + Linear projector β†’ DiT block cross-attends DiT bidir, VLM causal
StarVLA-GR00T Flow matching dual-system last-layer cross-attn (Γ  la GR00T) DiT bidir, VLM causal

The most distinctive variant is StarVLA-Ο€ because it differs from Ο€0/Ο€0.7 (same-stack) and from GR00T (last-layer cross-attn): it does multi-layer cross-attention from every Qwen layer, not just the last.

flowchart LR
  subgraph Qwen[Qwen3-VL backbone β€” every layer projected]
    direction TB
    L0[Layer 0] --> P0[LN + Linear projector]
    L1[Layer 1] --> P1[LN + Linear projector]
    L2[Layer 2] --> P2[LN + Linear projector]
    LD[…]
    LN[Layer N] --> PN[LN + Linear projector]
  end

  subgraph KV[Multi-layer KV pool]
    direction TB
    K0[KV from layer 0]
    K1[KV from layer 1]
    K2[KV from layer 2]
    KN[KV from layer N]
  end

  P0 --> K0
  P1 --> K1
  P2 --> K2
  PN --> KN

  subgraph DiT[Action DiT β€” each block cross-attends to a different VLM layer]
    direction TB
    D0[DiT block 0]
    D1[DiT block 1]
    D2[DiT block 2]
    DN[DiT block N]
  end

  K0 -.cross-attn KV.-> D0
  K1 -.cross-attn KV.-> D1
  K2 -.cross-attn KV.-> D2
  KN -.cross-attn KV.-> DN

  classDef vlm fill:#bbdefb,stroke:#1565c0,color:#000
  classDef proj fill:#fff9c4,stroke:#f57f17,color:#000
  classDef kv fill:#ffe0b2,stroke:#e65100,color:#000
  classDef dit fill:#c8e6c9,stroke:#2e7d32,color:#000
  class L0,L1,L2,LN,LD vlm
  class P0,P1,P2,PN proj
  class K0,K1,K2,KN kv
  class D0,D1,D2,DN dit
Loading

This is similar in spirit to ST4VLA's "k intermediate layers" approach (cited in Review-VLM-Action-Connection Β§8 as an open-question alternative to last-layer-only cross-attn).


3.4 NORA / Qwen-VL-based VLAs β€” inheriting Qwen2.5/3-VL attention

Qwen-VL-backed VLAs (NORA, NORA-1.5, anything else built on Qwen2.5-VL or Qwen3-VL) inherit the Qwen attention design wholesale. The interesting question is what the action head adds on top.

3.4.1 Qwen2.5-VL backbone attention (inherited by NORA, StarVLA, ChatVLA)

flowchart TB
  subgraph ViT[ViT β€” 32 layers total Β· native dynamic resolution]
    direction TB
    W1[Layers 0–6: windowed 8Γ—8 + 2D RoPE]
    G1[Layer 7: FULL GLOBAL]
    W2[Layers 8–14: windowed 8Γ—8 + 2D RoPE]
    G2[Layer 15: FULL GLOBAL]
    W3[Layers 16–22: windowed 8Γ—8 + 2D RoPE]
    G3[Layer 23: FULL GLOBAL]
    W4[Layers 24–30: windowed 8Γ—8 + 2D RoPE]
    G4[Layer 31: FULL GLOBAL]
    W1 --> G1 --> W2 --> G2 --> W3 --> G3 --> W4 --> G4
  end

  subgraph LLM[LLM β€” full SOFTMAX CAUSAL + M-RoPE]
    direction TB
    MR["M-RoPE β€” Multimodal Rotary PE<br/>text dim: 1D rotary<br/>image dim: 2D rotary (H + W separately)<br/>β†’ spatial structure preserved that 1D RoPE can't express"]
  end

  ViT --> LLM

  classDef win fill:#e1f5fe,stroke:#0277bd,color:#000
  classDef glob fill:#ffe0b2,stroke:#e65100,color:#000
  classDef llm fill:#c8e6c9,stroke:#2e7d32,color:#000
  class W1,W2,W3,W4 win
  class G1,G2,G3,G4 glob
  class MR llm
Loading

Pattern summary: 28 of 32 ViT layers use windowed local attention (8Γ—8 window with 2D RoPE per window), 4 layers (#7, 15, 23, 31) use full global attention. LLM uses full softmax causal with M-RoPE (text 1D + image H/W 2D rotary).

3.4.2 NORA v1 β€” pure AR, no new attention

flowchart LR
  subgraph Seq[Token sequence β€” single causal stream]
    direction LR
    I[Image patch tokens<br/>from ViT]
    T[Task text tokens]
    A[FAST action tokens<br/>new vocab extension]
  end
  I --> T --> A
  classDef img fill:#bbdefb,stroke:#1565c0,color:#000
  classDef txt fill:#fff9c4,stroke:#f57f17,color:#000
  classDef act fill:#c8e6c9,stroke:#2e7d32,color:#000
  class I img
  class T txt
  class A act
Loading

Causal mask over the entire sequence (standard Qwen). Action tokens are emitted left-to-right via next-token prediction:

image (K) text (K) action (K)
image (Q) CAUSAL within βœ— βœ—
text (Q) βœ“ CAUSAL within βœ—
action (Q) βœ“ βœ“ CAUSAL within

No new attention machinery. NORA's contribution is the FAST+ tokenizer + Qwen2.5-VL backbone choice, not a new attention mask.

3.4.3 NORA-1.5 β€” AR + flow expert with layer-wise self-attn

NORA-1.5 adds a 400M flow-matching action expert on top of NORA v1.

flowchart TB
  subgraph V1[NORA v1 stream β€” unchanged Qwen causal]
    direction LR
    I[image]
    T[text]
    F[FAST action tokens<br/>cross-entropy target]
    I --> T --> F
  end

  subgraph Expert[400M flow-matching action expert]
    direction TB
    EX[Expert tokens<br/>self-attn layer-wise with VLM hidden states]
    OUT[Continuous action chunk<br/>flow-matching loss]
    EX --> OUT
  end

  V1 -.layer-wise hidden states.-> EX

  classDef img fill:#bbdefb,stroke:#1565c0,color:#000
  classDef txt fill:#fff9c4,stroke:#f57f17,color:#000
  classDef fast fill:#ef9a9a,stroke:#b71c1c,color:#000
  classDef expert fill:#c8e6c9,stroke:#2e7d32,color:#000
  class I img
  class T txt
  class F fast
  class EX,OUT expert
Loading

Expert-side mask β€” FAST tokens are masked out to prevent leakage (expert can't peek at its own AR target):

Expert Q attends ↓ image hidden text hidden FAST hidden expert tokens
expert tokens (Q) βœ“ βœ“ βœ— MASKED BIDIR

This is structurally close to Ο€0.5's KI motivation β€” joint AR-on-FAST + flow-matching-on-continuous, with explicit attention masking to prevent the continuous head from cheating off the discrete FAST targets.

3.4.4 GR00T N1.7 (Qwen3-VL inherited)

GR00T N1.7 uses Cosmos-Reason2-2B, which is a Qwen3-VL-2B-Instruct derivative pretrained by NVIDIA on physical-AI scenes. So GR00T N1.7 also inherits Qwen3-VL's windowed-ViT + causal-LLM + M-RoPE + native-aspect-ratio, just consumed via cross-attention from a separate DiT (not native AR like NORA).

Key Qwen3-VL features GR00T N1.7 exploits:

  • Native aspect ratio β€” no image padding, useful for robot wrist cams with non-square aspect
  • M-RoPE β€” spatial structure preserved in cross-attn KV
  • Video pathway β€” used for EgoScale 20kh ego-video pretraining

4. Attention pattern primer (block-causal vs causal vs bidirectional vs prefix-LM)

Four canonical mask patterns, illustrated on the same 5-token sequence [tok 0, tok 1, tok 2, tok 3, tok 4]. Green cell = ATTEND, red cell = BLOCKED.

4.1 Causal mask (GPT-style)

kβ‚€ k₁ kβ‚‚ k₃ kβ‚„
qβ‚€ βœ“ βœ— βœ— βœ— βœ—
q₁ βœ“ βœ“ βœ— βœ— βœ—
qβ‚‚ βœ“ βœ“ βœ“ βœ— βœ—
q₃ βœ“ βœ“ βœ“ βœ“ βœ—
qβ‚„ βœ“ βœ“ βœ“ βœ“ βœ“

Lower-triangular: each query attends only to itself and prior positions. Used by GPT, OpenVLA, NORA v1, StarVLA-FAST/OFT, and the LLM portion of every modern VLM backbone (Qwen, Gemma, Eagle).

4.2 Bidirectional mask (BERT-style)

kβ‚€ k₁ kβ‚‚ k₃ kβ‚„
qβ‚€ βœ“ βœ“ βœ“ βœ“ βœ“
q₁ βœ“ βœ“ βœ“ βœ“ βœ“
qβ‚‚ βœ“ βœ“ βœ“ βœ“ βœ“
q₃ βœ“ βœ“ βœ“ βœ“ βœ“
qβ‚„ βœ“ βœ“ βœ“ βœ“ βœ“

All-to-all attention. Used by BERT, ViT, Ο€0.5 globally, and inside DiT action chunks (GR00T, StarVLA-Ο€).

4.3 Prefix-LM mask (PaliGemma-style)

Prefix = [tok 0, 1, 2] (bidir within), suffix = [tok 3, 4] (causal within, attends prefix).

kβ‚€ k₁ kβ‚‚ k₃ kβ‚„
qβ‚€ (prefix) βœ“ βœ“ βœ“ βœ— βœ—
q₁ (prefix) βœ“ βœ“ βœ“ βœ— βœ—
qβ‚‚ (prefix) βœ“ βœ“ βœ“ βœ— βœ—
q₃ (suffix) βœ“ βœ“ βœ“ βœ“ βœ—
qβ‚„ (suffix) βœ“ βœ“ βœ“ βœ“ βœ“

Used by PaliGemma (the Ο€0 backbone), which lets image+task act as a bidirectional prefix while generation is causal.

4.4 Block-causal mask (Ο€0.7 Mode B family)

Groups: obs = [0,1] (bidir), subgoal = [2] (bidir within, attends obs), text = [3] (causal within, attends prior), action = [4] (bidir within, attends all).

kβ‚€ obs k₁ obs kβ‚‚ subg k₃ text kβ‚„ act
qβ‚€ obs βœ“ βœ“ βœ— βœ— βœ—
q₁ obs βœ“ βœ“ βœ— βœ— βœ—
qβ‚‚ subgoal βœ“ βœ“ βœ“ βœ— βœ—
q₃ text βœ“ βœ“ βœ“ βœ“ βœ—
qβ‚„ action βœ“ βœ“ βœ“ βœ“ βœ“

Legend: dark green = bidir within block Β· light green = attend prior block Β· orange = causal within block Β· red = blocked.

The block-causal pattern is the most flexible β€” it lets you specify per-group mask semantics (bidir within, causal between, or attend-prior-only). Ο€0.7's Mode B is exactly this; PaliGemma's prefix-LM is a 2-block special case (one prefix bidir block + one suffix causal block).

Note that "block-causal" is not the same as causal. In a strict causal mask, position 4 cannot attend its own block bidirectionally β€” each query strictly looks at past positions only. Block-causal allows full bidir within each block while preserving a causal order between blocks.


5. VLM backbone attention characteristics (cheat sheet)

VLM Used in ViT attention LLM attention Special
PaliGemma-3B Ο€0 full softmax causal softmax + 1D RoPE SigLIP 400M vision
Gemma3-4B Ο€0.6 / Ο€0.7 full softmax 5:1 sliding-window : global hybrid + GQA Gemma3 mixed-attn
Eagle-2 (1.34B) GR00T N1 softmax causal softmax NVIDIA in-house
Eagle-2.5 (2.1B) GR00T N1.5 softmax causal softmax freezeable backbone
Cosmos-Reason-2B GR00T N1.6 softmax causal softmax CoT reasoning pretrain
Cosmos-Reason2-2B (= Qwen3-VL-2B) GR00T N1.7 windowed (28 local + 4 global) + native aspect ratio causal + M-RoPE + GQA + QK-Norm (Qwen3) 256k context, 2D/3D point loc
Qwen2.5-VL-3B NORA, StarVLA windowed (28 local + 4 global), 2D RoPE per window causal + M-RoPE + GQA native dynamic res
Qwen3-VL-{2,4,8}B StarVLA-Ξ±, GR00T N1.7 (via Cosmos-Reason2) windowed + native aspect ratio causal + M-RoPE + GQA + QK-Norm smaller sizes
Qwen3.5-{0.8,2,4,9}B StarVLA latest (2026/03) (Qwen3-VL inherits) causal + M-RoPE + 3:1 Gated DeltaNet : Gated Attention hybrid first VLA-deployed gated linear

Two notable backbone-level details:

  1. Gemma3's sliding-window-heavy hybrid in Ο€0.6/Ο€0.7 means the Ο€ policy already has efficient long-sequence attention even before MEM history compression β€” the sliding window covers local context, the 1-in-6 global layer carries the long-range signal.
  2. Qwen3.5's Gated DeltaNet (used by StarVLA's latest variant) is the first time a gated linear-recurrent attention has appeared in the VLA backbone. It's not yet evaluated head-to-head against full softmax for action generation. See ML-Attention for the gating mechanism.

6. Cross-cutting trade-offs

6.1 Same-stack (Ο€) vs cross-attention (GR00T) β€” the foundational choice

Same-stack MoE (Ο€) Cross-attn separate DiT (GR00T)
Latency per chunk Lower (KV cache shared) β€” Ο€0.6: 63ms / 50-chunk H100 Higher (separate DiT pass) β€” N1: 64ms / 16-chunk L40
VLM corruption risk High β€” needs Knowledge Insulation (stop-grad) Low β€” DiT is a separate parameter set
Architectural simplicity One transformer Two transformers + interface
Capacity scaling MoE expert size knob (Ο€0 300M β†’ Ο€0.7 860M) DiT depth knob (N1 12L β†’ N1.6 32L)
Where VLM features enter All layers (KV pool) Last-layer (or LN+self-attn preprocessed)

6.2 Last-layer vs multi-layer cross-attn

Approach Used by Trade-off
Last layer of (truncated) VLM only GR00T (all versions), most NORA-1.5 setups Simplest interface, but loses intermediate abstractions
Layer-wise from every VLM layer StarVLA-Ο€, ST4VLA Richer signal, more parameters in projectors, slower forward
Mid-layer tap (truncate at layer N) GR00T N1 (select_layer=12) Empirical compromise β€” N1.5+ abandoned this

6.3 AlternateVLDiT (RDT-1B / GR00T N1.6+) β€” necessary because of token count imbalance

The reason every-layer image+text cross-attn was a problem: in a typical batch, image tokens number in the hundreds (multi-view Γ— patches) while text tokens number in the tens. With softmax cross-attn, the gradient signal from text tokens is dominated by the softmax denominator pulled down by image attention scores. Alternation per-layer keeps text in the loop without diluting image-grounding.

6.4 Why no linear / gated attention in production VLAs (yet)

  • Action chunks are short (16–50 tokens). Linear attention's main win is O(N) over O(NΒ²), which doesn't matter when N=50.
  • VLM backbone tokens are the long sequence, but VLMs were pretrained with softmax β€” re-pretraining a backbone with linear attention is expensive and hasn't been done at production scale.
  • Qwen3.5's Gated DeltaNet is the first crack in this β€” a Qwen-3.5-backed VLA could eventually be the first linear-attention production VLA. As of May 2026, StarVLA's Qwen3.5 integration is the closest thing.

7. Open questions

  1. Does Ο€0.7's conditional block-causal hurt zero-shot generalization on prompts that mix image goals + free text? The mask switches at the prompt-construction boundary β€” has anyone ablated whether always-bidirectional or always-block-causal would be better than the conditional?
  2. Is layer-wise cross-attn (StarVLA-Ο€, ST4VLA) actually worth the projector parameters vs. last-layer cross-attn (GR00T)? No published controlled ablation matches param-count and data.
  3. Will gated-linear attention (Qwen3.5 / Gated DeltaNet) ever replace softmax for the VLM backbone of a production VLA? Action chunks won't benefit, but VLM-side prefill latency would.
  4. AlternateVLDiT's 2N cadence (default N=2 β†’ text every 4 blocks) is a hyperparameter that nobody has published a sweep for.
  5. Does causal text in Ο€0.7 actually need to be causal, or is it just "non-bidirectional"? Prefix-LM (PaliGemma-style) on the text block would be an in-between option β€” no one has tried.

8. Sources

Primary papers (read directly)

Companion wiki pages

  • ML-Attention β€” foundational primer on attention variants (causal / bidir / windowed / GQA / linear / gated)
  • Review-VLM-Action-Connection β€” interface mechanism taxonomy (KV-share / cross-attn / FiLM / latent / etc.)
  • Review-VLA-Architecture β€” action-head paradigm taxonomy (AR / flow / diffusion / VAM / hierarchical)
  • Review-pi07 β€” Ο€0.7 long-form review (the authoritative attention description per Β§4.1)
  • Review-GR00T-Series β€” GR00T N1β†’N1.7 code-level evolution (the authoritative cross-attn description)
  • pi-series-evolution β€” Ο€ series side-by-side architecture deltas (corrected attention row as of 2026-05-01)

πŸ—“ State of the Field (updated Aug 2026)

Verdict: attention design now follows the runtime (RTC/streaming/prefix-KV), not the backbone β€” and the biggest new variable, linear attention, is entirely unablated for control.

πŸ“ˆ Trend

Since late 2025 the mask and cache are chosen for RTC-compatibility, chunk streaming, and prefix-KV reuse rather than backbone convention. 2026 additions: hybrid linear-attention backbones entered VLAs via the Qwen suite (Qwen3.5's gated DeltaNet + interval GQA) with zero published ablation of the effect on control quality; and history/context conditioning was shown to interact with attention budgets (RobotManip's in-context variant needs 10 denoising steps where the base needs 4).

βš–οΈ Approaches & trade-offs

Design Pros Cons
Prefix-KV + parallel expert branch (Ο€) Cache reuse across denoise steps Backbone coupling
Full cross-attention re-read (GR00T / RobotManip) Fresh conditioning per chunk Re-encode cost each step
Streaming-token designs (Stream-to-Act-class) Continuous execution Young, little benchmark coverage
Hybrid linear attention (Qwen3.5-based VLAs) Long-context cost Control-quality impact unknown

ICML 2026 β€” the deployment cluster arrives. The strongest "make VLAs runnable" wave of any 2026 venue: Reflex exploits flow-timestep invariance to partition attention into static/sliding/dynamic regions with O(1) cache updates (2.58Γ— speedup, 50 Hz stable streaming); See What Matters's differentiable grid sampling cuts 76% of FLOPs with no success drop; SpecPrune-VLA (1.57–1.70Γ—) and EcoVLA (1.60Γ—, 2.18Γ— combined) prune training-free; and the XPU characterization study establishes the field's first hardware profile β€” VLM phase is compute-bound, action-expert phase is memory-bound β€” enabling 2.9–3.3Γ— via phase-aware scheduling.

⚠️ Limitations & open problems

  • Latency (ms/chunk) is the least-reported number in the field β€” see Review-Realtime-Execution for what little exists; ICML's deployment cluster is the first systematic push to close this.
  • Linear-attention-for-control is the largest open experimental gap on this axis.
  • KV-cache behavior under multi-view, high-rate vision remains folklore; no benchmark isolates the attention axis.

← Back to Home

⚠️ **GitHub.com Fallback** ⚠️