Review VLA Attention - Heungwoo/research GitHub Wiki
Compiled May 2026 Β· Focus: per-family attention structure (mask pattern, cross vs self, block-causal details, softmax vs linear) for the four dominant VLA families of 2025β2026 β Ο series, GR00T series, StarVLA, and Qwen-VL-based VLAs.
This page is the attention-centric companion to the broader VLA Architectures review (which catalogs categories by action-head paradigm) and VLMβAction Connection review (which catalogs interface mechanisms β KV-share / cross-attn / FiLM / latent / shared-params). Where those pages classify, this page diagrams β for each major family, what attention pattern is actually used, with block-causal masks drawn out and source citations to the paper section.
For the foundational primer on attention variants (causal, bidirectional, GQA-MQA-MHA, sliding-window, linear, gated), see ML-Attention. This page assumes you've read it.
The 2025β2026 VLA landscape has converged on softmax attention β no production VLA uses linear or gated-linear attention in the action path (RoboMamba SSM is the only mainstream exception, and it's small/efficient-class). The interesting variation is in mask patterns and VLMβaction coupling, not the attention kernel.
The 4 families separate cleanly:
- Ο series β same-stack MoE; conditional block-causal mask that switches between global bidirectional (no image goals) and block-causal with causal text (when subgoal images present).
- GR00T series β separate DiT with softmax cross-attention into truncated VLM features; DiT internally bidirectional. N1.6+ adds RDT-1B-style alternating image/text cross-attn.
- StarVLA β Lego-like; 4 head variants from pure causal AR (FAST) to layer-wise multi-layer cross-DiT (Ο-style on Qwen).
- NORA / Qwen-VL-based β inherits Qwen2.5-VL's windowed ViT + full-causal LLM + M-RoPE; AR head uses the native causal mask, optional flow expert adds bidir on action chunk.
The single most important correction over earlier wiki notes: Ο0.5 used global bidirectional attention (image + text), and Ο0.7's "block-causal" mask is conditional on subgoal images / metadata being in the prompt β without them it falls back to Ο0.5-style global bidir.
| Model | Action head location | Mask pattern | Bidir region | Cross-attn? | Linear? | VLM backbone attn |
|---|---|---|---|---|---|---|
| Ο0 | Same stack (MoE) | causal | obs only | β | β | PaliGemma full causal |
| Ο0.5 | Same stack (MoE) | global bidirectional (image + text) | image + text + actions | β | β | full causal LLM, bidir mask added |
| Ο0.6 | Same stack (MoE 860M) | (per Ο0.7 App. B implication) global bidirectional | image + text + actions | β | β | Gemma3-4B (sliding window + global hybrid) |
| Ο0.7 | Same stack (MoE) | conditional block-causal (see Β§3.1) | obs + subgoal + actions; text causal if image goals present | β | β | Gemma3-4B |
| GR00T N1 | Separate DiT | DiT bidir, VLM causal | DiT internal (state + action) | β softmax into truncated Eagle-2 (layer 12) | β | Eagle-2 causal |
| GR00T N1.5 | Separate DiT | same | same | β into Eagle-2.5 last layer + adapter | β | Eagle-2.5 causal (frozen) |
| GR00T N1.6 | Separate DiT (32 layer) | same + AlternateVLDiT (image-only / image+text every 2N blocks) | DiT internal | β + alternation | β | Cosmos-Reason-2B causal |
| GR00T N1.7 | Separate DiT | same + vlln LN + vl_self_attention pre-DiT |
DiT internal | β | β | Qwen3-VL-2B (windowed ViT + causal LLM + M-RoPE) |
| NORA v1 | Same stack AR | Qwen2.5-VL native causal | none | β | β | windowed ViT + causal LLM + M-RoPE |
| NORA-1.5 | + flow expert | causal + bidir on action chunk | action expert internal | partial layer-wise | β | Qwen2.5-VL |
| StarVLA-FAST | Same stack AR | Qwen native causal | none | β | β | Qwen2.5/3-VL |
| StarVLA-OFT | MLP on hidden | Qwen native causal | none | β | β | Qwen2.5/3-VL |
| StarVLA-Ο | Layer-wise cross-DiT | DiT bidir, VLM causal | DiT internal | β multi-layer (every Qwen layer β DiT) | β | Qwen3-VL |
| StarVLA-GR00T | Separate DiT | DiT bidir, VLM causal | DiT internal | β last-layer | β | Qwen3-VL |
The Ο series is architecturally one transformer: the VLM backbone (PaliGemma-3B β Gemma3-4B) and the action expert (~300M β 860M) sit in the same attention pool as a Mixture-of-Experts pair, with separate weights but a shared sequence and shared KV cache (prefix-KV reuse). The cross-version evolution is in the mask, not the topology.
Mode A β no image goals in prompt (Ο0.7 falls back to Ο0.5):
flowchart LR
subgraph Seq[Token sequence]
direction LR
I[Image tokens]
T[Text tokens<br/>task / subtask / metadata]
A[Action tokens Γ 50]
end
I <-->|bidir| T
T <-->|bidir| A
I <-->|bidir| A
classDef img fill:#bbdefb,stroke:#1565c0,color:#000
classDef txt fill:#fff9c4,stroke:#f57f17,color:#000
classDef act fill:#c8e6c9,stroke:#2e7d32,color:#000
class I img
class T txt
class A act
Mask matrix (token-by-block, all green = attend):
| image | text | action | |
|---|---|---|---|
| image | β | β | β |
| text | β | β | β |
| action | β | β | β |
This is the pattern Ο0.5 used. Paper Appendix B: "in absence of image goals we use the same attention patterns as in Ο0.5, with global bidirectional attention between embeddings for all".
Mode B β with subgoal images / metadata (the new Ο0.7 mask):
flowchart LR
subgraph Seq[Token sequence β 4 blocks]
direction LR
O[Block 1: obs<br/>image + proprio]
SG[Block 2: subgoal images]
T[Block 3: text<br/>task / subtask / metadata]
A[Block 4: action Γ 50]
end
O <-->|bidir within| O
SG <-->|bidir within| SG
SG -->|attend| O
T -->|causal within| T
T -->|attend| O
T -->|attend| SG
A <-->|bidir within| A
A -->|attend| O
A -->|attend| SG
A -->|attend| T
classDef obs fill:#bbdefb,stroke:#1565c0,color:#000
classDef sg fill:#ffcc80,stroke:#e65100,color:#000
classDef txt fill:#fff9c4,stroke:#f57f17,color:#000
classDef act fill:#c8e6c9,stroke:#2e7d32,color:#000
class O obs
class SG sg
class T txt
class A act
Block-level mask matrix (Q rows β, K cols β):
| obs | subgoal | text | action | |
|---|---|---|---|---|
| obs | BIDIR | β | β | β |
| subgoal | β | BIDIR | β | β |
| text | β | β | CAUSAL | β |
| action | β | β | β | BIDIR |
Legend: green BIDIR = bidirectional within block Β· green β = attend to prior block Β· orange CAUSAL = causal within block Β· red β = blocked.
Paper Β§VI.B (Model architecture): "We employ a block-causal masking scheme, such that the observation tokens and the subgoal image tokens use bidirectional attention within themselves, and goal-image tokens can additionally attend the observations. The following text tokens use causal attention. [...] The 50 [action] tokens attend bidirectionally to each other and can also attend to the VLM backbone activations."
Why the switch? When the prompt grew from {task} (Ο0/Ο0.5) to {task, subtask, subgoal images, metadata, control mode} (Ο0.7) with independent dropout per component + CFG on metadata, full bidirectional would let metadata leak into subtask, and CFG would have no defined "before/after". Causal text imposes an order that makes per-component dropout and CFG well-defined.
flowchart TB
subgraph Prompt[Prompt tokens]
direction TB
O[Obs: 4cam Γ 6hist 448Β²<br/>+ proprio]
SG[Subgoal images<br/>BAGEL-14B generated]
T[Text: task + subtask + metadata + ctrl]
end
subgraph SameStack[Single transformer stack β shared attention pool]
direction TB
L1[Layer 1: Gemma3-4B weights β Action-expert 860M weights<br/>shared KV cache, MoE routing per token type]
L2[Layer 2: same]
LN[Layer N: same]
L1 --> L2 --> LN
end
Prompt --> L1
N[Noise z] --> L1
LN --> AT[50 action tokens<br/>flow matching, 5 Euler steps]
AT --> Robot[Real robot 50Hz]
KI[Knowledge Insulation:<br/>action gradient stops at MoE boundary,<br/>VLM weights protected]
KI -.governs.-> L1
classDef shared fill:#e3f2fd,stroke:#1565c0,color:#000
class L1,L2,LN shared
Key point: there is no cross-attention in the Ο series. Action tokens and VLM tokens share the same attention layers (different MoE weights), so action queries reach VLM features via the same KV cache, not a separate cross-attention pass. This is what enables prefix-KV reuse β the VLM portion of the sequence is computed once per chunk, then 5 flow-matching denoising steps reuse the same KV.
Gemma3-4B uses a 5:1 sliding-window : global layer ratio internally. So the Ο0.7 backbone has two levels of attention masking stacked:
- Gemma3 layer-internal: 5 of every 6 layers use sliding window, 1 in 6 is full global (Gemma3 design)
- Ο policy-level: the block-causal mask described above is applied on top of whichever Gemma3 layer is computing
Both use softmax. No linear attention anywhere.
GR00T's structural choice is the opposite of Ο's: two transformers. The VLM backbone (Eagle-2 β Eagle-2.5 β Cosmos-Reason β Qwen3-VL) produces hidden states; a separate DiT (Diffusion Transformer) is the action head. The DiT cross-attends into the VLM hidden states at every block.
flowchart TB
IN[Input: state tokens + action tokens<br/>concat as one sequence inside DiT]
subgraph Block[One DiT block β BasicTransformerBlock]
direction TB
SA[1. Self-attention<br/>Q, K, V from state + action<br/>mask = BIDIRECTIONAL within chunk<br/>softmax]
CA[2. Cross-attention<br/>Q from state + action<br/>K, V from VLM backbone features vl_embeds<br/>softmax]
FF[3. FFN + AdaLN timestep t<br/>timestep injected via adaptive LayerNorm]
SA --> CA --> FF
end
IN --> Block
VLE[vl_embeds<br/>from truncated VLM] -.KV source.-> CA
Block --> NEXT[to next DiT block]
classDef sa fill:#c8e6c9,stroke:#2e7d32,color:#000
classDef ca fill:#ffe0b2,stroke:#e65100,color:#000
classDef ff fill:#e1bee7,stroke:#6a1b9a,color:#000
class SA sa
class CA ca
class FF ff
So inside DiT: bidirectional self-attention on the action chunk (no causal mask β chunks are denoised in parallel) + softmax cross-attention into VLM features.
flowchart LR
V[Multi-view images] --> ViT[Vision encoder]
L[Language] --> Tok[Tokenizer]
ViT --> LLM
Tok --> LLM
LLM[VLM LLM layers β physically truncated at select_layer]
LLM --> BF[backbone_features]
BF --> N1[N1: identity<br/>raw features]
BF --> N15[N1.5: adapter MLP + LN]
BF --> N16[N1.6: vlln LayerNorm]
BF --> N17[N1.7: vlln + vl_self_attention<br/>SelfAttentionTransformer<br/>Q-former-lite]
N1 -.cross-attn KV.-> DiT
N15 -.cross-attn KV.-> DiT
N16 -.cross-attn KV.-> DiT
N17 -.cross-attn KV.-> DiT
DiT[DiT action head]
classDef new fill:#ffe8c2,stroke:#b47820,color:#000
class N16,N17 new
The VLM is physically truncated with layers.pop() rather than tapping a mid-layer β N1's select_layer=12 actually drops layers 13+ and uses the new last layer. N1.5+ uses -1 (full stack last layer).
Problem: standard DiT has every block cross-attend to ALL VLM tokens (image + text). Image tokens (~hundreds) drown out text tokens (~tens) in the softmax.
Solution: every 2N-th block (default N=2 β text every 4 blocks) attends text + image; the other blocks attend image only.
| DiT block index | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | β¦ |
|---|---|---|---|---|---|---|---|---|---|---|
| Cross-attn to TEXT | β | β | β | β | β | β | β | β | β | β¦ |
| Cross-attn to IMAGE | β | β | β | β | β | β | β | β | β | β¦ |
Code reference (gr00t/model/modules/dit.py):
if idx % (2 * self.attend_text_every_n_blocks) == 0:
# cross-attend to TEXT (+ image)
else:
# cross-attend to IMAGE onlyThis is the direct analogue of RDT-1B's Alternating Condition Injection, adapted for VL inputs. Without alternation, image tokens overwhelm the text gradient signal.
flowchart TB
subgraph S2[System 2 β VLM]
direction LR
V[Images] --> ViT[ViT β windowed for Qwen3-VL N1.7,<br/>full softmax for Eagle-2/2.5]
L[Language] --> Tk[Tokenizer]
ViT --> LM[LLM causal softmax + M-RoPE for Qwen3-VL]
Tk --> LM
end
LM --> BF[backbone_features]
BF --> Pre[vlln + vl_self_attention<br/>N1.6/N1.7 only]
Pre --> VLE[vl_embeds β cross-attn KV source]
subgraph S1[System 1 β DiT 32 layers in N1.6/N1.7]
direction TB
SA[Self-attn:<br/>state + action<br/>BIDIRECTIONAL softmax]
CA[Cross-attn:<br/>Q from state+action, KV from vl_embeds<br/>softmax + AlternateVLDiT alternation]
FFN[FFN + AdaLN timestep injection]
SA --> CA --> FFN
end
VLE -.KV source.-> CA
ST[State history] --> SA
AN[Noised action chunk Γ 16] --> SA
FFN --> OUT[Action chunk H=16<br/>flow matching K=4β5 steps]
Compared to Ο series: GR00T pays a separate cross-attn pass (extra latency, but VLM features stay clean), while Ο pays MoE branching cost (cheaper per pass but needs Knowledge Insulation to keep VLM gradients safe).
StarVLA (arXiv 2604.05014) by Jiaya Jia's group is a modular codebase, not a single architecture. The same Qwen2.5-VL / Qwen3-VL / Qwen3.5 backbone can host 4 different attention/head configurations:
flowchart TB
subgraph Backbone[Qwen-VL backbone β common to all 4 variants]
ViT[ViT: windowed attn 28 local + 4 global, 2D RoPE per window]
LLM[LLM: full softmax causal + M-RoPE multimodal RoPE]
ViT --> LLM
end
LLM --> H1[Head: FAST AR tokens<br/>causal next-token]
LLM --> H2[Head: OFT MLP<br/>parallel regression on placeholder hidden states]
LLM --> H3[Head: Ο flow expert<br/>layer-wise cross-DiT]
LLM --> H4[Head: GR00T-style DiT<br/>last-layer cross-attn]
H1 --> A1[Action]
H2 --> A2[Action]
H3 --> A3[Action]
H4 --> A4[Action]
classDef variant fill:#f3e5f5,stroke:#6a1b9a,color:#000
class H1,H2,H3,H4 variant
Per-variant attention pattern:
| Variant | Action paradigm | Attention added on top of Qwen | Mask |
|---|---|---|---|
| StarVLA-FAST | AR FAST tokens (cross-entropy) | none β uses Qwen native causal | causal |
| StarVLA-OFT | Parallel MLP regression on placeholder action token hidden states | none β placeholder tokens enter Qwen causal stream | causal |
| StarVLA-Ο | Flow matching | layer-wise cross-DiT: every Qwen layer β LayerNorm + Linear projector β DiT block cross-attends | DiT bidir, VLM causal |
| StarVLA-GR00T | Flow matching dual-system | last-layer cross-attn (Γ la GR00T) | DiT bidir, VLM causal |
The most distinctive variant is StarVLA-Ο because it differs from Ο0/Ο0.7 (same-stack) and from GR00T (last-layer cross-attn): it does multi-layer cross-attention from every Qwen layer, not just the last.
flowchart LR
subgraph Qwen[Qwen3-VL backbone β every layer projected]
direction TB
L0[Layer 0] --> P0[LN + Linear projector]
L1[Layer 1] --> P1[LN + Linear projector]
L2[Layer 2] --> P2[LN + Linear projector]
LD[β¦]
LN[Layer N] --> PN[LN + Linear projector]
end
subgraph KV[Multi-layer KV pool]
direction TB
K0[KV from layer 0]
K1[KV from layer 1]
K2[KV from layer 2]
KN[KV from layer N]
end
P0 --> K0
P1 --> K1
P2 --> K2
PN --> KN
subgraph DiT[Action DiT β each block cross-attends to a different VLM layer]
direction TB
D0[DiT block 0]
D1[DiT block 1]
D2[DiT block 2]
DN[DiT block N]
end
K0 -.cross-attn KV.-> D0
K1 -.cross-attn KV.-> D1
K2 -.cross-attn KV.-> D2
KN -.cross-attn KV.-> DN
classDef vlm fill:#bbdefb,stroke:#1565c0,color:#000
classDef proj fill:#fff9c4,stroke:#f57f17,color:#000
classDef kv fill:#ffe0b2,stroke:#e65100,color:#000
classDef dit fill:#c8e6c9,stroke:#2e7d32,color:#000
class L0,L1,L2,LN,LD vlm
class P0,P1,P2,PN proj
class K0,K1,K2,KN kv
class D0,D1,D2,DN dit
This is similar in spirit to ST4VLA's "k intermediate layers" approach (cited in Review-VLM-Action-Connection Β§8 as an open-question alternative to last-layer-only cross-attn).
Qwen-VL-backed VLAs (NORA, NORA-1.5, anything else built on Qwen2.5-VL or Qwen3-VL) inherit the Qwen attention design wholesale. The interesting question is what the action head adds on top.
flowchart TB
subgraph ViT[ViT β 32 layers total Β· native dynamic resolution]
direction TB
W1[Layers 0β6: windowed 8Γ8 + 2D RoPE]
G1[Layer 7: FULL GLOBAL]
W2[Layers 8β14: windowed 8Γ8 + 2D RoPE]
G2[Layer 15: FULL GLOBAL]
W3[Layers 16β22: windowed 8Γ8 + 2D RoPE]
G3[Layer 23: FULL GLOBAL]
W4[Layers 24β30: windowed 8Γ8 + 2D RoPE]
G4[Layer 31: FULL GLOBAL]
W1 --> G1 --> W2 --> G2 --> W3 --> G3 --> W4 --> G4
end
subgraph LLM[LLM β full SOFTMAX CAUSAL + M-RoPE]
direction TB
MR["M-RoPE β Multimodal Rotary PE<br/>text dim: 1D rotary<br/>image dim: 2D rotary (H + W separately)<br/>β spatial structure preserved that 1D RoPE can't express"]
end
ViT --> LLM
classDef win fill:#e1f5fe,stroke:#0277bd,color:#000
classDef glob fill:#ffe0b2,stroke:#e65100,color:#000
classDef llm fill:#c8e6c9,stroke:#2e7d32,color:#000
class W1,W2,W3,W4 win
class G1,G2,G3,G4 glob
class MR llm
Pattern summary: 28 of 32 ViT layers use windowed local attention (8Γ8 window with 2D RoPE per window), 4 layers (#7, 15, 23, 31) use full global attention. LLM uses full softmax causal with M-RoPE (text 1D + image H/W 2D rotary).
flowchart LR
subgraph Seq[Token sequence β single causal stream]
direction LR
I[Image patch tokens<br/>from ViT]
T[Task text tokens]
A[FAST action tokens<br/>new vocab extension]
end
I --> T --> A
classDef img fill:#bbdefb,stroke:#1565c0,color:#000
classDef txt fill:#fff9c4,stroke:#f57f17,color:#000
classDef act fill:#c8e6c9,stroke:#2e7d32,color:#000
class I img
class T txt
class A act
Causal mask over the entire sequence (standard Qwen). Action tokens are emitted left-to-right via next-token prediction:
| image (K) | text (K) | action (K) | |
|---|---|---|---|
| image (Q) | CAUSAL within | β | β |
| text (Q) | β | CAUSAL within | β |
| action (Q) | β | β | CAUSAL within |
No new attention machinery. NORA's contribution is the FAST+ tokenizer + Qwen2.5-VL backbone choice, not a new attention mask.
NORA-1.5 adds a 400M flow-matching action expert on top of NORA v1.
flowchart TB
subgraph V1[NORA v1 stream β unchanged Qwen causal]
direction LR
I[image]
T[text]
F[FAST action tokens<br/>cross-entropy target]
I --> T --> F
end
subgraph Expert[400M flow-matching action expert]
direction TB
EX[Expert tokens<br/>self-attn layer-wise with VLM hidden states]
OUT[Continuous action chunk<br/>flow-matching loss]
EX --> OUT
end
V1 -.layer-wise hidden states.-> EX
classDef img fill:#bbdefb,stroke:#1565c0,color:#000
classDef txt fill:#fff9c4,stroke:#f57f17,color:#000
classDef fast fill:#ef9a9a,stroke:#b71c1c,color:#000
classDef expert fill:#c8e6c9,stroke:#2e7d32,color:#000
class I img
class T txt
class F fast
class EX,OUT expert
Expert-side mask β FAST tokens are masked out to prevent leakage (expert can't peek at its own AR target):
| Expert Q attends β | image hidden | text hidden | FAST hidden | expert tokens |
|---|---|---|---|---|
| expert tokens (Q) | β | β | β MASKED | BIDIR |
This is structurally close to Ο0.5's KI motivation β joint AR-on-FAST + flow-matching-on-continuous, with explicit attention masking to prevent the continuous head from cheating off the discrete FAST targets.
GR00T N1.7 uses Cosmos-Reason2-2B, which is a Qwen3-VL-2B-Instruct derivative pretrained by NVIDIA on physical-AI scenes. So GR00T N1.7 also inherits Qwen3-VL's windowed-ViT + causal-LLM + M-RoPE + native-aspect-ratio, just consumed via cross-attention from a separate DiT (not native AR like NORA).
Key Qwen3-VL features GR00T N1.7 exploits:
- Native aspect ratio β no image padding, useful for robot wrist cams with non-square aspect
- M-RoPE β spatial structure preserved in cross-attn KV
- Video pathway β used for EgoScale 20kh ego-video pretraining
Four canonical mask patterns, illustrated on the same 5-token sequence [tok 0, tok 1, tok 2, tok 3, tok 4]. Green cell = ATTEND, red cell = BLOCKED.
| kβ | kβ | kβ | kβ | kβ | |
|---|---|---|---|---|---|
| qβ | β | β | β | β | β |
| qβ | β | β | β | β | β |
| qβ | β | β | β | β | β |
| qβ | β | β | β | β | β |
| qβ | β | β | β | β | β |
Lower-triangular: each query attends only to itself and prior positions. Used by GPT, OpenVLA, NORA v1, StarVLA-FAST/OFT, and the LLM portion of every modern VLM backbone (Qwen, Gemma, Eagle).
| kβ | kβ | kβ | kβ | kβ | |
|---|---|---|---|---|---|
| qβ | β | β | β | β | β |
| qβ | β | β | β | β | β |
| qβ | β | β | β | β | β |
| qβ | β | β | β | β | β |
| qβ | β | β | β | β | β |
All-to-all attention. Used by BERT, ViT, Ο0.5 globally, and inside DiT action chunks (GR00T, StarVLA-Ο).
Prefix = [tok 0, 1, 2] (bidir within), suffix = [tok 3, 4] (causal within, attends prefix).
| kβ | kβ | kβ | kβ | kβ | |
|---|---|---|---|---|---|
| qβ (prefix) | β | β | β | β | β |
| qβ (prefix) | β | β | β | β | β |
| qβ (prefix) | β | β | β | β | β |
| qβ (suffix) | β | β | β | β | β |
| qβ (suffix) | β | β | β | β | β |
Used by PaliGemma (the Ο0 backbone), which lets image+task act as a bidirectional prefix while generation is causal.
Groups: obs = [0,1] (bidir), subgoal = [2] (bidir within, attends obs), text = [3] (causal within, attends prior), action = [4] (bidir within, attends all).
| kβ obs | kβ obs | kβ subg | kβ text | kβ act | |
|---|---|---|---|---|---|
| qβ obs | β | β | β | β | β |
| qβ obs | β | β | β | β | β |
| qβ subgoal | β | β | β | β | β |
| qβ text | β | β | β | β | β |
| qβ action | β | β | β | β | β |
Legend: dark green = bidir within block Β· light green = attend prior block Β· orange = causal within block Β· red = blocked.
The block-causal pattern is the most flexible β it lets you specify per-group mask semantics (bidir within, causal between, or attend-prior-only). Ο0.7's Mode B is exactly this; PaliGemma's prefix-LM is a 2-block special case (one prefix bidir block + one suffix causal block).
Note that "block-causal" is not the same as causal. In a strict causal mask, position 4 cannot attend its own block bidirectionally β each query strictly looks at past positions only. Block-causal allows full bidir within each block while preserving a causal order between blocks.
| VLM | Used in | ViT attention | LLM attention | Special |
|---|---|---|---|---|
| PaliGemma-3B | Ο0 | full softmax | causal softmax + 1D RoPE | SigLIP 400M vision |
| Gemma3-4B | Ο0.6 / Ο0.7 | full softmax | 5:1 sliding-window : global hybrid + GQA | Gemma3 mixed-attn |
| Eagle-2 (1.34B) | GR00T N1 | softmax | causal softmax | NVIDIA in-house |
| Eagle-2.5 (2.1B) | GR00T N1.5 | softmax | causal softmax | freezeable backbone |
| Cosmos-Reason-2B | GR00T N1.6 | softmax | causal softmax | CoT reasoning pretrain |
| Cosmos-Reason2-2B (= Qwen3-VL-2B) | GR00T N1.7 | windowed (28 local + 4 global) + native aspect ratio | causal + M-RoPE + GQA + QK-Norm (Qwen3) | 256k context, 2D/3D point loc |
| Qwen2.5-VL-3B | NORA, StarVLA | windowed (28 local + 4 global), 2D RoPE per window | causal + M-RoPE + GQA | native dynamic res |
| Qwen3-VL-{2,4,8}B | StarVLA-Ξ±, GR00T N1.7 (via Cosmos-Reason2) | windowed + native aspect ratio | causal + M-RoPE + GQA + QK-Norm | smaller sizes |
| Qwen3.5-{0.8,2,4,9}B | StarVLA latest (2026/03) | (Qwen3-VL inherits) | causal + M-RoPE + 3:1 Gated DeltaNet : Gated Attention hybrid | first VLA-deployed gated linear |
Two notable backbone-level details:
- Gemma3's sliding-window-heavy hybrid in Ο0.6/Ο0.7 means the Ο policy already has efficient long-sequence attention even before MEM history compression β the sliding window covers local context, the 1-in-6 global layer carries the long-range signal.
- Qwen3.5's Gated DeltaNet (used by StarVLA's latest variant) is the first time a gated linear-recurrent attention has appeared in the VLA backbone. It's not yet evaluated head-to-head against full softmax for action generation. See ML-Attention for the gating mechanism.
| Same-stack MoE (Ο) | Cross-attn separate DiT (GR00T) | |
|---|---|---|
| Latency per chunk | Lower (KV cache shared) β Ο0.6: 63ms / 50-chunk H100 | Higher (separate DiT pass) β N1: 64ms / 16-chunk L40 |
| VLM corruption risk | High β needs Knowledge Insulation (stop-grad) | Low β DiT is a separate parameter set |
| Architectural simplicity | One transformer | Two transformers + interface |
| Capacity scaling | MoE expert size knob (Ο0 300M β Ο0.7 860M) | DiT depth knob (N1 12L β N1.6 32L) |
| Where VLM features enter | All layers (KV pool) | Last-layer (or LN+self-attn preprocessed) |
| Approach | Used by | Trade-off |
|---|---|---|
| Last layer of (truncated) VLM only | GR00T (all versions), most NORA-1.5 setups | Simplest interface, but loses intermediate abstractions |
| Layer-wise from every VLM layer | StarVLA-Ο, ST4VLA | Richer signal, more parameters in projectors, slower forward |
| Mid-layer tap (truncate at layer N) | GR00T N1 (select_layer=12) |
Empirical compromise β N1.5+ abandoned this |
The reason every-layer image+text cross-attn was a problem: in a typical batch, image tokens number in the hundreds (multi-view Γ patches) while text tokens number in the tens. With softmax cross-attn, the gradient signal from text tokens is dominated by the softmax denominator pulled down by image attention scores. Alternation per-layer keeps text in the loop without diluting image-grounding.
- Action chunks are short (16β50 tokens). Linear attention's main win is O(N) over O(NΒ²), which doesn't matter when N=50.
- VLM backbone tokens are the long sequence, but VLMs were pretrained with softmax β re-pretraining a backbone with linear attention is expensive and hasn't been done at production scale.
- Qwen3.5's Gated DeltaNet is the first crack in this β a Qwen-3.5-backed VLA could eventually be the first linear-attention production VLA. As of May 2026, StarVLA's Qwen3.5 integration is the closest thing.
- Does Ο0.7's conditional block-causal hurt zero-shot generalization on prompts that mix image goals + free text? The mask switches at the prompt-construction boundary β has anyone ablated whether always-bidirectional or always-block-causal would be better than the conditional?
- Is layer-wise cross-attn (StarVLA-Ο, ST4VLA) actually worth the projector parameters vs. last-layer cross-attn (GR00T)? No published controlled ablation matches param-count and data.
- Will gated-linear attention (Qwen3.5 / Gated DeltaNet) ever replace softmax for the VLM backbone of a production VLA? Action chunks won't benefit, but VLM-side prefill latency would.
-
AlternateVLDiT's
2Ncadence (defaultN=2β text every 4 blocks) is a hyperparameter that nobody has published a sweep for. - Does causal text in Ο0.7 actually need to be causal, or is it just "non-bidirectional"? Prefix-LM (PaliGemma-style) on the text block would be an in-between option β no one has tried.
- Ο0.7 paper β https://www.pi.website/download/pi07.pdf Β· Β§VI.B (Model architecture) + Appendix B (Attention pattern, Fig. 19) are the authoritative sources for the conditional block-causal mask. Direct quote in Β§VI.B: "We employ a block-causal masking scheme..."; Appendix B: "in absence of image goals we use the same attention patterns as in Ο0.5, with global bidirectional attention between embeddings for all".
- GR00T N1 paper β https://arxiv.org/abs/2503.14734 Β· Β§4 (architecture) for the original mid-layer cross-attn.
-
GR00T N1.5/1.6/1.7 β research blog https://research.nvidia.com/labs/gear/ + code at https://github.com/NVIDIA/Isaac-GR00T (tags
n1.5-release,n1.6-release,n1.7-release) βgr00t/model/modules/dit.py,eagle_backbone.py,qwen3_backbone.py. - StarVLA β https://arxiv.org/abs/2604.05014 Β· code https://github.com/starVLA/starVLA Β· StarVLA-Ξ±: https://arxiv.org/abs/2604.11757
- NORA β https://arxiv.org/abs/2504.19854 Β· NORA-1.5: https://arxiv.org/abs/2511.14659
- Qwen2.5-VL tech report β https://arxiv.org/abs/2502.13923 Β· Qwen3-VL: https://github.com/QwenLM/Qwen3-VL
- Cosmos-Reason2-2B β https://huggingface.co/nvidia/Cosmos-Reason2-2B
- RDT-1B (Alternating Condition Injection origin) β https://arxiv.org/abs/2410.07864
- ML-Attention β foundational primer on attention variants (causal / bidir / windowed / GQA / linear / gated)
- Review-VLM-Action-Connection β interface mechanism taxonomy (KV-share / cross-attn / FiLM / latent / etc.)
- Review-VLA-Architecture β action-head paradigm taxonomy (AR / flow / diffusion / VAM / hierarchical)
- Review-pi07 β Ο0.7 long-form review (the authoritative attention description per Β§4.1)
- Review-GR00T-Series β GR00T N1βN1.7 code-level evolution (the authoritative cross-attn description)
- pi-series-evolution β Ο series side-by-side architecture deltas (corrected attention row as of 2026-05-01)
Verdict: attention design now follows the runtime (RTC/streaming/prefix-KV), not the backbone β and the biggest new variable, linear attention, is entirely unablated for control.
Since late 2025 the mask and cache are chosen for RTC-compatibility, chunk streaming, and prefix-KV reuse rather than backbone convention. 2026 additions: hybrid linear-attention backbones entered VLAs via the Qwen suite (Qwen3.5's gated DeltaNet + interval GQA) with zero published ablation of the effect on control quality; and history/context conditioning was shown to interact with attention budgets (RobotManip's in-context variant needs 10 denoising steps where the base needs 4).
| Design | Pros | Cons |
|---|---|---|
| Prefix-KV + parallel expert branch (Ο) | Cache reuse across denoise steps | Backbone coupling |
| Full cross-attention re-read (GR00T / RobotManip) | Fresh conditioning per chunk | Re-encode cost each step |
| Streaming-token designs (Stream-to-Act-class) | Continuous execution | Young, little benchmark coverage |
| Hybrid linear attention (Qwen3.5-based VLAs) | Long-context cost | Control-quality impact unknown |
ICML 2026 β the deployment cluster arrives. The strongest "make VLAs runnable" wave of any 2026 venue: Reflex exploits flow-timestep invariance to partition attention into static/sliding/dynamic regions with O(1) cache updates (2.58Γ speedup, 50 Hz stable streaming); See What Matters's differentiable grid sampling cuts 76% of FLOPs with no success drop; SpecPrune-VLA (1.57β1.70Γ) and EcoVLA (1.60Γ, 2.18Γ combined) prune training-free; and the XPU characterization study establishes the field's first hardware profile β VLM phase is compute-bound, action-expert phase is memory-bound β enabling 2.9β3.3Γ via phase-aware scheduling.
- Latency (ms/chunk) is the least-reported number in the field β see Review-Realtime-Execution for what little exists; ICML's deployment cluster is the first systematic push to close this.
- Linear-attention-for-control is the largest open experimental gap on this axis.
- KV-cache behavior under multi-view, high-rate vision remains folklore; no benchmark isolates the attention axis.
β Back to Home