ML Normalization - Heungwoo/research GitHub Wiki
Part of the ML Foundations section. Compiled April 2026. Scope: every normalization layer that shows up inside a modern VLA or Qwen-class VLM, with diagrams, formulas, and concrete VLA/Qwen3/Qwen3.5 cross-references.
Normalization layers exist to fix three separate problems, and each "variant" trades one fix for another:
- Internal covariate shift — activations drift during training, destabilizing gradients. (BatchNorm / LayerNorm.)
- Scale mismatch across batch — small or variable batches break BatchNorm's population estimates. (LayerNorm / GroupNorm / InstanceNorm.)
- Conditioning injection — diffusion / generative heads need to modulate features by a continuous condition (timestep, class, caption embedding) without adding a cross-attention block everywhere. (AdaLN / adaptive RMSNorm / FiLM.)
flowchart TB
ROOT["Normalization"]
ROOT --> STAT["Statistical<br/>reshape distribution"]
ROOT --> COND["Conditional<br/>inject a signal"]
STAT --> BN["BatchNorm<br/>over batch dim"]
STAT --> LN["LayerNorm<br/>over feature dim"]
STAT --> RMSN["RMSNorm"]
STAT --> GN["GroupNorm<br/>over feature groups"]
STAT --> INORM["InstanceNorm<br/>over H and W"]
COND --> ADALN["AdaLN scale-shift"]
COND --> ADALNZ["AdaLN-Zero"]
COND --> ADARMS["Adaptive RMSNorm"]
COND --> FILM["FiLM<br/>affine per-channel"]
COND --> QKN["QK-Norm"]
RMSN:::shipped
ADALN:::shipped
ADALNZ:::shipped
ADARMS:::shipped
QKN:::shipped
classDef shipped fill:#fff3b0,stroke:#b58900,color:#333
Legend. Yellow-filled boxes are actually shipped in a 2026 VLA or Qwen release (details in §8): RMSNorm is standard in Qwen3 and Llama; AdaLN and AdaLN-Zero are the DiT-family modulation used in GR00T; adaptive RMSNorm is π0.7's timestep injection; QK-Norm is the Qwen2.5 to Qwen3 stability upgrade.
Normalize over the batch and spatial dims, per feature channel. Requires large batches and is broken by distribution shift between train and eval, which is why learned running-averages exist.
μ_c = mean over (B, H, W) for channel c
σ_c = stddev over (B, H, W) for channel c
y_{b,c,h,w} = γ_c · (x_{b,c,h,w} − μ_c) / sqrt(σ_c² + ε) + β_c
Normalize over the feature dim, per token. Batch-size-independent.
μ_t = mean over d for token t
σ_t = stddev over d for token t
y_{t,d} = γ_d · (x_{t,d} − μ_t) / sqrt(σ_t² + ε) + β_d
LayerNorm with the mean-removal step dropped, and no bias. Cheaper, empirically matches or beats LN in transformers.
rms_t = sqrt(mean(x_{t,d}² over d) + ε)
y_{t,d} = γ_d · x_{t,d} / rms_t
Code:
def rms_norm(x, gamma, eps=1e-6):
rms = x.pow(2).mean(-1, keepdim=True).add(eps).sqrt()
return gamma * x / rmsGroupNorm splits channels into G groups and LayerNorms inside each group. InstanceNorm = GroupNorm with G = C. Both are batch-size-independent and widely used in vision CNNs and diffusion U-Nets.
flowchart LR
subgraph Tensor["A 4-D activation tensor: [Batch, Channel, Height, Width]"]
direction TB
N["N = batch"]
C["C = channel/feature"]
H["H × W = spatial"]
end
BN2["BatchNorm → normalize over (N, H, W) per C"]
LN2["LayerNorm → normalize over (C) per (N)<br/>(for a token sequence: over feature dim per token)"]
GN2["GroupNorm → split C into G groups, LayerNorm inside each"]
IN2["InstanceNorm → normalize over (H, W) per (N, C)"]
RMSN2["RMSNorm → LayerNorm without mean removal or bias"]
For a transformer token sequence of shape [B, N, d], LayerNorm / RMSNorm both act over the last axis d, independent per (b, token). The only difference is whether μ is subtracted and β is added.
Placement matters. It is the single most important stability choice in a transformer.
flowchart LR
subgraph Post["Post-norm: original GPT-1 and BERT"]
PIN["x"] --> PATT["attention"] --> PADD1["residual add"] --> PLN1["LayerNorm"] --> PFF["FFN"] --> PADD2["residual add"] --> PLN2["LayerNorm"] --> POUT["out"]
end
subgraph Pre["Pre-norm: GPT-2 onward, Qwen3, Llama, all 2024+ LLMs"]
RIN["x"] --> RLN1["LayerNorm"] --> RATT["attention"] --> RADD1["residual add"]
RADD1 --> RLN2["LayerNorm"] --> RFF["FFN"] --> RADD2["residual add"] --> ROUT["out"]
end
Pre-norm wins at scale because the residual path is unnormalized all the way through, so gradient magnitudes do not explode/vanish with depth. Every 2024+ production LLM is pre-norm.
Discussed also on the Attention page. The full norm story:
flowchart LR
X["Input x"] --> WQ["W_Q times x"] --> Qraw["Q_raw"] --> QNORM["RMSNorm per-head<br/>over head_dim"] --> Q["Q"]
X --> WK["W_K times x"] --> Kraw["K_raw"] --> KNORM["RMSNorm per-head<br/>over head_dim"] --> K["K"]
X --> WV["W_V times x"] --> V["V"]
Q --> DOT["Q · K_transposed<br/>scaled by sqrt d_head"]
K --> DOT
DOT --> SOFT["softmax"]
SOFT --> MUL["times V"]
V --> MUL
MUL --> Y["Y"]
Qwen3 specifics (verbatim from transformers/models/qwen3/modeling_qwen3.py):
self.q_norm = Qwen3RMSNorm(self.head_dim, eps=config.rms_norm_eps)
self.k_norm = Qwen3RMSNorm(self.head_dim, eps=config.rms_norm_eps)
# V is NOT normalizedWhy RMSNorm and not LayerNorm? Because Qwen3 uses RMSNorm everywhere else in the block; consistency means one norm primitive, one codepath, one kernel. Quality-wise, on QK they are indistinguishable.
Why per-head and not per-token? Because the pathology is per-head. One head's Q/K space can blow up while others are fine; a whole-token norm would wash that out.
In diffusion transformers, you need to condition every block on a timestep (and optionally class or text) without exploding compute. The trick from DiT (Peebles & Xie 2022) is to replace LayerNorm's static γ/β with timestep-conditioned γ, β, and gating α:
(γ, β, α) = MLP(timestep_embedding) # 3·d numbers
h = α · [γ · LayerNorm(x) + β] # scale, shift, gate
Equivalent mermaid:
flowchart LR
T["timestep t"] --> EMB["sinusoidal embed<br/>then MLP"] --> THREE["produce gamma, beta, alpha<br/>each of dim d"]
X["x"] --> LN["LayerNorm"] --> SCALE["scale by gamma"]
THREE --> SCALE
SCALE --> SHIFT["add beta"]
THREE --> SHIFT
SHIFT --> GATE["gate by alpha"]
THREE --> GATE
GATE --> OUT["out"]
AdaLN-Zero (the DiT block design used across all DiT sizes, including the flagship DiT-XL/2) initializes α to zero, so the block is the identity at init; this dramatically stabilizes training because the conditioning path has to earn its weight.
π0.7 keeps the AdaLN trick but swaps LayerNorm for RMSNorm — same reason Qwen3 did: consistency with the rest of the model, and RMSNorm is cheaper.
(γ, β, α) = MLP(timestep_embedding)
h = α · [γ · RMSNorm(x) + β]
The 860M flow-matching action expert in π0.7 uses this for timestep injection. No mean-subtraction, no bias on the static norm, just the adaptive γ/β/α.
RDT-1B's authors argued AdaLN is wrong for VLA because the condition is not a fixed-dim vector but a variable-length token sequence (image + language). You cannot produce a single γ/β from a variable sequence without collapsing it, and collapsing loses spatial/temporal structure. RDT-1B therefore uses cross-attention throughout — see Review-VLM-Action-Connection.
This is a pattern: AdaLN is ideal when the condition is a short vector (timestep, class), poor when the condition is a long sequence (image patches, text).
FiLM (Perez et al. 2018) is AdaLN without the LayerNorm step:
(γ, β) = MLP(condition)
h = γ · x + β # affine modulation only
Less powerful than AdaLN on a per-block basis, but cheap and famously effective when the condition is low-dim. CogVLA uses a double-FiLM design for VLM→action conditioning (see Review-VLM-Action-Connection).
| Layer | Type | Scope |
|---|---|---|
| Block-level pre-norm on attention input | RMSNorm, ε=1e-6 | over feature dim |
| Block-level pre-norm on MLP input | RMSNorm, ε=1e-6 | over feature dim |
| Q projection output | RMSNorm (QK-Norm) | per-head, over head_dim |
| K projection output | RMSNorm (QK-Norm) | per-head, over head_dim |
| V projection output | no norm | — |
| Final pre-lm_head norm | RMSNorm | over feature dim |
Qwen3's two Qwen2.5 → Qwen3 deltas are both in this table: QK-Norm added, and QKV bias removed.
Same RMSNorm recipe as Qwen3 across ViT and LLM. No AdaLN / scale-shift / GroupNorm anywhere — confirmed from HuggingFace transformers/models/qwen3_vl/modeling_qwen3_vl.py.
Pre-RMSNorm on each block, same ε=1e-6. The Gated DeltaNet blocks internally have their own per-state normalization required by the delta-rule recurrence (state normalization inside the recurrence), plus output gating — but at the block boundary it is still RMSNorm pre-norm. This means a Qwen3.5 block and a Qwen3 block look identical from the outside (same tensor shapes and norm type); only the core attention operator differs.
| VLA | VLM-trunk norm | Action-head / DiT norm | Special |
|---|---|---|---|
| π0.6 / π0.7 | Gemma3-4B: RMSNorm pre-norm | Adaptive RMSNorm for timestep injection in the 860M flow-matching action expert | π0.7 is the most explicit adoption of adaptive-RMSNorm in a production VLA |
| GR00T N1 | Eagle-2 LLM: LayerNorm pre-norm | DiT with AdaLN (standard DiT recipe) | Raw backbone_features → DiT cross-attention, no interface norm |
| GR00T N1.5 | Eagle-2.5 frozen: LayerNorm | DiT with AdaLN | Post-VLM adapter MLP + LayerNorm before DiT |
| GR00T N1.6 | Cosmos-Reason: RMSNorm (inherits Qwen-family base) | DiT with AdaLN, 32 layers |
vlln LayerNorm inserted on backbone_features before DiT cross-attention (replaces N1.5's adapter) |
| GR00T N1.7 | Cosmos-Reason2-2B = Qwen3-VL-2B → RMSNorm + QK-Norm | DiT with AdaLN, 32 layers |
vlln LayerNorm + vl_self_attention SelfAttentionTransformer in sequence; N1.7 is the first GR00T that inherits Qwen3's QK-Norm through its backbone |
| RDT-1B | VLM-native | Cross-attention only; AdaLN explicitly rejected for variable-length conditions | — |
| DDVLA | Prismatic-7B: Llama 2 RMSNorm + SigLIP/DINOv2 LayerNorm | Single unified transformer, no separate action-head norm | Bidirectional action span shares the backbone's RMSNorm |
| Fast-in-Slow | LLaVA-class LayerNorm | S1 shares the last 2 blocks with S2 → shares their LayerNorms | No additional norms for dual-frequency paths |
This is the normalization choice that is most instructive for a working VLA engineer, because it shows up differently in each release.
flowchart LR
subgraph N1["N1 — no interface norm"]
V1["VLM hidden states"] --> DITCA1["DiT cross-attention"]
end
subgraph N15["N1.5 — adapter MLP and LayerNorm"]
V2["VLM hidden states<br/>frozen"] --> ADAPTER["adapter MLP"] --> LN15["LayerNorm"] --> DITCA2["DiT cross-attention"]
end
subgraph N16["N1.6 — vlln replaces the adapter"]
V3["VLM hidden states<br/>top-4 layers tuned"] --> VLLN["vlln LayerNorm"] --> DITCA3["DiT cross-attention"]
end
subgraph N17["N1.7 — vlln and vl_self_attention"]
V4["VLM hidden states<br/>Qwen3-VL-2B"] --> VLLN2["vlln LayerNorm"] --> VSA["vl_self_attention<br/>SelfAttentionTransformer"] --> DITCA4["DiT cross-attention"]
end
Why the progression:
- N1 assumed raw VLM features were fine.
- N1.5 wanted to freeze the VLM and still let gradients reshape the features → adapter MLP + LN.
- N1.6 found the adapter MLP was redundant once top-4 layers were tunable → kept only
vllnLayerNorm as a final reshape. - N1.7 added a small self-attention transformer (
vl_self_attention, configurable depth) between vlln and DiT, letting VL features self-mix (image+text within itself) before being handed to the action path. This is effectively a mini Q-Former.
N1.7's vlln is still LayerNorm, not RMSNorm, despite the backbone (Qwen3-VL) being RMSNorm-native. The reason: vlln is an interface layer bolted onto the backbone by NVIDIA's action-head team, and DiT-family code tends to default to LayerNorm.
Two design forks from the same DiT ancestor:
- GR00T kept DiT's LayerNorm-based AdaLN because the DiT code they inherited uses it, and action-head training is stable enough.
- π0.7 swapped LayerNorm for RMSNorm throughout the action expert, because the rest of the π stack (Gemma3-4B backbone) is RMSNorm-native and they wanted one primitive. The math is identical minus mean-subtraction, which is a no-op on centered activations anyway.
| If you are building… | Pick |
|---|---|
| A 2024+ open-source LLM (Qwen / Llama style) | Pre-RMSNorm + per-head QK-Norm (RMS) |
| A ViT or an encoder with short fixed-length inputs | Pre-LayerNorm (still the default) |
| A diffusion or flow-matching action head conditioned on a scalar timestep | AdaLN or AdaLN-Zero (if LN backbone) / adaptive RMSNorm (if RMSNorm backbone) |
| An action head conditioned on a variable-length token sequence (image + text) | Cross-attention + pre-norm, not AdaLN (RDT-1B's argument) |
| A two-transformer system where an action DiT reads a VLM's hidden states |
vlln LayerNorm interface layer (GR00T N1.6/N1.7), optionally with a small vl_self_attention mixer |
| A new CNN-ish vision tower with small batches | GroupNorm (not BatchNorm) |
- LayerNorm — arXiv:1607.06450
- RMSNorm — arXiv:1910.07467
- GroupNorm — arXiv:1803.08494
- DiT / AdaLN-Zero — arXiv:2212.09748
- FiLM — arXiv:1709.07871
- QK-Norm — arXiv:2010.04245
- Qwen3 tech report — arXiv:2505.09388
- Qwen3-VL tech report — arXiv:2511.21631
-
GR00T code —
nvidia/Isaac-GR00Ton GitHub (gr00t/model/action_head/*.py, especiallyGr00tN1d6ActionHead.process_backbone_outputandGr00tN1d7ActionHead.process_backbone_output)
- Attention Variants — sister page. QK-Norm lives on both.
-
Review-GR00T-Series — full N1 → N1.7 evolution including
vllnandvl_self_attention. - Review-VLM-Action-Connection — where AdaLN, FiLM, and cross-attention are compared at the VLA-architecture level.
- Review-pi07 — adaptive RMSNorm timestep injection in the π0.7 action expert.
← Back to ML Foundations · Home