ML Normalization - Heungwoo/research GitHub Wiki

ML Foundations — Normalization Variants

Part of the ML Foundations section. Compiled April 2026. Scope: every normalization layer that shows up inside a modern VLA or Qwen-class VLM, with diagrams, formulas, and concrete VLA/Qwen3/Qwen3.5 cross-references.


1. Why normalization has a zoo of variants

Normalization layers exist to fix three separate problems, and each "variant" trades one fix for another:

  1. Internal covariate shift — activations drift during training, destabilizing gradients. (BatchNorm / LayerNorm.)
  2. Scale mismatch across batch — small or variable batches break BatchNorm's population estimates. (LayerNorm / GroupNorm / InstanceNorm.)
  3. Conditioning injection — diffusion / generative heads need to modulate features by a continuous condition (timestep, class, caption embedding) without adding a cross-attention block everywhere. (AdaLN / adaptive RMSNorm / FiLM.)
flowchart TB
  ROOT["Normalization"]
  ROOT --> STAT["Statistical<br/>reshape distribution"]
  ROOT --> COND["Conditional<br/>inject a signal"]

  STAT --> BN["BatchNorm<br/>over batch dim"]
  STAT --> LN["LayerNorm<br/>over feature dim"]
  STAT --> RMSN["RMSNorm"]
  STAT --> GN["GroupNorm<br/>over feature groups"]
  STAT --> INORM["InstanceNorm<br/>over H and W"]

  COND --> ADALN["AdaLN scale-shift"]
  COND --> ADALNZ["AdaLN-Zero"]
  COND --> ADARMS["Adaptive RMSNorm"]
  COND --> FILM["FiLM<br/>affine per-channel"]
  COND --> QKN["QK-Norm"]

  RMSN:::shipped
  ADALN:::shipped
  ADALNZ:::shipped
  ADARMS:::shipped
  QKN:::shipped
  classDef shipped fill:#fff3b0,stroke:#b58900,color:#333
Loading

Legend. Yellow-filled boxes are actually shipped in a 2026 VLA or Qwen release (details in §8): RMSNorm is standard in Qwen3 and Llama; AdaLN and AdaLN-Zero are the DiT-family modulation used in GR00T; adaptive RMSNorm is π0.7's timestep injection; QK-Norm is the Qwen2.5 to Qwen3 stability upgrade.


2. The four core formulas

2.1 BatchNorm (Ioffe & Szegedy 2015)

Normalize over the batch and spatial dims, per feature channel. Requires large batches and is broken by distribution shift between train and eval, which is why learned running-averages exist.

μ_c = mean over (B, H, W) for channel c
σ_c = stddev over (B, H, W) for channel c
y_{b,c,h,w} = γ_c · (x_{b,c,h,w} − μ_c) / sqrt(σ_c² + ε) + β_c

2.2 LayerNorm (Ba et al. 2016)

Normalize over the feature dim, per token. Batch-size-independent.

μ_t = mean over d for token t
σ_t = stddev over d for token t
y_{t,d} = γ_d · (x_{t,d} − μ_t) / sqrt(σ_t² + ε) + β_d

2.3 RMSNorm (Zhang & Sennrich 2019)

LayerNorm with the mean-removal step dropped, and no bias. Cheaper, empirically matches or beats LN in transformers.

rms_t = sqrt(mean(x_{t,d}² over d) + ε)
y_{t,d} = γ_d · x_{t,d} / rms_t

Code:

def rms_norm(x, gamma, eps=1e-6):
    rms = x.pow(2).mean(-1, keepdim=True).add(eps).sqrt()
    return gamma * x / rms

2.4 GroupNorm / InstanceNorm

GroupNorm splits channels into G groups and LayerNorms inside each group. InstanceNorm = GroupNorm with G = C. Both are batch-size-independent and widely used in vision CNNs and diffusion U-Nets.


3. The normalization axis, visualized

flowchart LR
  subgraph Tensor["A 4-D activation tensor: [Batch, Channel, Height, Width]"]
    direction TB
    N["N = batch"]
    C["C = channel/feature"]
    H["H × W = spatial"]
  end
  BN2["BatchNorm → normalize over (N, H, W) per C"]
  LN2["LayerNorm → normalize over (C) per (N)<br/>(for a token sequence: over feature dim per token)"]
  GN2["GroupNorm → split C into G groups, LayerNorm inside each"]
  IN2["InstanceNorm → normalize over (H, W) per (N, C)"]
  RMSN2["RMSNorm → LayerNorm without mean removal or bias"]
Loading

For a transformer token sequence of shape [B, N, d], LayerNorm / RMSNorm both act over the last axis d, independent per (b, token). The only difference is whether μ is subtracted and β is added.


4. Pre-norm vs post-norm — where does LN go in the block?

Placement matters. It is the single most important stability choice in a transformer.

flowchart LR
  subgraph Post["Post-norm: original GPT-1 and BERT"]
    PIN["x"] --> PATT["attention"] --> PADD1["residual add"] --> PLN1["LayerNorm"] --> PFF["FFN"] --> PADD2["residual add"] --> PLN2["LayerNorm"] --> POUT["out"]
  end
  subgraph Pre["Pre-norm: GPT-2 onward, Qwen3, Llama, all 2024+ LLMs"]
    RIN["x"] --> RLN1["LayerNorm"] --> RATT["attention"] --> RADD1["residual add"]
    RADD1 --> RLN2["LayerNorm"] --> RFF["FFN"] --> RADD2["residual add"] --> ROUT["out"]
  end
Loading

Pre-norm wins at scale because the residual path is unnormalized all the way through, so gradient magnitudes do not explode/vanish with depth. Every 2024+ production LLM is pre-norm.


5. QK-Norm — why Qwen3 added it

Discussed also on the Attention page. The full norm story:

flowchart LR
  X["Input x"] --> WQ["W_Q times x"] --> Qraw["Q_raw"] --> QNORM["RMSNorm per-head<br/>over head_dim"] --> Q["Q"]
  X --> WK["W_K times x"] --> Kraw["K_raw"] --> KNORM["RMSNorm per-head<br/>over head_dim"] --> K["K"]
  X --> WV["W_V times x"] --> V["V"]
  Q --> DOT["Q · K_transposed<br/>scaled by sqrt d_head"]
  K --> DOT
  DOT --> SOFT["softmax"]
  SOFT --> MUL["times V"]
  V --> MUL
  MUL --> Y["Y"]
Loading

Qwen3 specifics (verbatim from transformers/models/qwen3/modeling_qwen3.py):

self.q_norm = Qwen3RMSNorm(self.head_dim, eps=config.rms_norm_eps)
self.k_norm = Qwen3RMSNorm(self.head_dim, eps=config.rms_norm_eps)
# V is NOT normalized

Why RMSNorm and not LayerNorm? Because Qwen3 uses RMSNorm everywhere else in the block; consistency means one norm primitive, one codepath, one kernel. Quality-wise, on QK they are indistinguishable.

Why per-head and not per-token? Because the pathology is per-head. One head's Q/K space can blow up while others are fine; a whole-token norm would wash that out.


6. AdaLN — the DiT modulation trick

In diffusion transformers, you need to condition every block on a timestep (and optionally class or text) without exploding compute. The trick from DiT (Peebles & Xie 2022) is to replace LayerNorm's static γ/β with timestep-conditioned γ, β, and gating α:

(γ, β, α) = MLP(timestep_embedding)       # 3·d numbers
h = α · [γ · LayerNorm(x) + β]            # scale, shift, gate

Equivalent mermaid:

flowchart LR
  T["timestep t"] --> EMB["sinusoidal embed<br/>then MLP"] --> THREE["produce gamma, beta, alpha<br/>each of dim d"]
  X["x"] --> LN["LayerNorm"] --> SCALE["scale by gamma"]
  THREE --> SCALE
  SCALE --> SHIFT["add beta"]
  THREE --> SHIFT
  SHIFT --> GATE["gate by alpha"]
  THREE --> GATE
  GATE --> OUT["out"]
Loading

AdaLN-Zero (the DiT block design used across all DiT sizes, including the flagship DiT-XL/2) initializes α to zero, so the block is the identity at init; this dramatically stabilizes training because the conditioning path has to earn its weight.

6.1 Adaptive RMSNorm — π0.7's variant

π0.7 keeps the AdaLN trick but swaps LayerNorm for RMSNorm — same reason Qwen3 did: consistency with the rest of the model, and RMSNorm is cheaper.

(γ, β, α) = MLP(timestep_embedding)
h = α · [γ · RMSNorm(x) + β]

The 860M flow-matching action expert in π0.7 uses this for timestep injection. No mean-subtraction, no bias on the static norm, just the adaptive γ/β/α.

6.2 Why RDT-1B rejects AdaLN

RDT-1B's authors argued AdaLN is wrong for VLA because the condition is not a fixed-dim vector but a variable-length token sequence (image + language). You cannot produce a single γ/β from a variable sequence without collapsing it, and collapsing loses spatial/temporal structure. RDT-1B therefore uses cross-attention throughout — see Review-VLM-Action-Connection.

This is a pattern: AdaLN is ideal when the condition is a short vector (timestep, class), poor when the condition is a long sequence (image patches, text).


7. FiLM — the underrated cousin

FiLM (Perez et al. 2018) is AdaLN without the LayerNorm step:

(γ, β) = MLP(condition)
h = γ · x + β          # affine modulation only

Less powerful than AdaLN on a per-block basis, but cheap and famously effective when the condition is low-dim. CogVLA uses a double-FiLM design for VLM→action conditioning (see Review-VLM-Action-Connection).


8. What Qwen3 / Qwen3.5 / the VLAs actually use

8.1 Qwen3 (dense and MoE)

Layer Type Scope
Block-level pre-norm on attention input RMSNorm, ε=1e-6 over feature dim
Block-level pre-norm on MLP input RMSNorm, ε=1e-6 over feature dim
Q projection output RMSNorm (QK-Norm) per-head, over head_dim
K projection output RMSNorm (QK-Norm) per-head, over head_dim
V projection output no norm —
Final pre-lm_head norm RMSNorm over feature dim

Qwen3's two Qwen2.5 → Qwen3 deltas are both in this table: QK-Norm added, and QKV bias removed.

8.2 Qwen3-VL

Same RMSNorm recipe as Qwen3 across ViT and LLM. No AdaLN / scale-shift / GroupNorm anywhere — confirmed from HuggingFace transformers/models/qwen3_vl/modeling_qwen3_vl.py.

8.3 Qwen3-Next / Qwen3.5

Pre-RMSNorm on each block, same ε=1e-6. The Gated DeltaNet blocks internally have their own per-state normalization required by the delta-rule recurrence (state normalization inside the recurrence), plus output gating — but at the block boundary it is still RMSNorm pre-norm. This means a Qwen3.5 block and a Qwen3 block look identical from the outside (same tensor shapes and norm type); only the core attention operator differs.

8.4 The 2026 VLAs, side-by-side

VLA VLM-trunk norm Action-head / DiT norm Special
π0.6 / π0.7 Gemma3-4B: RMSNorm pre-norm Adaptive RMSNorm for timestep injection in the 860M flow-matching action expert π0.7 is the most explicit adoption of adaptive-RMSNorm in a production VLA
GR00T N1 Eagle-2 LLM: LayerNorm pre-norm DiT with AdaLN (standard DiT recipe) Raw backbone_features → DiT cross-attention, no interface norm
GR00T N1.5 Eagle-2.5 frozen: LayerNorm DiT with AdaLN Post-VLM adapter MLP + LayerNorm before DiT
GR00T N1.6 Cosmos-Reason: RMSNorm (inherits Qwen-family base) DiT with AdaLN, 32 layers vlln LayerNorm inserted on backbone_features before DiT cross-attention (replaces N1.5's adapter)
GR00T N1.7 Cosmos-Reason2-2B = Qwen3-VL-2B → RMSNorm + QK-Norm DiT with AdaLN, 32 layers vlln LayerNorm + vl_self_attention SelfAttentionTransformer in sequence; N1.7 is the first GR00T that inherits Qwen3's QK-Norm through its backbone
RDT-1B VLM-native Cross-attention only; AdaLN explicitly rejected for variable-length conditions —
DDVLA Prismatic-7B: Llama 2 RMSNorm + SigLIP/DINOv2 LayerNorm Single unified transformer, no separate action-head norm Bidirectional action span shares the backbone's RMSNorm
Fast-in-Slow LLaVA-class LayerNorm S1 shares the last 2 blocks with S2 → shares their LayerNorms No additional norms for dual-frequency paths

8.5 GR00T N1.6/N1.7's vlln + vl_self_attention, in detail

This is the normalization choice that is most instructive for a working VLA engineer, because it shows up differently in each release.

flowchart LR
  subgraph N1["N1 — no interface norm"]
    V1["VLM hidden states"] --> DITCA1["DiT cross-attention"]
  end
  subgraph N15["N1.5 — adapter MLP and LayerNorm"]
    V2["VLM hidden states<br/>frozen"] --> ADAPTER["adapter MLP"] --> LN15["LayerNorm"] --> DITCA2["DiT cross-attention"]
  end
  subgraph N16["N1.6 — vlln replaces the adapter"]
    V3["VLM hidden states<br/>top-4 layers tuned"] --> VLLN["vlln LayerNorm"] --> DITCA3["DiT cross-attention"]
  end
  subgraph N17["N1.7 — vlln and vl_self_attention"]
    V4["VLM hidden states<br/>Qwen3-VL-2B"] --> VLLN2["vlln LayerNorm"] --> VSA["vl_self_attention<br/>SelfAttentionTransformer"] --> DITCA4["DiT cross-attention"]
  end
Loading

Why the progression:

  • N1 assumed raw VLM features were fine.
  • N1.5 wanted to freeze the VLM and still let gradients reshape the features → adapter MLP + LN.
  • N1.6 found the adapter MLP was redundant once top-4 layers were tunable → kept only vlln LayerNorm as a final reshape.
  • N1.7 added a small self-attention transformer (vl_self_attention, configurable depth) between vlln and DiT, letting VL features self-mix (image+text within itself) before being handed to the action path. This is effectively a mini Q-Former.

N1.7's vlln is still LayerNorm, not RMSNorm, despite the backbone (Qwen3-VL) being RMSNorm-native. The reason: vlln is an interface layer bolted onto the backbone by NVIDIA's action-head team, and DiT-family code tends to default to LayerNorm.

8.6 Why π0.7 uses adaptive RMSNorm but GR00T uses AdaLN

Two design forks from the same DiT ancestor:

  • GR00T kept DiT's LayerNorm-based AdaLN because the DiT code they inherited uses it, and action-head training is stable enough.
  • π0.7 swapped LayerNorm for RMSNorm throughout the action expert, because the rest of the π stack (Gemma3-4B backbone) is RMSNorm-native and they wanted one primitive. The math is identical minus mean-subtraction, which is a no-op on centered activations anyway.

9. Decision guide

If you are building… Pick
A 2024+ open-source LLM (Qwen / Llama style) Pre-RMSNorm + per-head QK-Norm (RMS)
A ViT or an encoder with short fixed-length inputs Pre-LayerNorm (still the default)
A diffusion or flow-matching action head conditioned on a scalar timestep AdaLN or AdaLN-Zero (if LN backbone) / adaptive RMSNorm (if RMSNorm backbone)
An action head conditioned on a variable-length token sequence (image + text) Cross-attention + pre-norm, not AdaLN (RDT-1B's argument)
A two-transformer system where an action DiT reads a VLM's hidden states vlln LayerNorm interface layer (GR00T N1.6/N1.7), optionally with a small vl_self_attention mixer
A new CNN-ish vision tower with small batches GroupNorm (not BatchNorm)

Links

Related pages

← Back to ML Foundations · Home

⚠️ **GitHub.com Fallback** ⚠️