ML - Heungwoo/research GitHub Wiki
ML Foundations — Building-Block Reviews
Compiled April 2026 · Focus: the transformer building blocks that sit underneath every modern VLA/VLM. This section is written as a primer with concrete VLA/Qwen cross-references, not a general deep-learning textbook.
The other wiki sections (architectures, RL, memory, …) assume you already know what "GQA" or "RMSNorm" or "AdaLN" mean. This section is the glossary-with-diagrams that backs those assumptions. Each page follows the same shape: (1) a taxonomy of the family, (2) per-variant mermaid diagrams + formulas, (3) a concrete mapping to what Qwen3 / Qwen3.5 and the major VLAs (π, GR00T, DDVLA, RDT, Fast-in-Slow) actually use.
Pages
| Page | Scope |
|---|---|
| [Attention Variants](/Heungwoo/research/wiki/ML-Attention) | Full / causal / sliding-window / cross / bidirectional / GQA (& MQA / MHA) / linear (DeltaNet, Mamba-style) / gated (Gated Attention, Gated DeltaNet). Concrete: Qwen3 = GQA + QK-Norm, Qwen3-Next / Qwen3.5 = 3:1 Gated DeltaNet : Gated Attention hybrid, π-series = same-stack MoE + prefix-KV + block-causal, GR00T = cross-attention from DiT + AlternateVLDiT, DDVLA = bidirectional over action tokens. |
| [Normalization Variants](/Heungwoo/research/wiki/ML-Normalization) | BatchNorm / LayerNorm / RMSNorm / GroupNorm / InstanceNorm / QK-Norm / AdaLN (DiT-style, scale-shift-gate) / adaptive RMSNorm / pre-norm vs post-norm. Concrete: Qwen3 = pre-RMSNorm + per-head QK-Norm, GR00T N1.6/N1.7 = vlln LayerNorm before DiT cross-attention, π0.7 = adaptive RMSNorm for timestep injection, RDT-1B = rejects AdaLN (variable-length conditions). |
How to read this section
If you are already familiar with transformers, skim §1–3 of each page for the taxonomy table, then jump to the "what Qwen/VLA actually uses" section at the bottom — that is the part that is novel to this wiki rather than to a standard textbook.
If you are new, read top-to-bottom. Every mechanism comes with a mermaid diagram and the formula needed to implement it from scratch in a few lines of PyTorch.
Why put this in a VLA wiki?
Because the attention and norm choices in a VLA are not cosmetic: they change latency (FlashAttention vs dense, linear vs softmax), they change training stability (QK-Norm was added to Qwen3 explicitly to fix large-scale training blow-ups), and they change which VLM→action-head connection pattern is even possible (prefix-KV caching in the π-series requires matched head-dim, pre-norm, and causal masking on the prefix). Picking a VLM backbone commits you to its attention and norm family — so you should at least know what you are committing to.
Related reviews
- Review-VLA-Architecture — the architectural-family-level view (AR / flow / diffusion / dual-system / …).
- Review-VLM-Action-Connection — how the VLM wires into the action head (same-stack MoE, cross-attn, FiLM, shared blocks). This ML section explains what the wire is made of; the VLM↔Action page explains where the wire goes.
- Review-GR00T-Series — GR00T N1 → N1.7 code-level evolution. Uses
vllnLayerNorm andvl_self_attentionterminology that is explained in the Normalization page.
← Back to Home