Training Methodology2016intermediate9 min read
Layer Normalization
تسوية الطبقة
Ba, J. L. · Kiros, J. R. · Hinton, G. E. — arXiv
The problem
dramatically accelerated of deep feed-forward networks by normalizing neuron inputs across a mini-batch. But it has two critical limitations: it depends on mini-batch size — small batches yield noisy statistics — and it is unclear how to apply it to recurrent neural networks, where each time step may see different input distributions and the sequence length varies across examples.
The contribution
computes the mean and from all neurons within a single layer on a single training example — not across the batch. Each example normalizes itself independently. This makes it batch-size agnostic and naturally applicable to recurrent networks: the same formula works identically at every time step, during training and , with no need for running statistics.
The impact
Layer normalization became a core building block of the architecture. Every GPT, BERT, and Claude model uses layer normalization after (or before) each and feed-forward sub-layer. It enabled stable training of very deep sequence models where batch normalization could not operate, and inspired subsequent variants like and .
Imagine a classroom exam. Batch normalization is like the teacher grading each question by comparing everyone's answer — she needs the whole class present to compute the curve. If only three students show up, the curve is unreliable.
Layer normalization works differently: each student normalizes their own answers by comparing all their question scores against each other. If one student tends to score high across all questions, that personal baseline is subtracted out. No other students are needed — each paper is self-calibrating.
This is why layer normalization works for recurrent networks processing one sentence at a time: it never needs to peek at other examples in the batch.
The problem: batch normalization needs a batch
Batch normalization was a breakthrough: it normalized each neuron's input by subtracting the mean and dividing by the standard deviation computed across the mini-batch. This smoothed the landscape, allowed higher learning rates, and cut training time dramatically.
But batch normalization carries a structural dependency: it needs multiple examples to compute meaningful statistics. This creates three problems:
-
Small batches yield noisy estimates. When GPU memory is tight and the batch size drops to 2 or 4, the mean and variance become unreliable, and training destabilizes.
-
Recurrent networks are a mismatch. In an , each time step produces activations from different input distributions. Batch normalization would need to compute and store separate statistics for every time step, and sequences of different lengths make this awkward — what statistics do you use at time step 100 if most sequences ended at step 50?
-
Training ≠ inference. During training, batch normalization uses the current mini-batch statistics. During inference, it uses stored running averages. This train–test mismatch can degrade performance, especially for small datasets.
The idea: normalize across features, not examples
The insight is a simple axis swap. Instead of computing statistics across different training examples (the batch dimension), compute them across all neurons within the same layer on a single example (the feature dimension).
Picture a spreadsheet where each row is a training example and each column is a neuron. Batch normalization standardizes each column — it looks down. Layer normalization standardizes each row — it looks across. Because each row is one example, no other examples are needed.
After normalizing, a learned scale () and shift () are applied per neuron, so the network can still represent any distribution it finds useful. The normalization removes unwanted variation; the affine parameters re-introduce only the variation that helps.
The math: mean, variance, normalize, transform
Layer normalization follows four clean steps. Given a single training example whose activations in a layer form a vector with neurons, we first compute the mean and variance across the entire vector, then normalize, then apply a learned .
The purpose is clear before we see the symbols: we want every neuron in the layer to land on a common scale (zero mean, unit variance), then let the network choose the scale and shift that produce the best output. The four steps proceed as follows.
Why it works for recurrent networks
Recurrent neural networks process sequences by token. At each time step , the depends on the current input and the previous hidden state. The distribution of can drift wildly as the sequence progresses — a phenomenon directly related to the that motivated batch normalization in the first place.
Applying batch normalization to an RNN would require computing statistics across the batch at each time step, and storing separate running means for time steps 1, 2, 3, and so on. This is impractical: sequences vary in length, and time step 100 might see only a handful of examples, producing unreliable statistics.
Layer normalization sidesteps the issue entirely. At every time step, it normalizes across its dimensions using only that single vector. No batch statistics, no time-step-dependent storage. The formula is identical at step 1 and step 1000, during training and during inference. This makes it a natural fit for LSTMs, GRUs, and later for the Transformer.
Layer normalization in the Transformer
When the Transformer was introduced in 2017, it adopted layer normalization as its stabilization method — not batch normalization. In the original Transformer each and block follows the pattern: sub-layer → add residual → layer norm. This is called Post-LN (normalize after the residual).
Later work (GPT-2 and many successors) moved the layer normalization before the sub-layer: layer norm → sub-layer → add residual. This Pre-LN variant makes training more stable at very large depths because the residual path stays clean — gradients flow through the addition without first passing through normalization.
In both variants, layer normalization does the same job: it re-centers and re-scales each token's representation so that the next sub-layer receives well-conditioned inputs. Without it, residual connections accumulate magnitude with depth, and training becomes fragile.
The normalization family
Layer normalization is one member of a family that varies by which dimensions are used to compute the mean and variance. Given a 4D of activations (batch × channels × height × width), each method normalizes along a different slice:
- Batch Norm normalizes across the batch for each channel — same formula, different examples.
- Layer Norm normalizes across all channels (and spatial dimensions) for each example — same example, different features.
- Instance Norm normalizes across spatial dimensions per channel per example — used in style transfer.
- Group Norm splits channels into groups and normalizes each group independently — a middle ground between Layer and Instance Norm.
The choice depends on the task. Batch Norm excels in large-batch CNNs. Layer Norm dominates in Transformers and RNNs. Group Norm is preferred when batch sizes are small but spatial structure matters (e.g. object detection). Understanding what axis you normalize over is the key to choosing the right method.
The same idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
def layer_norm(x, gamma, beta, eps=1e-5):
"""
x: (batch, features) — activations for one layer
gamma: (features,) — learned scale, initialized to 1
beta: (features,) — learned shift, initialized to 0
"""
# Step 1: mean across features (each example independently)
mu = x.mean(axis=-1, keepdims=True) # (batch, 1)
# Step 2: variance across features
var = x.var(axis=-1, keepdims=True) # (batch, 1)
# Step 3: normalize to zero mean, unit variance
x_hat = (x - mu) / np.sqrt(var + eps) # (batch, features)
# Step 4: learned affine transform
return gamma * x_hat + beta # (batch, features)
# Compare with batch normalization:
# BN computes mu and var across axis=0 (the batch).
# LN computes mu and var across axis=-1 (the features).
# That single axis change is the entire difference.Why it mattered
2015
Batch Normalization
Ioffe & Szegedy introduce batch normalization, dramatically accelerating feed-forward network training by normalizing across mini-batches.
2016
Layer Normalization
Ba, Kiros & Hinton propose layer normalization — normalizing across features instead of the batch, enabling RNNs and batch-size-independent training.
2017
Transformer adopts Layer Norm
Vaswani et al.'s "Attention Is All You Need" uses layer normalization in every encoder and decoder block, establishing it as the Transformer's standard normalization.
2018
Group Normalization
Wu & He propose Group Norm — a middle ground that splits channels into groups, working well for CNNs with small batches where both Batch Norm and Layer Norm struggle.
2019
Pre-LN Transformer
GPT-2 and subsequent models move layer norm before the sub-layer, making training more stable at extreme depths.
2021
RMSNorm
Zhang & Sennrich simplify layer norm by removing the mean subtraction — only dividing by the root mean square. Used in LLaMA and many modern LLMs for faster computation.
Layer normalization is one of those ideas whose importance was revealed by what came after it. At publication, it was a practical fix for RNN training. A year later, the Transformer made it indispensable. Today every large language model — GPT, Claude, Gemini, LLaMA — runs layer normalization (or its descendant RMSNorm) at every single layer, billions of times per inference. A simple axis swap became infrastructure.
CitationBa, Kiros, Hinton. Layer Normalization. arXiv, 2016.
Terms in this paper
- Layer Normalizationالتسوية الطبقية
- Batch Normalizationتسوية الدفعات الحسابية
- Normalizationالمعايرة القياسية للبيانات
- Internal Covariate Shiftالانزياح الداخلي للتوزيعات
- Recurrent Neural Network (RNN)الشبكة العصبية التكرارية
- Affine Transformالتحويل التآلفي
- Varianceالتباين
- Activation Functionدالة التنشيط