Computer Vision2018intermediate9 min read

Group Normalization

التسوية بالمجموعات

Wu, Y. · He, K. — ECCV

The problem

(BN) revolutionized by stabilizing through normalizing features using batch statistics. But BN has an Achilles heel: it requires large batch sizes (e.g., 32 per ) for accurate mean and estimates. When memory constraints force small batches — as in (1–2 images), video classification, and — BN's error increases dramatically. A batch of 2 images gives a hopelessly noisy estimate of the population statistics, and the network's accuracy collapses.

The contribution

Group (GN): divide channels into groups (32 by default) and compute mean and variance within each group per sample. GN is entirely independent of — its statistics come from a single image's channels, not from the batch. With a batch size of 2, GN achieves 10.6% lower error than BN on ResNet-50 / ImageNet. At typical batch sizes, GN matches BN within 0.5%. GN transfers naturally from pre-training to and outperforms BN-based counterparts on COCO detection/ and Kinetics video classification.

The impact

GN unlocked high-capacity vision models that were previously impossible to train with small batches. It became the default normalization in Mask R- and Detectron2 for detection and segmentation. ConvNeXt (2022) adopted GN as part of its modernized ConvNet design that matches performance. GN demonstrated that the choice of which dimensions to normalize over is a fundamental design axis, opening the door to task-specific normalization strategies.

Imagine a school where the principal evaluates teachers by averaging their students' test scores. With 30 students per class, the average is meaningful. But budget cuts shrink some classes to just 2 students — now one bad test day makes the teacher look terrible.

Batch Normalization is the principal averaging across the batch. With large batches it works beautifully, but with tiny batches the average is noise.

Group Normalization changes the strategy entirely: instead of averaging across students (the batch), it groups subjects (channels) — math with physics, history with geography — and normalizes within each subject group for every individual student. The class size no longer matters.

The problem: Batch Normalization breaks with small batches

Batch Normalization computes the mean and variance of features across the batch dimension. With 32 images per GPU, these statistics are stable estimates of the true population statistics. But many critical vision tasks cannot afford large batches:

  • Object detection (Faster R-CNN, Mask R-CNN): high-resolution images limit batches to 1–2 images per GPU.
  • Video classification (3D convolutions): spatiotemporal features eat memory, forcing small batches.
  • Semantic segmentation: dense predictions on full-resolution images leave little room for large batches.

When the batch size drops to 2, BN's error on ResNet-50 jumps by 10.6% compared to its performance at batch size 32. The statistics become so noisy that the network essentially trains on random normalization — the very problem BN was designed to solve.

Open in Lab
Drag the batch size slider and watch BN's error explode at small batches while GN stays rock-steady.
The demo wakes as you arrive…

The normalization landscape: what changes is *which pixels* you average

Every normalization method applies the same formula — subtract the mean, divide by the standard deviation, then apply a learnable scale and shift. The only difference is the set of pixels Si\mathcal{S}_i over which the mean and variance are computed. Think of the feature as a cube with axes (N,C,H,W)(N, C, H, W) — batch, channels, height, width — and each method paints a different slice of that cube blue:

  • Batch Norm: one slice per , across all samples in the batch — the (N, H, W) axes. Depends on batch size.
  • Layer Norm: one slice per sample, across all channels — the (C, H, W) axes. Independent of batch size, but assumes all channels contribute equally.
  • Instance Norm: one slice per sample per channel — only the (H, W) axes. Ignores relationships between channels.
  • Group Norm: one slice per sample per group of channels — the (H, W) axes plus C/G channels. The sweet spot: batch-independent, but channels within each group are normalized together, preserving inter-channel structure.
Open in Lab
Toggle between BN, LN, IN, and GN to see which pixels (blue) share statistics. Notice how GN operates per-sample but groups channels together.
The demo wakes as you arrive…

The formula: same engine, different fuel

All normalization methods share the same two-step engine. First, standardize the features to zero mean and unit variance. Second, apply a learnable so the network can undo the normalization if it needs to. What makes GN special is how it defines the neighborhood of pixels that share statistics.

x^i=xi−μiσi,yi=γx^i+β\hat{x}_i = \frac{x_i - \mu_i}{\sigma_i}, \quad y_i = \gamma \hat{x}_i + \beta
Normalize then rescale — shared by BN, LN, IN, and GN — Subtract the mean μ, divide by the standard deviation σ, then apply learnable scale γ and shift β (one per channel). The only difference between methods is which set of pixels contributes to μ and σ.
Si={k  |  kN=iN,  ⌊kCC/G⌋=⌊iCC/G⌋}\mathcal{S}_i = \left\{ k \;\middle|\; k_N = i_N, \; \left\lfloor \frac{k_C}{C/G} \right\rfloor = \left\lfloor \frac{i_C}{C/G} \right\rfloor \right\}
Group Norm's pixel set — the key innovation — Two pixels i and k share statistics if they belong to the same sample (k_N = i_N) and the same group of channels (floor division assigns channels to groups). G groups, each with C/G channels.

Why grouping channels makes biological and computational sense

Channels in a convolutional network are not independent — filters that detect similar patterns (e.g., horizontal edges at different positions) naturally have correlated responses. Classical features like and HOG are explicitly designed as group-wise representations: each group is a histogram of oriented gradients within a spatial cell, and the histogram is normalized as a unit.

In neuroscience, a well-established computational model is divisive normalization — neurons normalize their responses across groups of cells with similar properties. This happens not just in the primary visual cortex but throughout the visual system. GN brings this principle to deep networks: channels that likely encode related features (edges at different orientations, textures at different frequencies) are normalized together.

Open in Lab
Explore how 64 channels are divided into groups. Click a group to see which feature maps share statistics.
The demo wakes as you arrive…

Implementation: reshape, normalize, reshape back

The beauty of GN is its simplicity. The entire implementation is a reshape trick: take the (N, C, H, W) tensor, reshape it to (N, G, C//G, H, W), compute mean and variance along axes (2, 3, 4) — which are the channels-within-group and spatial dimensions — normalize, then reshape back. Three lines of code. No running averages, no train-vs-test discrepancy, no synchronization across GPUs.

Group Normalization in 7 lines — the complete algorithmpython

Simplified to show the idea — not the real implementation.

def group_norm(x, gamma, beta, G=32, eps=1e-5):
    """x: input features with shape [N, C, H, W]
       gamma, beta: learnable scale and shift, shape [1, C, 1, 1]
       G: number of groups (default 32)"""
    N, C, H, W = x.shape
    x = x.reshape(N, G, C // G, H, W)        # split channels into G groups
    mean = x.mean(axis=(2, 3, 4), keepdims=True)  # mean within each group
    var  = x.var(axis=(2, 3, 4), keepdims=True)   # variance within each group
    x = (x - mean) / np.sqrt(var + eps)       # standardize
    x = x.reshape(N, C, H, W)                 # restore original shape
    return x * gamma + beta                    # learnable affine transform

Results: stable where BN crumbles

The paper evaluates GN on three fronts — ImageNet classification, COCO object detection/segmentation, and Kinetics video classification — systematically comparing against BN, LN, and IN.

On ImageNet with ResNet-50 at batch size 32, GN trails BN by only 0.5% (24.1% vs. 23.6%). LN trails by 1.7% and IN by 4.8%. GN strikes the best balance among batch-independent methods.

The dramatic advantage emerges at small batch sizes: at batch size 2, BN's validation error jumps from 23.6% to 34.7% — a catastrophic 11.1% increase. GN at batch size 2 achieves 24.1%, virtually identical to its batch-size-32 performance. The error curve is essentially flat across all batch sizes from 2 to 32.

On COCO with Mask R-CNN (ResNet-50 backbone, batch size 1 per GPU), GN outperforms a frozen-BN baseline in both box AP and mask AP. On Kinetics video classification with 3D ResNet-50, GN again outperforms BN when batch sizes are constrained.

Open in Lab
Compare validation error rates across normalization methods and batch sizes. Toggle between ImageNet classification, COCO detection, and Kinetics video.
The demo wakes as you arrive…

Sensitivity to the number of groups G

How many groups should you use? The paper tests G = 1, 2, 4, 8, 16, 32, and 64 on ResNet-50 / ImageNet. The results show remarkable insensitivity: all values from G=8 to G=64 perform within 0.5% of each other, with G=32 being slightly optimal. The two extremes confirm expectations: G=1 (Layer Norm) is 1.2% worse, and G=C (Instance Norm) is 4.7% worse.

This insensitivity is practical gold — you can pick G=32 and not worry about tuning it. The default works across ResNet-50, ResNet-101, and different tasks (classification, detection, segmentation, video).

Open in Lab
Drag G from 1 to 64 and see how validation error barely changes in the sweet spot. Extremes (G=1 and G=C) are labeled.
The demo wakes as you arrive…

Transfer learning: GN's hidden advantage

BN has a subtle problem. During pre-training on ImageNet, BN accumulates running statistics (population mean and variance) from the training set. When you fine-tune on a different dataset or task, those frozen statistics may not match the new data . Practitioners either freeze BN layers (losing adaptability) or re-estimate statistics (adding complexity).

GN has no running statistics — it computes everything from the current input. This means GN transfers seamlessly: the same computation happens during pre-training and fine-tuning, with no distribution mismatch. The paper shows this matters in practice: GN achieves higher AP than frozen-BN baselines on COCO, partly because it can adapt its normalization to the new task's feature distributions.

Legacy: from ECCV 2018 to ConvNeXt and beyond

GN's immediate impact was in object detection and segmentation — Mask R-CNN with GN became the standard baseline in Detectron2. But GN's deeper lesson is architectural: the dimensions you normalize over are a design choice, not a fixed rule.

When Liu et al. designed ConvNeXt (2022) — a pure ConvNet that matches Vision Transformer accuracy — they deliberately chose GN (replacing BN) as one of the modernization steps. The rationale was clean: GN aligns with the per-sample computation philosophy of Transformers (which use ) while preserving the channel-grouping structure that benefits convolutions.

  1. 2015

    Batch Normalization

    Ioffe & Szegedy introduce BN, enabling training of very deep networks by normalizing across the batch. Becomes ubiquitous but requires large batch sizes.

  2. 2016

    Layer Normalization

    Ba et al. normalize across all channels per sample. Batch- independent, works well for RNNs and Transformers but underperforms on visual tasks.

  3. 2018

    Group Normalization

    Wu & He find the sweet spot: group channels together, normalize per sample per group. Batch-independent, stable from batch 2 to 32, natural transfer learning.

  4. 2022

    ConvNeXt adopts GN

    A pure ConvNet redesigned with modern principles. Replacing BN with GN is one of the key changes that lets ConvNeXt match Vision Transformers.

CitationWu, He. Group Normalization. ECCV, 2018.

Terms in this paper