Training Methodology2018intermediate10 min read

Mixed Precision Training

التدريب بالدقة المختلطة

Micikevicius, P. · Narang, S. · Alben, J. · Diamos, G. · Elsen, E. · Garcia, D. · Ginsburg, B. · Houston, M. · Kuchaiev, O. · Venkatesh, G. · Wu, H. — ICLR

The problem

Modern deep neural networks keep growing — more layers, more parameters, more data. them in standard 32-bit floating point (FP32) is expensive: every , activation, and occupies 4 bytes of GPU memory, and arithmetic on 32-bit numbers is slower than it needs to be. GPUs can do () math 2–8× faster, but FP16 has a much narrower : values below 2⁻²⁴ become zero, and small gradient updates vanish. Simply switching everything to FP16 causes models to diverge.

The contribution

Three techniques that let you train in FP16 without losing . First, keep an FP32 "master copy" of the weights and round it to FP16 for each forward/backward pass — tiny updates accumulate safely in full precision. Second, scale the loss by a large constant before so that small gradients shift into FP16's representable range, then unscale before the optimizer step. Third, accumulate FP16 dot-product partial results into FP32 before writing back — on Volta GPUs do this natively. Together these techniques nearly halve memory and speed up training 2–3× with no hyperparameter changes.

The impact

became the standard way to train large models. Every major framework — PyTorch AMP, TensorFlow mixed precision, DeepSpeed — adopted it. Without this technique, training billion-parameter models like GPT-3 would have required roughly twice the GPU memory and time, making the scaling revolution far more expensive. It directly enabled Megatron-LM, ZeRO, and the entire large-model ecosystem.

Imagine a team of architects designing a skyscraper. The lead architect keeps one pristine, high-resolution blueprint (FP32 ). Each morning, she hands the construction crew a photocopied sketch (FP16 copy) — good enough for cutting steel and pouring concrete all day. At sunset the crew reports what changed; she pencils those tiny adjustments onto her master blueprint where the fine grid lines let her place them precisely.

Without the master blueprint, rounding errors in each day's photocopy would drift the building off-center by millimeters that compound into meters. With it, the crew works twice as fast with half the paper, and the finished tower is indistinguishable from one built entirely with fine blueprints.

The problem: FP32 is accurate but expensive

Standard neural network training stores every weight, activation, and gradient as a 32-bit float. That means 4 bytes per number. A model with 100 million parameters already needs 400 MB just for the weights — and activations saved for backpropagation multiply that cost by the and depth of the network.

Modern GPUs offer a faster lane: half-precision (FP16) arithmetic runs 2–8× faster and uses half the memory. But FP16 has only 10 mantissa bits (vs. 23 in FP32) and a much smaller exponent range — values below about 6×10⁻⁸ collapse to zero. Naively converting everything to FP16 causes two failures:

  • Vanishing updates: when a weight gradient is multiplied by a small , the result often falls below FP16's minimum representable value and rounds to zero. About 5% of all gradient values in a typical model lie in this danger zone.
  • Swamped additions: even when a gradient is representable, adding it to a weight that is 2048× larger causes the smaller value to be right-shifted out of the mantissa and discarded.
Open in Lab
Compare the bit layouts of FP32 and FP16. Notice how FP16's smaller exponent and mantissa limit both range and precision.
The demo wakes as you arrive…

The recipe: three techniques, zero accuracy loss

The paper proposes three complementary techniques that together let you do almost all arithmetic in FP16 while matching FP32 accuracy. Think of them as a safety net with three layers — each layer catches errors that slip through the others.

Open in Lab
Step through a complete mixed precision training iteration. Watch how the three techniques work together in each phase.
The demo wakes as you arrive…

Loss scaling: why gradients vanish in FP16

The heart of the problem is a mismatch between where gradient values live and what FP16 can represent. The paper shows a histogram of activation gradients from an SSD object detector trained in FP32: 67% of gradient values were exactly zero, and many of the remaining values had exponents in the range [−32, −20] — well below FP16's minimum normalized exponent of −14.

Without loss scaling, these small-but-important gradients round to zero in FP16 and the model diverges. With a scaling factor of just 8 (shifting exponents up by 3), the gradients slide into FP16's sweet spot and training matches FP32 perfectly.

Open in Lab
Drag the loss scale slider to see how gradients shift from the "underflow zone" into FP16's representable range.
The demo wakes as you arrive…

How to choose the scale factor? The simplest approach is a fixed constant — the paper found that values between 8 and 32K work across most networks. A more robust approach is : start with a large scale, and if an overflow (infinity or NaN) is detected in the gradients, skip that update and halve the scale. If several consecutive steps succeed, double it again. This way the scale automatically adapts to the model's gradient distribution without any manual tuning.

L^=S⋅L⇒∂L^∂w=S⋅∂L∂w\hat{L} = S \cdot L \quad \Rightarrow \quad \frac{\partial \hat{L}}{\partial w} = S \cdot \frac{\partial L}{\partial w}
Loss scaling — multiplying the loss shifts all gradients equally — S = scale factor · L = original loss · Ŝ·∂L/∂w = the scaled gradient that stays in FP16 range · before the optimizer step, divide by S to restore the true gradient magnitude

Memory savings — where does the benefit come from?

A common concern: if you keep FP32 master weights plus FP16 copies, aren't you using more memory than pure FP32? For the weights alone, yes — 50% more. But weights are a small fraction of total training memory. The dominant cost is activations: each layer's output must be saved for backpropagation, and with large batch sizes these dwarf the weights. Since activations and their gradients are now stored in FP16 (2 bytes instead of 4), total memory drops by roughly 40–50%.

The net result: you can either train the same model faster, or fit a larger batch size into the same GPU memory — both improve wall-clock training time.

Open in Lab
Compare GPU memory usage between pure FP32 training and mixed precision. Drag the model size slider to see how the gap grows.
The demo wakes as you arrive…

Which operations stay in FP32?

Not every operation can safely run in pure FP16. The paper categorizes neural network arithmetic into three buckets and prescribes different precision for each:

  • Dot products (convolutions, matrix multiplies in fully-connected and recurrent layers): multiply in FP16, accumulate partial products in FP32. Tensor Cores on Volta GPUs do this natively — no speed penalty.
  • Large reductions (batch normalization statistics, softmax sums): compute in FP32. These layers are memory-bandwidth-limited anyway, so arithmetic precision doesn't affect speed. They still read/write FP16 tensors.
  • Point-wise operations (ReLU, element-wise multiply, bias addition): either precision works. These are memory-bound, so FP16 halves the bandwidth cost with no accuracy risk.
Open in Lab
Click each operation category to see why it uses a specific precision strategy.
The demo wakes as you arrive…

Mixed precision in code

Modern frameworks make mixed precision training a three-line addition. PyTorch's torch.cuda.amp module provides autocast (which selects FP16 or FP32 per operation automatically) and GradScaler (which handles dynamic loss scaling). Here's the essential pattern:

PyTorch mixed precision training looppython

Simplified to show the idea — not the real implementation.

import torch
from torch.cuda.amp import autocast, GradScaler

model = MyModel().cuda()
optimizer = torch.optim.Adam(model.parameters(), lr=1e-3)
scaler = GradScaler()       # handles dynamic loss scaling

for inputs, targets in dataloader:
    optimizer.zero_grad()

    with autocast():         # FP16 forward pass, FP32 where needed
        outputs = model(inputs)
        loss = criterion(outputs, targets)

    scaler.scale(loss).backward()   # scaled FP16 backward pass
    scaler.step(optimizer)          # unscale gradients, update FP32 weights
    scaler.update()                # adjust scale factor for next step

Results — no accuracy loss, across everything

The paper demonstrates mixed precision training across six application domains. In every case, mixed precision matched or slightly exceeded FP32 accuracy with identical hyperparameters:

  • Image classification: AlexNet, VGG, GoogLeNet, Inception v2/v3, ResNet-50 — all matched FP32 on ILSVRC. No loss scaling needed.
  • Object detection: Faster R-CNN and Multibox SSD. SSD diverged without loss scaling but matched FP32 with a scale factor of just 8.
  • Speech recognition: DeepSpeech 2 models with 115M–215M parameters. Mixed precision actually achieved 5–10% better error rates, suggesting FP16 acts as a regularizer.
  • Machine translation: 3-layer and 5-layer LSTM -decoders. Required loss scaling to match FP32.
  • Language modeling: bigLSTM with 8192-unit layers on 1B Word dataset. Diverged without loss scaling; factor of 128 recovered full accuracy.
  • Image generation: DCGAN for 128×128 face generation. Matched FP32 quality with no loss scaling.
Open in Lab
FP32 vs Mixed Precision accuracy across all tested models. Hover over any bar to see the exact numbers.
The demo wakes as you arrive…

Why it changed everything

After this paper, mixed precision became a non-negotiable part of the training stack. PyTorch integrated it as torch.cuda.amp, TensorFlow added tf.keras.mixed_precision, and NVIDIA's Apex library provided early adopters with plug-and-play support. Systems like Megatron-LM and ZeRO built directly on mixed precision as a foundational layer.

The paper's insight generalizes beyond FP16: the same master-weight + scaling pattern applies to , FP8, and other reduced-precision formats that power today's largest models. Every time you hear about a model trained on thousands of GPUs, mixed precision is running underneath.

  1. 2015

    Gupta et al. — 16-bit fixed point

    Demonstrated 16-bit fixed point training on small datasets (MNIST, CIFAR-10), but the approach didn't scale to large CNNs or RNNs.

  2. 2018

    Mixed Precision Training (this paper)

    Master weights + loss scaling + FP32 accumulation. First method to match FP32 accuracy across CNNs, RNNs, GANs, and multiple domains with no hyperparameter changes.

  3. 2019

    NVIDIA Apex & PyTorch AMP

    Framework-level integration made mixed precision a three-line code change. Automatic mixed precision became the standard way to train.

  4. 2020

    Megatron-LM & ZeRO

    Large-model training systems built on mixed precision as a foundational layer, enabling models with billions of parameters on clusters of GPUs.

  5. 2022

    FP8 Training

    The same master-weight + scaling pattern extended to 8-bit floats on NVIDIA Hopper GPUs, further doubling throughput.

CitationMicikevicius, Narang, Alben, Diamos, Elsen, Garcia, Ginsburg, Houston, Kuchaiev, Venkatesh, Wu. Mixed Precision Training. ICLR, 2018.

Terms in this paper