Training Methodology2018intermediate10 min read
Mixed Precision Training
التدريب بالدقة المختلطة
Micikevicius, P. · Narang, S. · Alben, J. · Diamos, G. · Elsen, E. · Garcia, D. · Ginsburg, B. · Houston, M. · Kuchaiev, O. · Venkatesh, G. · Wu, H. — ICLR
The problem
Modern deep neural networks keep growing — more layers, more parameters, more data. them in standard 32-bit floating point (FP32) is expensive: every , activation, and occupies 4 bytes of GPU memory, and arithmetic on 32-bit numbers is slower than it needs to be. GPUs can do () math 2–8× faster, but FP16 has a much narrower : values below 2⁻²⁴ become zero, and small gradient updates vanish. Simply switching everything to FP16 causes models to diverge.
The contribution
Three techniques that let you train in FP16 without losing . First, keep an FP32 "master copy" of the weights and round it to FP16 for each forward/backward pass — tiny updates accumulate safely in full precision. Second, scale the loss by a large constant before so that small gradients shift into FP16's representable range, then unscale before the optimizer step. Third, accumulate FP16 dot-product partial results into FP32 before writing back — on Volta GPUs do this natively. Together these techniques nearly halve memory and speed up training 2–3× with no hyperparameter changes.
The impact
became the standard way to train large models. Every major framework — PyTorch AMP, TensorFlow mixed precision, DeepSpeed — adopted it. Without this technique, training billion-parameter models like GPT-3 would have required roughly twice the GPU memory and time, making the scaling revolution far more expensive. It directly enabled Megatron-LM, ZeRO, and the entire large-model ecosystem.
Imagine a team of architects designing a skyscraper. The lead architect keeps one pristine, high-resolution blueprint (FP32 ). Each morning, she hands the construction crew a photocopied sketch (FP16 copy) — good enough for cutting steel and pouring concrete all day. At sunset the crew reports what changed; she pencils those tiny adjustments onto her master blueprint where the fine grid lines let her place them precisely.
Without the master blueprint, rounding errors in each day's photocopy would drift the building off-center by millimeters that compound into meters. With it, the crew works twice as fast with half the paper, and the finished tower is indistinguishable from one built entirely with fine blueprints.
The problem: FP32 is accurate but expensive
Standard neural network training stores every weight, activation, and gradient as a 32-bit float. That means 4 bytes per number. A model with 100 million parameters already needs 400 MB just for the weights — and activations saved for backpropagation multiply that cost by the and depth of the network.
Modern GPUs offer a faster lane: half-precision (FP16) arithmetic runs 2–8× faster and uses half the memory. But FP16 has only 10 mantissa bits (vs. 23 in FP32) and a much smaller exponent range — values below about 6×10⁻⁸ collapse to zero. Naively converting everything to FP16 causes two failures:
- Vanishing updates: when a weight gradient is multiplied by a small , the result often falls below FP16's minimum representable value and rounds to zero. About 5% of all gradient values in a typical model lie in this danger zone.
- Swamped additions: even when a gradient is representable, adding it to a weight that is 2048× larger causes the smaller value to be right-shifted out of the mantissa and discarded.
The recipe: three techniques, zero accuracy loss
The paper proposes three complementary techniques that together let you do almost all arithmetic in FP16 while matching FP32 accuracy. Think of them as a safety net with three layers — each layer catches errors that slip through the others.
Loss scaling: why gradients vanish in FP16
The heart of the problem is a mismatch between where gradient values live and what FP16 can represent. The paper shows a histogram of activation gradients from an SSD object detector trained in FP32: 67% of gradient values were exactly zero, and many of the remaining values had exponents in the range [−32, −20] — well below FP16's minimum normalized exponent of −14.
Without loss scaling, these small-but-important gradients round to zero in FP16 and the model diverges. With a scaling factor of just 8 (shifting exponents up by 3), the gradients slide into FP16's sweet spot and training matches FP32 perfectly.
How to choose the scale factor? The simplest approach is a fixed constant — the paper found that values between 8 and 32K work across most networks. A more robust approach is : start with a large scale, and if an overflow (infinity or NaN) is detected in the gradients, skip that update and halve the scale. If several consecutive steps succeed, double it again. This way the scale automatically adapts to the model's gradient distribution without any manual tuning.
Memory savings — where does the benefit come from?
A common concern: if you keep FP32 master weights plus FP16 copies, aren't you using more memory than pure FP32? For the weights alone, yes — 50% more. But weights are a small fraction of total training memory. The dominant cost is activations: each layer's output must be saved for backpropagation, and with large batch sizes these dwarf the weights. Since activations and their gradients are now stored in FP16 (2 bytes instead of 4), total memory drops by roughly 40–50%.
The net result: you can either train the same model faster, or fit a larger batch size into the same GPU memory — both improve wall-clock training time.
Which operations stay in FP32?
Not every operation can safely run in pure FP16. The paper categorizes neural network arithmetic into three buckets and prescribes different precision for each:
- Dot products (convolutions, matrix multiplies in fully-connected and recurrent layers): multiply in FP16, accumulate partial products in FP32. Tensor Cores on Volta GPUs do this natively — no speed penalty.
- Large reductions (batch normalization statistics, softmax sums): compute in FP32. These layers are memory-bandwidth-limited anyway, so arithmetic precision doesn't affect speed. They still read/write FP16 tensors.
- Point-wise operations (ReLU, element-wise multiply, bias addition): either precision works. These are memory-bound, so FP16 halves the bandwidth cost with no accuracy risk.
Mixed precision in code
Modern frameworks make mixed precision training a three-line addition. PyTorch's torch.cuda.amp module provides autocast (which selects FP16 or FP32 per operation automatically) and GradScaler (which handles dynamic loss scaling). Here's the essential pattern:
Simplified to show the idea — not the real implementation.
import torch
from torch.cuda.amp import autocast, GradScaler
model = MyModel().cuda()
optimizer = torch.optim.Adam(model.parameters(), lr=1e-3)
scaler = GradScaler() # handles dynamic loss scaling
for inputs, targets in dataloader:
optimizer.zero_grad()
with autocast(): # FP16 forward pass, FP32 where needed
outputs = model(inputs)
loss = criterion(outputs, targets)
scaler.scale(loss).backward() # scaled FP16 backward pass
scaler.step(optimizer) # unscale gradients, update FP32 weights
scaler.update() # adjust scale factor for next stepResults — no accuracy loss, across everything
The paper demonstrates mixed precision training across six application domains. In every case, mixed precision matched or slightly exceeded FP32 accuracy with identical hyperparameters:
- Image classification: AlexNet, VGG, GoogLeNet, Inception v2/v3, ResNet-50 — all matched FP32 on ILSVRC. No loss scaling needed.
- Object detection: Faster R-CNN and Multibox SSD. SSD diverged without loss scaling but matched FP32 with a scale factor of just 8.
- Speech recognition: DeepSpeech 2 models with 115M–215M parameters. Mixed precision actually achieved 5–10% better error rates, suggesting FP16 acts as a regularizer.
- Machine translation: 3-layer and 5-layer LSTM -decoders. Required loss scaling to match FP32.
- Language modeling: bigLSTM with 8192-unit layers on 1B Word dataset. Diverged without loss scaling; factor of 128 recovered full accuracy.
- Image generation: DCGAN for 128×128 face generation. Matched FP32 quality with no loss scaling.
Why it changed everything
After this paper, mixed precision became a non-negotiable part of the training stack. PyTorch integrated it as torch.cuda.amp, TensorFlow added tf.keras.mixed_precision, and NVIDIA's Apex library provided early adopters with plug-and-play support. Systems like Megatron-LM and ZeRO built directly on mixed precision as a foundational layer.
The paper's insight generalizes beyond FP16: the same master-weight + scaling pattern applies to , FP8, and other reduced-precision formats that power today's largest models. Every time you hear about a model trained on thousands of GPUs, mixed precision is running underneath.
2015
Gupta et al. — 16-bit fixed point
Demonstrated 16-bit fixed point training on small datasets (MNIST, CIFAR-10), but the approach didn't scale to large CNNs or RNNs.
2018
Mixed Precision Training (this paper)
Master weights + loss scaling + FP32 accumulation. First method to match FP32 accuracy across CNNs, RNNs, GANs, and multiple domains with no hyperparameter changes.
2019
NVIDIA Apex & PyTorch AMP
Framework-level integration made mixed precision a three-line code change. Automatic mixed precision became the standard way to train.
2020
Megatron-LM & ZeRO
Large-model training systems built on mixed precision as a foundational layer, enabling models with billions of parameters on clusters of GPUs.
2022
FP8 Training
The same master-weight + scaling pattern extended to 8-bit floats on NVIDIA Hopper GPUs, further doubling throughput.
CitationMicikevicius, Narang, Alben, Diamos, Elsen, Garcia, Ginsburg, Houston, Kuchaiev, Venkatesh, Wu. Mixed Precision Training. ICLR, 2018.
Terms in this paper
- Mixed Precision Trainingالتدريب بالدقة المختلطة
- FP16الدقة العائمة بنظام 16 بت
- Loss Scalingتحجيم الخسارة
- Master Weightsالأوزان الرئيسية
- Gradient Underflowالطفح السفلي للتدرّجات
- Tensor Coresنوى Tensor Cores
- Half-Precisionنصف الدقة
- Dynamic Rangeالنطاق الديناميكي