Core ML2019intermediate11 min read
When Does Label Smoothing Help?
متى يُفيد تنعيم التسميات؟
Müller, R. · Kornblith, S. · Hinton, G. — NeurIPS
The problem
— replacing hard one-hot targets with a weighted mixture of the true label and a uniform distribution — had become a standard trick in state-of-the-art image classifiers, speech recognizers, and machine translation systems. Yet no one understood why it worked, what it did to internal representations, or when it could actually hurt performance.
The contribution
Three key findings. First, label smoothing encourages the to form tight, equally-spaced clusters — each class maps to a compact region equidistant from all other classes. Second, this clustering implicitly calibrates the model, reducing the without post-hoc scaling. Third, this same clustering erases inter-class similarity information from the logits, which makes from label-smoothed teachers significantly less effective. The authors introduced a novel visualization method for penultimate layer activations and quantified information erasure via estimation.
The impact
This paper transformed label smoothing from a mysterious recipe ingredient into a well-understood tool. It revealed that smoothing provides implicit — a finding that influenced how practitioners deploy models in safety-critical applications where prediction confidence must be trustworthy. The discovery that label smoothing hurts knowledge distillation directly influenced pipelines: teams now avoid smoothing when training teacher networks intended for distillation. The visualization technique for penultimate layer representations became a widely adopted diagnostic tool.
Imagine a driving instructor grading a student's parking. A strict instructor says "Spot 3 — correct!" and marks every other spot as completely wrong. The student learns to park perfectly in spot 3 — but has no idea that spot 2 is almost as good, while spot 47 on the other side of the lot would be terrible.
A label-smoothing instructor says "Spot 3 gets 95% credit, and every other spot gets a tiny sliver of credit." Now the student learns something subtler: all other spots are roughly equally wrong. The student parks more calmly, and their confidence matches reality.
But there's a catch. When this student becomes a driving instructor themselves (knowledge distillation), they have lost the knowledge that "spot 2 is almost right." They can only say "spot 3 is right, everything else is equally wrong" — and their own students learn less from them than from the strict instructor.
What is label smoothing?
In standard classification, we train neural networks using hard targets: for a cat image, the target vector is — full confidence in the correct class, zero everywhere else. This is . The network minimizes cross- between these targets and its outputs.
The problem is that hard targets push the network toward extreme confidence. To make the predicted probability of the correct class approach 1, the network must drive the corresponding toward infinity. This leads to : the network memorizes training examples rather than learning generalizable patterns, and its confidence becomes poorly calibrated — it says "99.9% cat" even when the evidence is ambiguous.
Label smoothing, introduced by Szegedy et al. in the Inception-v2 architecture, offers a simple fix. Instead of a hard target , we use a smoothed target that mixes the hard target with a uniform distribution. With a smoothing parameter , the correct class gets and each incorrect class gets , where is the total number of classes.
Tight clusters: what label smoothing does to representations
The paper's most striking finding is a geometric one. When a network is trained with hard targets, the penultimate layer representations form broad, overlapping clouds for each class. Examples from the same class can sit at very different distances from other class templates, preserving rich information about inter-class similarity.
With label smoothing, the picture changes dramatically. The smoothed targets encourage every example to be equally distant from all incorrect class templates. The result is that representations collapse into tight, compact clusters arranged in a regular geometric pattern — like vertices of an equilateral triangle (for 3 classes projected onto 2D).
Think of it as the difference between a parking lot where cars scatter freely within each section versus one with assigned spots — every car in its exact position, equidistant from neighboring sections. The parking lot is more orderly, but you lose the information about which cars were "almost" in the wrong section.
To visualize this, the authors proposed a novel technique. They pick three classes, find the 2D plane passing through the three class templates (the weight vectors of the last layer), and project the penultimate layer activations onto that plane. This reveals the clustering geometry directly.
The logit for class is , which can be related to the squared Euclidean distance between the activation vector and the template :
Implicit calibration: confidence that matches reality
A model is well-calibrated if, among all predictions it makes with 80% confidence, approximately 80% are actually correct. Modern neural networks are notoriously over-confident: they routinely predict 99% confidence on examples they get wrong.
Guo et al. (2017) showed that post-hoc temperature scaling — dividing logits by a learned temperature before the softmax — can fix calibration. But this requires a held-out validation set and a separate calibration step.
Müller et al. discovered that label smoothing provides a similar calibration benefit for free, as a side effect of training. The expected calibration error (ECE) drops substantially without any post-processing. On CIFAR-100 with ResNet-56, label smoothing with reduces ECE from 0.150 to 0.024 — nearly matching temperature scaling at which achieves 0.021.
The calibration benefit has a direct practical payoff in machine translation. Vaswani et al. had reported that label smoothing with improved BLEU scores in their model despite worse perplexity. This seems paradoxical — how can a model with worse predictions produce better translations?
The answer is calibration. , the decoding algorithm used in translation, relies on model confidence to rank candidate sequences. An over-confident model assigns disproportionately high probability to its first guess, preventing beam search from exploring better alternatives. A well-calibrated model spreads probability more truthfully, giving beam search better signal to work with.
The dark side: label smoothing hurts knowledge distillation
Knowledge distillation trains a smaller student network to mimic the soft outputs of a larger teacher network. The key idea is that the teacher's soft predictions carry "dark knowledge" — information about which incorrect classes are more similar to each other. A teacher that gives 10% probability to "dog" and 0.1% to "truck" when classifying a cat image is telling the student that cats look somewhat like dogs but nothing like trucks.
Label smoothing destroys exactly this information. By forcing the penultimate layer to place every example equidistant from all incorrect class templates, label smoothing ensures that the teacher's soft predictions treat all wrong classes as equally wrong. The "dark knowledge" about inter-class similarities is erased.
The experiment is striking: a ResNet-56 teacher trained with label smoothing on CIFAR-10 achieves better accuracy than one trained with hard targets. But when this better teacher distills knowledge into an AlexNet student, the student performs worse than one trained with the hard-target teacher. A better teacher produces a worse student — because the teacher's logits have lost the relative information that makes distillation effective.
Information erasure: quantifying what is lost
To go beyond visual intuition, the authors quantified the information erasure by estimating the mutual information between training examples and their logit differences. Mutual information measures how much knowing the input tells you about the output beyond just knowing the class label.
For a network trained with hard targets, the mutual information stays high throughout training. Different training examples from the same class produce different logit patterns — a cat lying in grass has different inter-class similarities than a cat sitting on a chair. This example-specific information is what makes distillation powerful.
For a network trained with label smoothing, the mutual information rises initially (as the network learns to classify) but then drops sharply. Eventually it approaches — the information content of a single binary label, meaning the logits carry almost no information beyond "correct class vs everything else." Every cat looks the same in the logit space, regardless of context.
Implementing label smoothing
Simplified to show the idea — not the real implementation.
import torch
import torch.nn.functional as F
def label_smoothing_loss(logits, targets, alpha=0.1):
"""Cross-entropy with label smoothing.
logits: (batch, num_classes) — raw model outputs
targets: (batch,) — integer class labels
alpha: smoothing factor (0 = hard targets, 1 = uniform)
"""
K = logits.size(-1) # number of classes
# Standard cross-entropy with the true label
hard_loss = F.cross_entropy(logits, targets)
# KL divergence from uniform distribution
# log_softmax gives log p_k, uniform has log(1/K) for each class
log_probs = F.log_softmax(logits, dim=-1)
uniform_loss = -log_probs.mean(dim=-1).mean()
# Blend: (1 - alpha) * hard + alpha * uniform
return (1 - alpha) * hard_loss + alpha * uniform_lossResults across domains
The paper validates its findings across three domains, three architectures, and three datasets. In image classification, AlexNet on CIFAR-10 and ResNet-56 on CIFAR-100 both showed the tight clustering phenomenon with label smoothing. In machine translation, the Transformer on English-German showed both improved BLEU scores and better calibration with label smoothing. On ImageNet with Inception-v4, label smoothing reduced ECE from 0.071 to 0.035.
An especially interesting case is semantically similar classes on ImageNet. When visualizing two poodle breeds (toy poodle and miniature poodle) together with a dissimilar class (tench fish), hard targets produce overlapping clusters for the similar classes with a continuous gradient of similarity toward the dissimilar class. Label smoothing forces even these similar classes into an arc shape, erasing the gradient of similarity. You can no longer measure "how much this poodle resembles a tench" — a piece of information that hard-target networks naturally preserve.
Timeline: label smoothing in context
2015
Knowledge Distillation (Hinton et al.)
Introduced the idea of training a small student network to mimic a large teacher's soft predictions. Temperature scaling of softmax outputs became the standard tool for extracting "dark knowledge."
2016
Label smoothing in Inception-v2 (Szegedy et al.)
First introduced label smoothing as a regularization technique. It improved ImageNet top-1 error from 23.1% to 22.8% — a small but consistent gain that became standard in subsequent architectures.
2017
Transformer with label smoothing (Vaswani et al.)
The Transformer architecture used label smoothing with α=0.1, achieving better BLEU scores despite worse perplexity — a puzzle that this paper later explained through the calibration lens.
2017
Confidence penalty (Pereyra et al.)
Showed that penalizing low-entropy output distributions is equivalent to label smoothing when the KL divergence direction is reversed. Proposed using non-uniform distributions as targets.
2019
This paper (Müller, Kornblith & Hinton)
Explained label smoothing through penultimate layer visualization, showed implicit calibration, and discovered the distillation-harming effect of information erasure.
2020
Mixup and beyond
Soft-target methods diversified. Mixup, CutMix, and other data augmentation techniques provided alternative ways to smooth targets while preserving inter-class information — partly addressing the distillation limitation identified by this paper.
CitationMüller, R., Kornblith, S. & Hinton, G.. When Does Label Smoothing Help?. NeurIPS, 2019.
Terms in this paper
- Label Smoothingتنعيم التسميات
- Soft Targetsالأهداف المرنة
- Calibrationالمعايرة
- Knowledge Distillationتقطير المعرفة
- Cross Entropyالعشوائية المتقاطعة
- Softmaxسوفت ماكس
- Temperatureالحرارة
- One-Hot Encodingالترميز الأحادي الساخن
- Logitالدرجة الخام
- Overfittingفرط التخصيص
- Regularizationالضبط الهيكلي
- Generalizationالتعميم
- Confidence Scoreدرجة الثقة
- Expected Calibration Errorخطأ المعايرة المتوقع
- Entropyالعشوائية الدلالية