Core ML2019intermediate11 min read

When Does Label Smoothing Help?

متى يُفيد تنعيم التسميات؟

Müller, R. · Kornblith, S. · Hinton, G. — NeurIPS

The problem

— replacing hard one-hot targets with a weighted mixture of the true label and a uniform distribution — had become a standard trick in state-of-the-art image classifiers, speech recognizers, and machine translation systems. Yet no one understood why it worked, what it did to internal representations, or when it could actually hurt performance.

The contribution

Three key findings. First, label smoothing encourages the to form tight, equally-spaced clusters — each class maps to a compact region equidistant from all other classes. Second, this clustering implicitly calibrates the model, reducing the without post-hoc scaling. Third, this same clustering erases inter-class similarity information from the logits, which makes from label-smoothed teachers significantly less effective. The authors introduced a novel visualization method for penultimate layer activations and quantified information erasure via estimation.

The impact

This paper transformed label smoothing from a mysterious recipe ingredient into a well-understood tool. It revealed that smoothing provides implicit — a finding that influenced how practitioners deploy models in safety-critical applications where prediction confidence must be trustworthy. The discovery that label smoothing hurts knowledge distillation directly influenced pipelines: teams now avoid smoothing when training teacher networks intended for distillation. The visualization technique for penultimate layer representations became a widely adopted diagnostic tool.

Imagine a driving instructor grading a student's parking. A strict instructor says "Spot 3 — correct!" and marks every other spot as completely wrong. The student learns to park perfectly in spot 3 — but has no idea that spot 2 is almost as good, while spot 47 on the other side of the lot would be terrible.

A label-smoothing instructor says "Spot 3 gets 95% credit, and every other spot gets a tiny sliver of credit." Now the student learns something subtler: all other spots are roughly equally wrong. The student parks more calmly, and their confidence matches reality.

But there's a catch. When this student becomes a driving instructor themselves (knowledge distillation), they have lost the knowledge that "spot 2 is almost right." They can only say "spot 3 is right, everything else is equally wrong" — and their own students learn less from them than from the strict instructor.

What is label smoothing?

In standard classification, we train neural networks using hard targets: for a cat image, the target vector is [1,0,0,…,0][1, 0, 0, \ldots, 0] — full confidence in the correct class, zero everywhere else. This is . The network minimizes cross- between these targets and its outputs.

The problem is that hard targets push the network toward extreme confidence. To make the predicted probability of the correct class approach 1, the network must drive the corresponding toward infinity. This leads to : the network memorizes training examples rather than learning generalizable patterns, and its confidence becomes poorly calibrated — it says "99.9% cat" even when the evidence is ambiguous.

Label smoothing, introduced by Szegedy et al. in the Inception-v2 architecture, offers a simple fix. Instead of a hard target yky_k, we use a smoothed target ykLSy_k^{LS} that mixes the hard target with a uniform distribution. With a smoothing parameter α\alpha, the correct class gets 1−α+α/K1 - \alpha + \alpha/K and each incorrect class gets α/K\alpha/K, where KK is the total number of classes.

ykLS=yk(1−α)+αKy_k^{LS} = y_k(1 - \alpha) + \frac{\alpha}{K}
Label smoothing target — mixing one-hot with uniform distribution — For a training example with true class cc: the correct class gets target 1−α+α/K1 - \alpha + \alpha/K (slightly less than 1), and each of the K−1K-1 incorrect classes gets α/K\alpha/K (slightly more than 0). With α=0.1\alpha = 0.1 and K=10K = 10, the correct class target becomes 0.91 and each incorrect class gets 0.01.
Open in Lab
Drag the smoothing slider α to see how label smoothing redistributes probability mass from the correct class to all other classes.
The demo wakes as you arrive…

Tight clusters: what label smoothing does to representations

The paper's most striking finding is a geometric one. When a network is trained with hard targets, the penultimate layer representations form broad, overlapping clouds for each class. Examples from the same class can sit at very different distances from other class templates, preserving rich information about inter-class similarity.

With label smoothing, the picture changes dramatically. The smoothed targets encourage every example to be equally distant from all incorrect class templates. The result is that representations collapse into tight, compact clusters arranged in a regular geometric pattern — like vertices of an equilateral triangle (for 3 classes projected onto 2D).

Think of it as the difference between a parking lot where cars scatter freely within each section versus one with assigned spots — every car in its exact position, equidistant from neighboring sections. The parking lot is more orderly, but you lose the information about which cars were "almost" in the wrong section.

To visualize this, the authors proposed a novel technique. They pick three classes, find the 2D plane passing through the three class templates (the weight vectors wkw_k of the last layer), and project the penultimate layer activations onto that plane. This reveals the clustering geometry directly.

The logit for class kk is xTwkx^T w_k, which can be related to the squared Euclidean distance between the activation vector xx and the template wkw_k:

∥x−wk∥2=xTx−2xTwk+wkTwk\|x - w_k\|^2 = x^T x - 2 x^T w_k + w_k^T w_k
Logit as distance — connecting activations to class templates — Since xTxx^T x factors out in the softmax and wkTwkw_k^T w_k is approximately constant across classes, the logit xTwkx^T w_k acts as a (negated) squared distance to the class template. Label smoothing forces all distances to incorrect templates to be equal — hence the tight, equidistant clustering.
Open in Lab
Toggle between hard targets and label smoothing to see how representations cluster. With hard targets, clusters are broad and overlap; with label smoothing, they snap into tight, equidistant formations.
The demo wakes as you arrive…

Implicit calibration: confidence that matches reality

A model is well-calibrated if, among all predictions it makes with 80% confidence, approximately 80% are actually correct. Modern neural networks are notoriously over-confident: they routinely predict 99% confidence on examples they get wrong.

Guo et al. (2017) showed that post-hoc temperature scaling — dividing logits by a learned temperature TT before the softmax — can fix calibration. But this requires a held-out validation set and a separate calibration step.

Müller et al. discovered that label smoothing provides a similar calibration benefit for free, as a side effect of training. The expected calibration error (ECE) drops substantially without any post-processing. On CIFAR-100 with ResNet-56, label smoothing with α=0.05\alpha = 0.05 reduces ECE from 0.150 to 0.024 — nearly matching temperature scaling at T=1.9T = 1.9 which achieves 0.021.

Open in Lab
Compare reliability diagrams: hard targets (over-confident), temperature scaling (post-hoc fix), and label smoothing (implicit calibration). The diagonal line represents perfect calibration.
The demo wakes as you arrive…

The calibration benefit has a direct practical payoff in machine translation. Vaswani et al. had reported that label smoothing with α=0.1\alpha = 0.1 improved BLEU scores in their model despite worse perplexity. This seems paradoxical — how can a model with worse predictions produce better translations?

The answer is calibration. , the decoding algorithm used in translation, relies on model confidence to rank candidate sequences. An over-confident model assigns disproportionately high probability to its first guess, preventing beam search from exploring better alternatives. A well-calibrated model spreads probability more truthfully, giving beam search better signal to work with.

The dark side: label smoothing hurts knowledge distillation

Knowledge distillation trains a smaller student network to mimic the soft outputs of a larger teacher network. The key idea is that the teacher's soft predictions carry "dark knowledge" — information about which incorrect classes are more similar to each other. A teacher that gives 10% probability to "dog" and 0.1% to "truck" when classifying a cat image is telling the student that cats look somewhat like dogs but nothing like trucks.

Label smoothing destroys exactly this information. By forcing the penultimate layer to place every example equidistant from all incorrect class templates, label smoothing ensures that the teacher's soft predictions treat all wrong classes as equally wrong. The "dark knowledge" about inter-class similarities is erased.

The experiment is striking: a ResNet-56 teacher trained with label smoothing on CIFAR-10 achieves better accuracy than one trained with hard targets. But when this better teacher distills knowledge into an AlexNet student, the student performs worse than one trained with the hard-target teacher. A better teacher produces a worse student — because the teacher's logits have lost the relative information that makes distillation effective.

Open in Lab
Explore the distillation paradox: a label-smoothed teacher has better accuracy but produces worse students. Toggle between teacher and student performance to see the tradeoff.
The demo wakes as you arrive…

Information erasure: quantifying what is lost

To go beyond visual intuition, the authors quantified the information erasure by estimating the mutual information between training examples and their logit differences. Mutual information measures how much knowing the input tells you about the output beyond just knowing the class label.

For a network trained with hard targets, the mutual information stays high throughout training. Different training examples from the same class produce different logit patterns — a cat lying in grass has different inter-class similarities than a cat sitting on a chair. This example-specific information is what makes distillation powerful.

For a network trained with label smoothing, the mutual information rises initially (as the network learns to classify) but then drops sharply. Eventually it approaches log⁡(2)\log(2) — the information content of a single binary label, meaning the logits carry almost no information beyond "correct class vs everything else." Every cat looks the same in the logit space, regardless of context.

Open in Lab
Watch mutual information evolve during training. With hard targets, it stays high — each example retains its identity. With label smoothing, it collapses toward the single-bit threshold.
The demo wakes as you arrive…

Implementing label smoothing

Label smoothing cross-entropy in PyTorchpython

Simplified to show the idea — not the real implementation.

import torch
import torch.nn.functional as F

def label_smoothing_loss(logits, targets, alpha=0.1):
    """Cross-entropy with label smoothing.
    logits:  (batch, num_classes) — raw model outputs
    targets: (batch,)            — integer class labels
    alpha:   smoothing factor (0 = hard targets, 1 = uniform)
    """
    K = logits.size(-1)  # number of classes

    # Standard cross-entropy with the true label
    hard_loss = F.cross_entropy(logits, targets)

    # KL divergence from uniform distribution
    # log_softmax gives log p_k, uniform has log(1/K) for each class
    log_probs = F.log_softmax(logits, dim=-1)
    uniform_loss = -log_probs.mean(dim=-1).mean()

    # Blend: (1 - alpha) * hard + alpha * uniform
    return (1 - alpha) * hard_loss + alpha * uniform_loss

Results across domains

The paper validates its findings across three domains, three architectures, and three datasets. In image classification, AlexNet on CIFAR-10 and ResNet-56 on CIFAR-100 both showed the tight clustering phenomenon with label smoothing. In machine translation, the Transformer on English-German showed both improved BLEU scores and better calibration with label smoothing. On ImageNet with Inception-v4, label smoothing reduced ECE from 0.071 to 0.035.

An especially interesting case is semantically similar classes on ImageNet. When visualizing two poodle breeds (toy poodle and miniature poodle) together with a dissimilar class (tench fish), hard targets produce overlapping clusters for the similar classes with a continuous gradient of similarity toward the dissimilar class. Label smoothing forces even these similar classes into an arc shape, erasing the gradient of similarity. You can no longer measure "how much this poodle resembles a tench" — a piece of information that hard-target networks naturally preserve.

Timeline: label smoothing in context

  1. 2015

    Knowledge Distillation (Hinton et al.)

    Introduced the idea of training a small student network to mimic a large teacher's soft predictions. Temperature scaling of softmax outputs became the standard tool for extracting "dark knowledge."

  2. 2016

    Label smoothing in Inception-v2 (Szegedy et al.)

    First introduced label smoothing as a regularization technique. It improved ImageNet top-1 error from 23.1% to 22.8% — a small but consistent gain that became standard in subsequent architectures.

  3. 2017

    Transformer with label smoothing (Vaswani et al.)

    The Transformer architecture used label smoothing with α=0.1, achieving better BLEU scores despite worse perplexity — a puzzle that this paper later explained through the calibration lens.

  4. 2017

    Confidence penalty (Pereyra et al.)

    Showed that penalizing low-entropy output distributions is equivalent to label smoothing when the KL divergence direction is reversed. Proposed using non-uniform distributions as targets.

  5. 2019

    This paper (Müller, Kornblith & Hinton)

    Explained label smoothing through penultimate layer visualization, showed implicit calibration, and discovered the distillation-harming effect of information erasure.

  6. 2020

    Mixup and beyond

    Soft-target methods diversified. Mixup, CutMix, and other data augmentation techniques provided alternative ways to smooth targets while preserving inter-class information — partly addressing the distillation limitation identified by this paper.

CitationMüller, R., Kornblith, S. & Hinton, G.. When Does Label Smoothing Help?. NeurIPS, 2019.

Terms in this paper