Information Theory1951intermediate8 min read

On Information and Sufficiency

في المعلومات والكفاية الإحصائية

Kullback, S. · Leibler, R. A. — Annals of Mathematical Statistics

The problem

By 1951, Shannon had defined — the minimum number of bits to encode a message — but there was no formal measure for how much information is lost when one is used to approximate another. Statisticians needed a way to quantify: how far is my from reality? And when does a summary of data (a ) preserve all the information the full dataset contains?

The contribution

Kullback and Leibler introduced a measure now called (or relative entropy): for two distributions P and Q, it equals the expected number of extra bits needed when coding data from P using a code optimized for Q. They proved it is always non-negative (Gibbs' inequality), equals zero only when P = Q, and showed that a statistic is sufficient if and only if it preserves the full KL between any two hypotheses. This unified information theory and statistical sufficiency in a single framework.

The impact

KL divergence became one of the most used quantities in modern AI. It is the foundation of cross-entropy in , the regularizer in Variational Autoencoders (VAEs), the objective in t-SNE visualization, a core component of (RLHF), and the theoretical basis for model selection criteria like AIC. Any time a system measures how far a learned distribution is from a target, KL divergence is likely involved.

Imagine you're packing an umbrella for every day of the year based on a weather forecast. If the forecast is perfect, you pack exactly right — an umbrella on rainy days, nothing on sunny ones. But if the forecast says "10% rain" when the truth is "60% rain", you'll get soaked repeatedly.

KL divergence measures this mismatch: the extra "surprise" you experience, on average, because you prepared for the wrong weather pattern. It's not about any single day — it's about the systematic cost of believing the wrong distribution.

Zero divergence means your forecast is reality. The bigger the divergence, the more umbrellas you're missing.

The question: how wrong is my model?

In 1948, Claude Shannon showed that a source producing symbols according to distribution P needs at least H(P)H(P) bits per symbol on average — this is entropy. But what if you don't know P and instead design your code for a different distribution Q? You'll need H(P,Q)H(P, Q) bits — cross-entropy — which is always at least as large as H(P)H(P).

The gap H(P,Q)−H(P)H(P, Q) - H(P) is the wasted bits: the cost of using the wrong model. Kullback and Leibler formalized this gap as a single quantity and gave it rigorous properties, connecting Shannon's information theory to the classical statistics of hypothesis testing and sufficient statistics.

Open in Lab
Drag the Q slider to change the approximate distribution. Watch the cross-entropy gap grow as Q deviates from P.
The demo wakes as you arrive…

The formula: measuring information loss

Before seeing the math, here's the idea in words: for each possible outcome, calculate how much more surprised you are under Q than under P. Weight that extra surprise by how often the outcome actually happens (under P). Sum it all up. That sum is KL divergence — the average wasted surprise.

DKL(P∥Q)=∑xP(x)log⁡P(x)Q(x)D_{\text{KL}}(P \| Q) = \sum_{x} P(x) \log \frac{P(x)}{Q(x)}
KL divergence — the cost of using the wrong distribution — P(x) = how often x really happens · log P(x)/Q(x) = extra surprise per event · the sum = average extra surprise across all events.

For continuous distributions the sum becomes an : DKL(P∥Q)=∫p(x)log⁡p(x)q(x) dxD_{\text{KL}}(P \| Q) = \int p(x) \log \frac{p(x)}{q(x)} \, dx, where pp and qq are the probability density functions. The interpretation is identical: average extra information per sample.

Open in Lab
Adjust P and Q and see the KL divergence computed in real time. Notice how it's not symmetric — D(P‖Q) ≠ D(Q‖P).
The demo wakes as you arrive…

Properties: what KL divergence is (and isn't)

KL divergence has four properties that matter in practice:

  • Non-negativity. DKL(P∥Q)≥0D_{\text{KL}}(P \| Q) \geq 0 always. This is Gibbs' inequality — you can never do better by using the wrong distribution.

  • Zero iff identical. DKL(P∥Q)=0D_{\text{KL}}(P \| Q) = 0 if and only if P=QP = Q almost everywhere. The mismatch cost vanishes only when the distributions match perfectly.

  • Asymmetry. DKL(P∥Q)≠DKL(Q∥P)D_{\text{KL}}(P \| Q) \neq D_{\text{KL}}(Q \| P) in general. KL divergence is not a true distance. The cost of approximating P with Q is different from approximating Q with P.

  • Additivity for independent variables. If P and Q factor over independent components, the total KL divergence is the sum of the individual divergences.

Open in Lab
Switch between D(P‖Q) and D(Q‖P) to see asymmetry in action. The two values can be dramatically different.
The demo wakes as you arrive…

Forward vs reverse KL: mode-covering vs mode-seeking

This asymmetry has a vivid intuition. Suppose P is a mixture of two Gaussians (two peaks) and Q is a single Gaussian.

Forward KL D(P∥Q)D(P \| Q): penalizes Q heavily wherever P has mass but Q doesn't. Q stretches wide to cover both peaks — even if that means putting probability in the valley between them. This is mode-covering behavior.

Reverse KL D(Q∥P)D(Q \| P): penalizes Q wherever Q has mass but P doesn't. Q collapses onto a single peak to avoid placing probability where P is zero. This is mode-seeking behavior.

Understanding this distinction is essential for choosing the right objective in generative models, variational inference, and optimization.

Open in Lab
Toggle between forward and reverse KL. Watch Q's shape change from mode-covering (spread wide) to mode-seeking (collapse onto one peak).
The demo wakes as you arrive…

Sufficiency: when a summary preserves all information

The second half of Kullback and Leibler's paper connects KL divergence to statistical sufficiency. A sufficient statistic is a summary of data that loses nothing relevant to distinguishing between hypotheses.

Think of it like compressing a photo. A lossy compression throws away detail. A sufficient compression throws away only — every bit of signal survives. In formal terms: a statistic T(X)T(X) is sufficient for distinguishing PP from QQ if and only if DKL(PT∥QT)=DKL(P∥Q)D_{\text{KL}}(P_{T} \| Q_{T}) = D_{\text{KL}}(P \| Q). Compressing the data through TT doesn't reduce the divergence — no information is lost.

DKL(PX∥QX)≥DKL(PT(X)∥QT(X))D_{\text{KL}}(P_X \| Q_X) \geq D_{\text{KL}}(P_{T(X)} \| Q_{T(X)})
Data processing inequality — information never increases — Any function T(X) applied to data can only reduce or maintain the divergence, never increase it. Equality holds when and only when T is sufficient.
Open in Lab
Click each compression step to see how divergence changes. The sufficient statistic preserves it fully; a lossy summary does not.
The demo wakes as you arrive…

KL divergence in modern machine learning

KL divergence appears — sometimes explicitly, sometimes disguised — in nearly every corner of modern machine learning:

  • Cross-entropy loss. Minimizing cross-entropy between labels and model predictions is minimizing KL divergence. Every classifier you've trained uses this.

  • Variational Autoencoders. The loss has two terms: a reconstruction loss and DKL(q(z∣x)∥p(z))D_{\text{KL}}(q(z|x) \| p(z)), which pushes the learned posterior toward the prior. This is what makes the smooth and generative.

  • t-SNE. The visualization algorithm minimizes KL divergence between high-dimensional and low-dimensional neighborhood probabilities — this asymmetric cost is why t-SNE preserves local but not global structure.

  • RLHF. A KL penalty term prevents the fine-tuned policy from drifting too far from the base model, balancing helpfulness with stability.

Open in Lab
Click each application to see how KL divergence appears in its loss function.
The demo wakes as you arrive…

The same idea in code

KL divergence for discrete distributions, from scratchpython

Simplified to show the idea — not the real implementation.

import numpy as np

def kl_divergence(p, q):
    """Compute D_KL(P || Q) for discrete distributions P and Q.
    Both must be valid probability arrays that sum to 1."""
    p = np.asarray(p, dtype=float)
    q = np.asarray(q, dtype=float)

    # Only sum where p > 0 (0 * log(0) = 0 by convention)
    mask = p > 0
    return np.sum(p[mask] * np.log(p[mask] / q[mask]))

def cross_entropy(p, q):
    """H(P, Q) = H(P) + D_KL(P || Q)."""
    p = np.asarray(p, dtype=float)
    q = np.asarray(q, dtype=float)
    mask = p > 0
    return -np.sum(p[mask] * np.log(q[mask]))

# Example: fair coin P vs biased coin Q
P = np.array([0.5, 0.5])       # true distribution
Q = np.array([0.9, 0.1])       # our wrong model

print(f"H(P)        = {-np.sum(P * np.log(P)):.4f} nats")
print(f"H(P, Q)     = {cross_entropy(P, Q):.4f} nats")
print(f"D_KL(P||Q)  = {kl_divergence(P, Q):.4f} nats")   # the gap
print(f"D_KL(Q||P)  = {kl_divergence(Q, P):.4f} nats")   # different!

Why it changed everything

  1. 1948

    Shannon's entropy

    Claude Shannon defines information entropy — the minimum bits to encode a source. The mathematical foundation on which KL divergence is built.

  2. 1951

    Kullback-Leibler divergence

    Kullback and Leibler formalize relative entropy and prove the sufficiency theorem. A 7-page paper that seeded decades of research.

  3. 1974

    Akaike Information Criterion (AIC)

    Akaike uses KL divergence as the theoretical basis for model selection — choosing between candidate models by estimating their divergence from truth.

  4. 2006

    Variational Autoencoders (conceptual roots)

    Variational inference methods using KL divergence mature. The KL term becomes a standard regularizer for latent variable models.

  5. 2008

    t-SNE

    van der Maaten and Hinton use KL divergence as the cost function for neighborhood-preserving visualization. The most cited visualization paper in ML.

  6. 2014

    VAE published

    Kingma and Welling introduce the VAE, with KL divergence as the explicit regularizer that keeps the latent space structured and generative.

  7. 2017

    PPO and KL-penalized RL

    Proximal Policy Optimization uses a KL constraint to keep policy updates stable. Later, RLHF for language models adopts the same principle.

  8. 2022

    RLHF for ChatGPT

    KL divergence penalty became central to aligning large language models — a direct descendant of a 1951 statistics paper now shapes AI safety research.

From a 7-page paper in the Annals of Mathematical Statistics to the loss function inside every major language model — KL divergence is the thread connecting Shannon's information theory, Fisher's sufficiency, and the deep learning revolution. Every time a model learns, it is — in some form — closing the KL gap.

CitationKullback, Leibler. On Information and Sufficiency. Annals of Mathematical Statistics, 1951.

Terms in this paper