Information Theory1951intermediate8 min read
On Information and Sufficiency
في المعلومات والكفاية الإحصائية
Kullback, S. · Leibler, R. A. — Annals of Mathematical Statistics
The problem
By 1951, Shannon had defined — the minimum number of bits to encode a message — but there was no formal measure for how much information is lost when one is used to approximate another. Statisticians needed a way to quantify: how far is my from reality? And when does a summary of data (a ) preserve all the information the full dataset contains?
The contribution
Kullback and Leibler introduced a measure now called (or relative entropy): for two distributions P and Q, it equals the expected number of extra bits needed when coding data from P using a code optimized for Q. They proved it is always non-negative (Gibbs' inequality), equals zero only when P = Q, and showed that a statistic is sufficient if and only if it preserves the full KL between any two hypotheses. This unified information theory and statistical sufficiency in a single framework.
The impact
KL divergence became one of the most used quantities in modern AI. It is the foundation of cross-entropy in , the regularizer in Variational Autoencoders (VAEs), the objective in t-SNE visualization, a core component of (RLHF), and the theoretical basis for model selection criteria like AIC. Any time a system measures how far a learned distribution is from a target, KL divergence is likely involved.
Imagine you're packing an umbrella for every day of the year based on a weather forecast. If the forecast is perfect, you pack exactly right — an umbrella on rainy days, nothing on sunny ones. But if the forecast says "10% rain" when the truth is "60% rain", you'll get soaked repeatedly.
KL divergence measures this mismatch: the extra "surprise" you experience, on average, because you prepared for the wrong weather pattern. It's not about any single day — it's about the systematic cost of believing the wrong distribution.
Zero divergence means your forecast is reality. The bigger the divergence, the more umbrellas you're missing.
The question: how wrong is my model?
In 1948, Claude Shannon showed that a source producing symbols according to distribution P needs at least bits per symbol on average — this is entropy. But what if you don't know P and instead design your code for a different distribution Q? You'll need bits — cross-entropy — which is always at least as large as .
The gap is the wasted bits: the cost of using the wrong model. Kullback and Leibler formalized this gap as a single quantity and gave it rigorous properties, connecting Shannon's information theory to the classical statistics of hypothesis testing and sufficient statistics.
The formula: measuring information loss
Before seeing the math, here's the idea in words: for each possible outcome, calculate how much more surprised you are under Q than under P. Weight that extra surprise by how often the outcome actually happens (under P). Sum it all up. That sum is KL divergence — the average wasted surprise.
For continuous distributions the sum becomes an : , where and are the probability density functions. The interpretation is identical: average extra information per sample.
Properties: what KL divergence is (and isn't)
KL divergence has four properties that matter in practice:
-
Non-negativity. always. This is Gibbs' inequality — you can never do better by using the wrong distribution.
-
Zero iff identical. if and only if almost everywhere. The mismatch cost vanishes only when the distributions match perfectly.
-
Asymmetry. in general. KL divergence is not a true distance. The cost of approximating P with Q is different from approximating Q with P.
-
Additivity for independent variables. If P and Q factor over independent components, the total KL divergence is the sum of the individual divergences.
Forward vs reverse KL: mode-covering vs mode-seeking
This asymmetry has a vivid intuition. Suppose P is a mixture of two Gaussians (two peaks) and Q is a single Gaussian.
Forward KL : penalizes Q heavily wherever P has mass but Q doesn't. Q stretches wide to cover both peaks — even if that means putting probability in the valley between them. This is mode-covering behavior.
Reverse KL : penalizes Q wherever Q has mass but P doesn't. Q collapses onto a single peak to avoid placing probability where P is zero. This is mode-seeking behavior.
Understanding this distinction is essential for choosing the right objective in generative models, variational inference, and optimization.
Sufficiency: when a summary preserves all information
The second half of Kullback and Leibler's paper connects KL divergence to statistical sufficiency. A sufficient statistic is a summary of data that loses nothing relevant to distinguishing between hypotheses.
Think of it like compressing a photo. A lossy compression throws away detail. A sufficient compression throws away only — every bit of signal survives. In formal terms: a statistic is sufficient for distinguishing from if and only if . Compressing the data through doesn't reduce the divergence — no information is lost.
KL divergence in modern machine learning
KL divergence appears — sometimes explicitly, sometimes disguised — in nearly every corner of modern machine learning:
-
Cross-entropy loss. Minimizing cross-entropy between labels and model predictions is minimizing KL divergence. Every classifier you've trained uses this.
-
Variational Autoencoders. The loss has two terms: a reconstruction loss and , which pushes the learned posterior toward the prior. This is what makes the smooth and generative.
-
t-SNE. The visualization algorithm minimizes KL divergence between high-dimensional and low-dimensional neighborhood probabilities — this asymmetric cost is why t-SNE preserves local but not global structure.
-
RLHF. A KL penalty term prevents the fine-tuned policy from drifting too far from the base model, balancing helpfulness with stability.
The same idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
def kl_divergence(p, q):
"""Compute D_KL(P || Q) for discrete distributions P and Q.
Both must be valid probability arrays that sum to 1."""
p = np.asarray(p, dtype=float)
q = np.asarray(q, dtype=float)
# Only sum where p > 0 (0 * log(0) = 0 by convention)
mask = p > 0
return np.sum(p[mask] * np.log(p[mask] / q[mask]))
def cross_entropy(p, q):
"""H(P, Q) = H(P) + D_KL(P || Q)."""
p = np.asarray(p, dtype=float)
q = np.asarray(q, dtype=float)
mask = p > 0
return -np.sum(p[mask] * np.log(q[mask]))
# Example: fair coin P vs biased coin Q
P = np.array([0.5, 0.5]) # true distribution
Q = np.array([0.9, 0.1]) # our wrong model
print(f"H(P) = {-np.sum(P * np.log(P)):.4f} nats")
print(f"H(P, Q) = {cross_entropy(P, Q):.4f} nats")
print(f"D_KL(P||Q) = {kl_divergence(P, Q):.4f} nats") # the gap
print(f"D_KL(Q||P) = {kl_divergence(Q, P):.4f} nats") # different!Why it changed everything
1948
Shannon's entropy
Claude Shannon defines information entropy — the minimum bits to encode a source. The mathematical foundation on which KL divergence is built.
1951
Kullback-Leibler divergence
Kullback and Leibler formalize relative entropy and prove the sufficiency theorem. A 7-page paper that seeded decades of research.
1974
Akaike Information Criterion (AIC)
Akaike uses KL divergence as the theoretical basis for model selection — choosing between candidate models by estimating their divergence from truth.
2006
Variational Autoencoders (conceptual roots)
Variational inference methods using KL divergence mature. The KL term becomes a standard regularizer for latent variable models.
2008
t-SNE
van der Maaten and Hinton use KL divergence as the cost function for neighborhood-preserving visualization. The most cited visualization paper in ML.
2014
VAE published
Kingma and Welling introduce the VAE, with KL divergence as the explicit regularizer that keeps the latent space structured and generative.
2017
PPO and KL-penalized RL
Proximal Policy Optimization uses a KL constraint to keep policy updates stable. Later, RLHF for language models adopts the same principle.
2022
RLHF for ChatGPT
KL divergence penalty became central to aligning large language models — a direct descendant of a 1951 statistics paper now shapes AI safety research.
From a 7-page paper in the Annals of Mathematical Statistics to the loss function inside every major language model — KL divergence is the thread connecting Shannon's information theory, Fisher's sufficiency, and the deep learning revolution. Every time a model learns, it is — in some form — closing the KL gap.
CitationKullback, Leibler. On Information and Sufficiency. Annals of Mathematical Statistics, 1951.
Terms in this paper
- KL Divergenceتباعد KL
- Entropyالعشوائية الدلالية
- Cross Entropyالعشوائية المتقاطعة
- Distributionالتوزيع الإحصائي
- Divergenceالتباعد
- Likelihoodالأرجحية
- Mutual Informationالمعلومات المتبادلة
- Probabilityالاحتمالية
- Information Gainالكسب المعلوماتي
- Sufficient Statisticالإحصاء الكافي