Mathematics1922foundational11 min read
On the Mathematical Foundations of Theoretical Statistics
في الأسس الرياضية للإحصاء النظري
Fisher, R. A. — Philosophical Transactions of the Royal Society A
The problem
By the early 1900s, statistics had no unified theory of estimation. Researchers used Karl Pearson's method of moments — matching sample moments to population moments — but nobody had proved it was optimal, or even defined what "optimal" means. The same word "mean" was used for both the population and the sample average, blurring the line between the unknown truth and the data-driven guess. There were no formal criteria to evaluate or compare estimation methods.
The contribution
Fisher built the conceptual vocabulary that all of statistics still uses. He distinguished parameters (fixed but unknown population quantities) from statistics (functions of the sample data). He defined three criteria for judging an — , , and — and proposed as the method that satisfies all three. He introduced the function (not a , but a measure of how well parameters explain the data) and showed that sufficient statistics capture all the information a sample contains about a parameter. He formalized — the curvature of the — as the measure of how much a sample tells you.
The impact
This single paper created the language of modern statistics. Every machine-learning loss function is a likelihood in disguise. Every run is maximum-likelihood estimation. Fisher Information appears in the , in natural gradient descent, in variational inference, and in information geometry. The concepts Fisher defined here — parameter, statistic, consistency, efficiency, sufficiency, likelihood — are so deeply embedded in science that most practitioners use them without knowing they all trace back to one 60-page paper from 1922.
Before Fisher, statisticians were like chefs who judged every soup by stirring it and tasting randomly — some took a spoonful from the top, others from the bottom, each confident their method was best, but nobody had defined what "best" actually means.
Fisher walked into the kitchen and said: first define the recipe you're trying to recover (the parameter). Then define what a taste-test is (the statistic). Now I'll tell you exactly which spoonful squeezes out every last drop of information about the recipe (sufficiency), wastes the least effort doing it (efficiency), and guarantees that with a big enough pot your guess converges to the truth (consistency).
The method he proposed — Maximum Likelihood — is the spoon that satisfies all three criteria at once.
The problem: statistics without a compass
By the early 1900s, Karl Pearson's method of moments was the dominant way to fit a to data: match the sample mean to the population mean, the sample to the population variance, and so on. It was practical and intuitive, but Fisher identified two deep flaws:
-
No optimality guarantee. Matching moments gives consistent estimates, but nobody had shown they are the most precise estimates possible. Fisher demonstrated concrete cases where moment-based estimates throw away information — their efficiency can drop to zero.
-
No separation of concepts. The word "mean" referred both to the fixed population quantity (the parameter ) and to the number computed from data (). Without this distinction, you cannot even state the estimation problem clearly, let alone solve it optimally.
Fisher's vocabulary: parameter, statistic, and the infinite population
Fisher's first move was linguistic: name the objects clearly, and the theory follows.
A parameter is a fixed, unknown quantity that fully describes the population — for example, the true mean or the true standard deviation of a normal distribution. It is a property of the infinite hypothetical population that generated your data. You never observe it directly.
A statistic is any function of the observed sample — , the sample median, the sample range — that you compute to estimate a parameter. It varies from sample to sample because it depends on which data you happened to collect.
This separation sounds obvious today, but before Fisher it was the single biggest source of confusion in statistics. Once you have it, you can ask precise questions: which statistic is the best estimator of which parameter, and what does "best" mean?
The three criteria: consistency, efficiency, sufficiency
Fisher proposed three criteria for evaluating any estimator. Think of them as three increasingly demanding tests:
Consistency — In the limit. When you calculate the statistic from the entire population (not a sample), it equals the parameter exactly. Equivalently, as the sample grows toward infinity, the estimate converges to the true value. This is the minimum requirement: an estimator that doesn't converge to the truth is useless.
Efficiency — Among all consistent estimators, the efficient one has the smallest possible variance in large samples. It squeezes the maximum precision out of each data point. Fisher showed that the variance of any unbiased estimator cannot be smaller than a lower bound — what we now call the Cramér-Rao bound — and the estimator that achieves this bound is efficient.
Sufficiency — A statistic is sufficient for a parameter if it captures all the information the sample contains about that parameter. Once you know the sufficient statistic, the rest of the data is noise — it cannot tell you anything more about the parameter. Sufficiency is about information preservation: it compresses the data without losing anything relevant.
Likelihood: not probability, but its mirror
Before Fisher, the dominant tool for reasoning backward from data to parameters was inverse probability — with a uniform prior. Fisher rejected this on philosophical grounds: if you truly know nothing about a parameter, you cannot assign it a probability distribution.
Instead he introduced the likelihood function: given observed data , the likelihood of a parameter value is . It uses the same formula as probability but reads it in the opposite direction. Probability asks: given parameters, how likely is this data? Likelihood asks: given data, how well does this parameter value explain it?
The crucial distinction is that likelihood is not a probability distribution over — it does not integrate to 1, and it does not obey the axioms of probability. It is a relative measure: tells you how much better explains the data than .
Read the likelihood function like a detective reviewing evidence: the data (the evidence) is fixed, and you're trying each suspect (each ) to see which one best explains the crime scene. The suspect who makes the evidence most probable is the maximum likelihood estimate.
Maximum Likelihood Estimation: finding the peak
The Method of Maximum Likelihood says: choose the parameter value that maximizes . In practice, it is easier to maximize the log-likelihood , because products become sums and the maximum stays in the same place.
For a normal distribution with unknown mean and known variance , the log-likelihood is:
Fisher claimed — and largely showed — that Maximum Likelihood Estimation produces estimators that are consistent, asymptotically efficient, and (when a sufficient statistic exists) sufficient. This is why MLE became the standard estimation method in all of science. In modern machine learning, "training" a usually means maximizing a log-likelihood (or equivalently, minimizing a cross-entropy loss).
Why not method of moments?
Fisher compared his Maximum Likelihood method to Pearson's method of moments and showed that moments can be inefficient: they sometimes use only a fraction of the available information.
Consider fitting a Pearson Type III distribution. The method of moments matches the first three sample moments to the population moments and solves for the parameters. Fisher showed that the resulting estimates can have variance dramatically larger than the MLE — in some cases, the efficiency of the moment estimator drops toward zero. The moment estimator is consistent (it converges) but wasteful (it converges slowly because it ignores information in the data).
This was Fisher's most devastating critique: consistency alone is not enough. You must also demand efficiency — and the method of moments fails this test.
Fisher Information: the curvature of knowledge
How much does a sample tell you about a parameter? Fisher answered with a single number: the Fisher Information, defined as the expected value of the squared (the of the log-likelihood), or equivalently, the negative expected value of the second derivative of the log-likelihood.
Imagine the log-likelihood as a hilltop. If the peak is sharp (high curvature), even a small shift in makes the likelihood drop dramatically — the data strongly "points" to one value. If the peak is flat (low curvature), many parameter values look almost equally plausible — the data is uninformative.
Fisher Information measures that curvature. It sets the ultimate precision limit for any estimator via what we now call the Cramér-Rao bound: .
The same ideas in code
Simplified to show the idea — not the real implementation.
import numpy as np
def log_likelihood_normal(data, mu, sigma):
"""Log-likelihood of data under Normal(mu, sigma²)."""
n = len(data)
return -n/2 * np.log(2 * np.pi * sigma**2) \
- np.sum((data - mu)**2) / (2 * sigma**2)
def mle_normal(data):
"""Maximum Likelihood Estimates for Normal distribution."""
mu_hat = np.mean(data) # sufficient statistic for mu
sigma2_hat = np.var(data, ddof=0) # MLE uses n, not n-1
return mu_hat, np.sqrt(sigma2_hat)
def fisher_information_normal(n, sigma):
"""Fisher Information for the mean of a Normal(mu, sigma²)."""
return n / sigma**2 # more data or less noise → more information
# Example: 100 observations from Normal(5, 2²)
np.random.seed(42)
data = np.random.normal(loc=5, scale=2, size=100)
mu_hat, sigma_hat = mle_normal(data)
I = fisher_information_normal(len(data), sigma_hat)
print(f"MLE mean: {mu_hat:.3f} (true = 5)")
print(f"MLE std: {sigma_hat:.3f} (true = 2)")
print(f"Fisher Info: {I:.2f}")
print(f"Cramér-Rao lower bound on Var(μ̂): {1/I:.5f}")
# That's it. Every ML training loop is doing exactly this:
# maximize log_likelihood(data, parameters).Why it mattered
1894
Method of Moments (Pearson)
Karl Pearson introduces the method of moments for fitting distributions — the first systematic estimation procedure. It works, but Fisher will show it can be wasteful.
1912
Fisher's first mention of maximum likelihood
In his undergraduate paper, Fisher proposes maximizing the likelihood — but calls it "absolute criterion" and justifies it with inverse probability, which he later rejects.
1922
This paper — the foundations
Fisher defines parameter, statistic, consistency, efficiency, sufficiency, likelihood, and Maximum Likelihood Estimation. The conceptual framework for all of statistics.
1925
Theory of Statistical Estimation
Fisher refines the 1922 ideas, proves efficiency more rigorously, and discovers that sufficient statistics need not always exist.
1935
Cramér-Rao Bound formalized
Cramér and Rao independently formalize the lower bound on estimator variance using Fisher Information — the bound Fisher had intuited in 1922.
1946
Rao-Blackwell Theorem
If you condition any estimator on a sufficient statistic, you get an estimator that is at least as good — formalizing Fisher's insight that sufficiency preserves all information.
1998
Fisher Information in deep learning
Amari introduces the natural gradient — using the Fisher Information matrix to scale gradient steps — bridging 1922 statistics to modern neural network optimization.
Every time a is trained by minimizing cross-entropy loss, it is performing Maximum Likelihood Estimation. Every time a researcher reports a confidence interval, the math traces to Fisher's definitions of efficiency and information. The concepts in this paper are so fundamental that they have become invisible — like the air science breathes.
CitationFisher, R. A.. On the Mathematical Foundations of Theoretical Statistics. Philosophical Transactions of the Royal Society A, 1922.
Terms in this paper
- Maximum Likelihood Estimationتقدير الأرجحية القصوى
- Fisher Informationمعلومات فيشر
- Likelihoodالأرجحية
- Sufficiencyالكفاية
- Efficiencyالكفاءة
- Consistencyالاتّساق
- Cramér-Rao Boundحدّ كرامير-راو
- Score Functionدالة الرصيد