Mathematics1922foundational11 min read

On the Mathematical Foundations of Theoretical Statistics

في الأسس الرياضية للإحصاء النظري

Fisher, R. A. — Philosophical Transactions of the Royal Society A

The problem

By the early 1900s, statistics had no unified theory of estimation. Researchers used Karl Pearson's method of moments — matching sample moments to population moments — but nobody had proved it was optimal, or even defined what "optimal" means. The same word "mean" was used for both the population and the sample average, blurring the line between the unknown truth and the data-driven guess. There were no formal criteria to evaluate or compare estimation methods.

The contribution

Fisher built the conceptual vocabulary that all of statistics still uses. He distinguished parameters (fixed but unknown population quantities) from statistics (functions of the sample data). He defined three criteria for judging an — , , and — and proposed as the method that satisfies all three. He introduced the function (not a , but a measure of how well parameters explain the data) and showed that sufficient statistics capture all the information a sample contains about a parameter. He formalized — the curvature of the — as the measure of how much a sample tells you.

The impact

This single paper created the language of modern statistics. Every machine-learning loss function is a likelihood in disguise. Every run is maximum-likelihood estimation. Fisher Information appears in the , in natural gradient descent, in variational inference, and in information geometry. The concepts Fisher defined here — parameter, statistic, consistency, efficiency, sufficiency, likelihood — are so deeply embedded in science that most practitioners use them without knowing they all trace back to one 60-page paper from 1922.

Before Fisher, statisticians were like chefs who judged every soup by stirring it and tasting randomly — some took a spoonful from the top, others from the bottom, each confident their method was best, but nobody had defined what "best" actually means.

Fisher walked into the kitchen and said: first define the recipe you're trying to recover (the parameter). Then define what a taste-test is (the statistic). Now I'll tell you exactly which spoonful squeezes out every last drop of information about the recipe (sufficiency), wastes the least effort doing it (efficiency), and guarantees that with a big enough pot your guess converges to the truth (consistency).

The method he proposed — Maximum Likelihood — is the spoon that satisfies all three criteria at once.

The problem: statistics without a compass

By the early 1900s, Karl Pearson's method of moments was the dominant way to fit a to data: match the sample mean to the population mean, the sample to the population variance, and so on. It was practical and intuitive, but Fisher identified two deep flaws:

  • No optimality guarantee. Matching moments gives consistent estimates, but nobody had shown they are the most precise estimates possible. Fisher demonstrated concrete cases where moment-based estimates throw away information — their efficiency can drop to zero.

  • No separation of concepts. The word "mean" referred both to the fixed population quantity (the parameter μ\mu) and to the number computed from data (xˉ=1n∑xi\bar{x} = \frac{1}{n}\sum x_i). Without this distinction, you cannot even state the estimation problem clearly, let alone solve it optimally.

Open in Lab
Toggle between the population (parameter) and a random sample (statistic) to see how they differ.
The demo wakes as you arrive…

Fisher's vocabulary: parameter, statistic, and the infinite population

Fisher's first move was linguistic: name the objects clearly, and the theory follows.

A parameter is a fixed, unknown quantity that fully describes the population — for example, the true mean μ\mu or the true standard deviation σ\sigma of a normal distribution. It is a property of the infinite hypothetical population that generated your data. You never observe it directly.

A statistic is any function of the observed sample — xˉ\bar{x}, the sample median, the sample range — that you compute to estimate a parameter. It varies from sample to sample because it depends on which data you happened to collect.

This separation sounds obvious today, but before Fisher it was the single biggest source of confusion in statistics. Once you have it, you can ask precise questions: which statistic is the best estimator of which parameter, and what does "best" mean?

The three criteria: consistency, efficiency, sufficiency

Fisher proposed three criteria for evaluating any estimator. Think of them as three increasingly demanding tests:

Consistency — In the limit. When you calculate the statistic from the entire population (not a sample), it equals the parameter exactly. Equivalently, as the sample grows toward infinity, the estimate converges to the true value. This is the minimum requirement: an estimator that doesn't converge to the truth is useless.

Efficiency — Among all consistent estimators, the efficient one has the smallest possible variance in large samples. It squeezes the maximum precision out of each data point. Fisher showed that the variance of any unbiased estimator cannot be smaller than a lower bound — what we now call the Cramér-Rao bound — and the estimator that achieves this bound is efficient.

Sufficiency — A statistic is sufficient for a parameter if it captures all the information the sample contains about that parameter. Once you know the sufficient statistic, the rest of the data is noise — it cannot tell you anything more about the parameter. Sufficiency is about information preservation: it compresses the data without losing anything relevant.

Open in Lab
Step through the three criteria to see what each one demands from an estimator.
The demo wakes as you arrive…
Open in Lab
Watch how a sufficient statistic compresses 50 data points into a single number — without losing any information about the mean.
The demo wakes as you arrive…

Likelihood: not probability, but its mirror

Before Fisher, the dominant tool for reasoning backward from data to parameters was inverse probability — with a uniform prior. Fisher rejected this on philosophical grounds: if you truly know nothing about a parameter, you cannot assign it a probability distribution.

Instead he introduced the likelihood function: given observed data xx, the likelihood of a parameter value θ\theta is L(θ)=P(x∣θ)L(\theta) = P(x \mid \theta). It uses the same formula as probability but reads it in the opposite direction. Probability asks: given parameters, how likely is this data? Likelihood asks: given data, how well does this parameter value explain it?

The crucial distinction is that likelihood is not a probability distribution over θ\theta — it does not integrate to 1, and it does not obey the axioms of probability. It is a relative measure: L(θ1)/L(θ2)L(\theta_1) / L(\theta_2) tells you how much better θ1\theta_1 explains the data than θ2\theta_2.

L(θ)=P(x1,x2,…,xn∣θ)=∏i=1nf(xi∣θ)L(\theta) = P(x_1, x_2, \ldots, x_n \mid \theta) = \prod_{i=1}^{n} f(x_i \mid \theta)
The likelihood function — the engine of Fisher's framework — For independent observations, the likelihood is the product of their individual densities evaluated at each data point · It measures how plausible a parameter value is given the observed data

Read the likelihood function like a detective reviewing evidence: the data (the evidence) is fixed, and you're trying each suspect (each θ\theta) to see which one best explains the crime scene. The suspect who makes the evidence most probable is the maximum likelihood estimate.

Open in Lab
Drag θ along the curve. The peak is the Maximum Likelihood Estimate.
The demo wakes as you arrive…

Maximum Likelihood Estimation: finding the peak

The Method of Maximum Likelihood says: choose the parameter value θ^\hat{\theta} that maximizes L(θ)L(\theta). In practice, it is easier to maximize the log-likelihood ℓ(θ)=ln⁡L(θ)\ell(\theta) = \ln L(\theta), because products become sums and the maximum stays in the same place.

For a normal distribution with unknown mean μ\mu and known variance σ2\sigma^2, the log-likelihood is:

ℓ(μ)=−n2ln⁡(2πσ2)−12σ2∑i=1n(xi−μ)2\ell(\mu) = -\frac{n}{2}\ln(2\pi\sigma^2) - \frac{1}{2\sigma^2}\sum_{i=1}^{n}(x_i - \mu)^2
Log-likelihood for a normal distribution — This objective measures how well a particular choice of the population mean explains the observed data under a normal-distribution assumption. Maximum likelihood estimation chooses the value that makes the observed samples appear most plausible. For a normal distribution with known variance, the optimal estimate turns out to be the sample mean. This result is especially important because the sample mean captures all the information needed about the population mean contained in the data.

Fisher claimed — and largely showed — that Maximum Likelihood Estimation produces estimators that are consistent, asymptotically efficient, and (when a sufficient statistic exists) sufficient. This is why MLE became the standard estimation method in all of science. In modern machine learning, "training" a usually means maximizing a log-likelihood (or equivalently, minimizing a cross-entropy loss).

Why not method of moments?

Fisher compared his Maximum Likelihood method to Pearson's method of moments and showed that moments can be inefficient: they sometimes use only a fraction of the available information.

Consider fitting a Pearson Type III distribution. The method of moments matches the first three sample moments to the population moments and solves for the parameters. Fisher showed that the resulting estimates can have variance dramatically larger than the MLE — in some cases, the efficiency of the moment estimator drops toward zero. The moment estimator is consistent (it converges) but wasteful (it converges slowly because it ignores information in the data).

This was Fisher's most devastating critique: consistency alone is not enough. You must also demand efficiency — and the method of moments fails this test.

Open in Lab
Compare the sampling distributions of MLE vs. Method of Moments. Notice how MLE clusters tightly around the truth.
The demo wakes as you arrive…

Fisher Information: the curvature of knowledge

How much does a sample tell you about a parameter? Fisher answered with a single number: the Fisher Information, defined as the expected value of the squared (the of the log-likelihood), or equivalently, the negative expected value of the second derivative of the log-likelihood.

Imagine the log-likelihood as a hilltop. If the peak is sharp (high curvature), even a small shift in θ\theta makes the likelihood drop dramatically — the data strongly "points" to one value. If the peak is flat (low curvature), many parameter values look almost equally plausible — the data is uninformative.

Fisher Information I(θ)I(\theta) measures that curvature. It sets the ultimate precision limit for any estimator via what we now call the Cramér-Rao bound: Var(θ^)≥1/I(θ)\text{Var}(\hat{\theta}) \geq 1 / I(\theta).

I(θ)=E[(∂∂θln⁡f(X∣θ))2]=−E[∂2∂θ2ln⁡f(X∣θ)]I(\theta) = E\left[\left(\frac{\partial}{\partial\theta} \ln f(X \mid \theta)\right)^2\right] = -E\left[\frac{\partial^2}{\partial\theta^2} \ln f(X \mid \theta)\right]
Fisher Information — the sharpness of the likelihood peak — Fisher Information measures how much information the observed data carries about an unknown parameter. When small changes in the parameter produce large changes in how well the model explains the data, the parameter can be estimated precisely. When many different parameter values explain the data almost equally well, estimation becomes more uncertain. As more independent observations are collected, the available information grows, leading to increasingly precise parameter estimates.
Open in Lab
Drag the sample size slider. Watch the log-likelihood sharpen and the Fisher Information grow — more data means a sharper peak.
The demo wakes as you arrive…

The same ideas in code

Maximum Likelihood Estimation for a Normal distributionpython

Simplified to show the idea — not the real implementation.

import numpy as np

def log_likelihood_normal(data, mu, sigma):
    """Log-likelihood of data under Normal(mu, sigma²)."""
    n = len(data)
    return -n/2 * np.log(2 * np.pi * sigma**2) \
           - np.sum((data - mu)**2) / (2 * sigma**2)

def mle_normal(data):
    """Maximum Likelihood Estimates for Normal distribution."""
    mu_hat = np.mean(data)            # sufficient statistic for mu
    sigma2_hat = np.var(data, ddof=0)  # MLE uses n, not n-1
    return mu_hat, np.sqrt(sigma2_hat)

def fisher_information_normal(n, sigma):
    """Fisher Information for the mean of a Normal(mu, sigma²)."""
    return n / sigma**2   # more data or less noise → more information

# Example: 100 observations from Normal(5, 2²)
np.random.seed(42)
data = np.random.normal(loc=5, scale=2, size=100)
mu_hat, sigma_hat = mle_normal(data)
I = fisher_information_normal(len(data), sigma_hat)

print(f"MLE mean:    {mu_hat:.3f}   (true = 5)")
print(f"MLE std:     {sigma_hat:.3f}   (true = 2)")
print(f"Fisher Info: {I:.2f}")
print(f"Cramér-Rao lower bound on Var(μ̂): {1/I:.5f}")

# That's it. Every ML training loop is doing exactly this:
# maximize log_likelihood(data, parameters).

Why it mattered

  1. 1894

    Method of Moments (Pearson)

    Karl Pearson introduces the method of moments for fitting distributions — the first systematic estimation procedure. It works, but Fisher will show it can be wasteful.

  2. 1912

    Fisher's first mention of maximum likelihood

    In his undergraduate paper, Fisher proposes maximizing the likelihood — but calls it "absolute criterion" and justifies it with inverse probability, which he later rejects.

  3. 1922

    This paper — the foundations

    Fisher defines parameter, statistic, consistency, efficiency, sufficiency, likelihood, and Maximum Likelihood Estimation. The conceptual framework for all of statistics.

  4. 1925

    Theory of Statistical Estimation

    Fisher refines the 1922 ideas, proves efficiency more rigorously, and discovers that sufficient statistics need not always exist.

  5. 1935

    Cramér-Rao Bound formalized

    Cramér and Rao independently formalize the lower bound on estimator variance using Fisher Information — the bound Fisher had intuited in 1922.

  6. 1946

    Rao-Blackwell Theorem

    If you condition any estimator on a sufficient statistic, you get an estimator that is at least as good — formalizing Fisher's insight that sufficiency preserves all information.

  7. 1998

    Fisher Information in deep learning

    Amari introduces the natural gradient — using the Fisher Information matrix to scale gradient steps — bridging 1922 statistics to modern neural network optimization.

Every time a is trained by minimizing cross-entropy loss, it is performing Maximum Likelihood Estimation. Every time a researcher reports a confidence interval, the math traces to Fisher's definitions of efficiency and information. The concepts in this paper are so fundamental that they have become invisible — like the air science breathes.

CitationFisher, R. A.. On the Mathematical Foundations of Theoretical Statistics. Philosophical Transactions of the Royal Society A, 1922.

Terms in this paper