Mathematics1933foundational13 min read

Foundations of the Theory of Probability

أُسُس نظرية الاحتمالات

Kolmogorov, A. N. — Springer (Grundbegriffe der Wahrscheinlichkeitsrechnung)

The problem

Before 1933, was a patchwork: combinatorial counting for dice, frequency interpretations for experiments, and subjective degrees of belief — each with its own inconsistencies and paradoxes. There was no unified mathematical framework. Concepts like "" and "expected value" were used intuitively but never formally defined. Bertrand's paradox showed that even simple geometric probability questions could yield contradictory answers depending on unstated assumptions. Mathematics needed a foundation for probability as rigorous as Euclid's axioms were for geometry.

The contribution

Kolmogorov built probability theory on measure theory using three axioms: (1) probabilities are non-negative, (2) the certain event has probability 1, and (3) probabilities of mutually exclusive events add up. He introduced the probability space (Ω, F, P) — , σ-algebra, and probability measure — then derived random variables as measurable functions, expectation as the Lebesgue , via the Radon-Nikodym theorem, and proved the strong . The entire edifice stands on those three axioms plus one extension for infinite collections of events.

The impact

This 60-page monograph became the universal foundation of probability. Every branch — statistics, stochastic processes, information theory, , quantum mechanics, financial mathematics — rests on Kolmogorov's axioms. It resolved decades of foundational debates, answered Hilbert's sixth problem for probability, and gave mathematicians a shared language. Without it, concepts central to AI — , loss functions, probabilistic graphical models — would lack rigorous grounding.

Before Kolmogorov, probability was like a building with no blueprints: each room — gambling, insurance, physics — was built by different craftsmen using different rules. The rooms didn't always connect, and sometimes the floors disagreed about which way was up.

Kolmogorov drew the blueprint: three rules that every room must follow. The rooms can still look different inside — a casino and a physics lab use probability differently — but they all share the same foundation, the same walls, the same guarantee that the building won't collapse.

Those three rules are the axioms of probability. Everything else — Bayes' theorem, the law of large numbers, random variables — is built from them, brick by logical brick.

The crisis: probability without a foundation

By the early 20th century, probability had become indispensable — actuaries priced insurance, physicists modeled Brownian motion, geneticists predicted trait inheritance — yet the subject lacked any formal definition everyone agreed on. Three rival camps each claimed ownership:

  • Classical (Laplace): probability = favorable outcomes / total outcomes. Works for dice, but breaks immediately for continuous problems. How do you count "outcomes" when a dart can land at any real-valued point?

  • Frequentist (von Mises): probability = limiting frequency in repeated experiments. But not all events can be repeated, and defining "random sequence" precisely turned out to be circular.

  • Subjective (de Finetti): probability = rational degree of belief. Powerful for decision theory, but without formal constraints, different people assign different numbers to the same event, and there's no way to say who's right.

Bertrand's paradox (1889) crystallized the problem: "What is the probability that a random chord of a circle is longer than the side of an inscribed equilateral triangle?" Depending on how you define "random chord," the answer is 1/2, 1/3, or 1/4 — all with seemingly valid reasoning. Without agreed-upon rules, probability questions could have no definite answer.

Open in Lab
Bertrand's paradox: three valid methods for choosing a "random chord" give three different probabilities. This is why axioms were needed.
The demo wakes as you arrive…

The solution: the probability space (Ω, F, P)

Kolmogorov's insight was radical in its simplicity: don't define what probability is — define what it does. He borrowed the machinery of measure theory, developed by Lebesgue and Borel, and recast probability as a special kind of measurement.

The construction has three parts, often called the probabilist's trinity:

  • Ω (sample space) — the set of all possible outcomes of an experiment. For a coin flip, Ω = {H, T}. For a dart throw, Ω is the entire dartboard. Think of Ω as the complete catalog of everything that could happen.

  • F (σ-algebra) — the collection of events we're allowed to ask about. Not every subset of Ω needs to be measurable, especially in continuous spaces. F tells us which questions are well-posed. Think of it as the list of valid questions you can put to the experiment.

  • P (probability measure) — a function that assigns a number between 0 and 1 to each event in F, following the three axioms. Think of P as the oracle that answers each valid question with a number.

Open in Lab
Build a probability space step by step: choose a sample space, see which events form a valid σ-algebra, and assign probabilities that satisfy the axioms.
The demo wakes as you arrive…

The three axioms: the entire foundation

Everything in probability theory flows from three axioms. Before presenting them formally, let's understand what each one demands and why it's needed.

Axiom 1 — Non-negativity: probabilities cannot be negative. You cannot have a −30% chance of rain. This anchors the scale: the minimum possible probability is zero.

Axiom 2 — Normalization: the probability that something happens is exactly 1. If you flip a coin, it will land on heads or tails (or the edge, if your sample space includes that). The certain event always gets probability 1.

Axiom 3 — Countable additivity: if events cannot happen simultaneously (they are mutually exclusive), then the probability that any one of them happens equals the sum of their individual probabilities. If there's a 30% chance of rain and a 20% chance of snow, and it can't rain and snow at the same time, there's a 50% chance of precipitation. This rule extends to countably infinite collections of disjoint events.

Axiom 1: P(A)≥0∀ A∈F\text{Axiom 1: } P(A) \geq 0 \quad \forall\, A \in \mathcal{F}
Axiom 1 — Non-negativity — Every event has a probability of at least zero.
Axiom 2: P(Ω)=1\text{Axiom 2: } P(\Omega) = 1
Axiom 2 — Normalization — The probability that some outcome occurs is exactly 1.
Axiom 3: P ⁣(⋃i=1∞Ai)=∑i=1∞P(Ai)when Ai∩Aj=∅  ∀ i≠j\text{Axiom 3: } P\!\left(\bigcup_{i=1}^{\infty} A_i\right) = \sum_{i=1}^{\infty} P(A_i) \quad \text{when } A_i \cap A_j = \emptyset \;\forall\, i \neq j
Axiom 3 — Countable additivity (σ-additivity) — For any countable collection of mutually exclusive events, the probability of their union equals the sum of their individual probabilities.
Open in Lab
Assign probabilities to events and watch the axioms check themselves in real-time. Violate one to see what breaks.
The demo wakes as you arrive…

What the axioms give us for free

From just three axioms, a surprising wealth of results follows logically. No additional assumptions are needed — these are theorems, not postulates:

  • P(∅) = 0 — the impossible event has zero probability.
  • P(Aᶜ) = 1 − P(A) — the probability of "not A" is one minus the probability of A.
  • P(A ∪ B) = P(A) + P(B) − P(A ∩ B) — the inclusion-exclusion principle.
  • If A ⊆ B, then P(A) ≤ P(B) — a subset can't be more likely than its superset.
  • 0 ≤ P(A) ≤ 1 — probabilities are always between 0 and 1.

These feel "obvious," but that's exactly the point. The axioms formalize our intuition so precisely that the obvious becomes provable and the non-obvious becomes discoverable.

P(A∪B)=P(A)+P(B)−P(A∩B)P(A \cup B) = P(A) + P(B) - P(A \cap B)
Inclusion-exclusion — derived from the axioms — To count the probability of A or B, add their individual probabilities then subtract the overlap to avoid double-counting. This is a theorem, not an axiom.

Random variables: from outcomes to numbers

A sample space can be anything — coin sides, weather states, DNA sequences. But mathematics works with numbers. A random variable is a function X:Ω→RX: \Omega \to \mathbb{R} that translates each outcome into a number. The key requirement: for any value xx, the set {ω∈Ω:X(ω)≤x}\lbrace\omega \in \Omega : X(\omega) \leq x\rbrace must be in the σ-algebra F. In other words, the question "is X less than or equal to x?" must be a valid question we can assign probability to. This requirement is called measurability.

Think of a random variable as a numerical reading from an experiment. A thermometer is a random variable: the outcome is the state of the atmosphere (an element of Ω), and the reading is a number on the scale. Different thermometers (different random variables) can extract different numbers from the same experiment.

FX(x)=P({ω∈Ω:X(ω)≤x})=P(X≤x)F_X(x) = P(\lbrace\omega \in \Omega : X(\omega) \leq x\rbrace) = P(X \leq x)
Cumulative distribution function (CDF) — The CDF summarizes everything about a random variable's probability behavior: for each threshold x, what is the probability of falling at or below it?
Open in Lab
Map outcomes in Ω to numerical values and see the resulting distribution function build itself.
The demo wakes as you arrive…

Expectation: the Lebesgue integral in disguise

The expected value of a random variable is its "long-run average" — the value you'd converge to if you repeated the experiment infinitely many times and averaged the results. Kolmogorov defined it formally as the Lebesgue integral of X with respect to the probability measure P.

Why Lebesgue and not Riemann? Because the Riemann integral slices the x-axis into equal intervals, which doesn't work for arbitrary probability distributions. The Lebesgue integral slices the y-axis instead: it groups together all outcomes that give the same value, then weights by their total probability. This works for any measurable function on any probability space — discrete, continuous, or anything in between.

E[X]=∫ΩX(ω) dP(ω)E[X] = \int_{\Omega} X(\omega)\, dP(\omega)
Expected value as Lebesgue integral — Integrate the random variable X over the entire sample space, weighted by the probability measure. This single formula covers both discrete sums and continuous integrals.

Conditional probability: updating belief with evidence

Conditional probability answers: "Given that B happened, what is the probability of A?" Kolmogorov defined it as a derived concept, not a primitive one. It falls out of the axioms naturally:

Think of it as zooming in on the sample space. Once you know B occurred, the universe of possibilities shrinks from Ω to B. The conditional probability rescales everything so that B becomes the new "certain event" — its probability becomes 1, and everything inside B is proportionally adjusted.

This definition works beautifully when P(B) > 0. For the case P(B) = 0 — which arises constantly in continuous probability, e.g., "given that the dart landed at exactly this point" — Kolmogorov introduced via the Radon-Nikodym theorem, one of the book's deepest contributions.

P(A∣B)=P(A∩B)P(B),P(B)>0P(A \mid B) = \frac{P(A \cap B)}{P(B)}, \quad P(B) > 0
Conditional probability — derived, not axiom — The probability of A given B equals the probability of both happening, divided by the probability of B. This is a definition derived from the axioms, not a fourth axiom.

Bayes' theorem follows immediately from the definition of conditional probability by simple algebra. It reverses the direction of conditioning: if you know P(B|A) and want P(A|B), Bayes tells you how:

P(A∣B)=P(B∣A) P(A)P(B)P(A \mid B) = \frac{P(B \mid A)\, P(A)}{P(B)}
Bayes' theorem — a direct consequence of the axioms — Prior belief P(A) is updated by evidence P(B|A) to produce posterior belief P(A|B). This single equation powers Bayesian inference, spam filters, medical diagnosis, and much of modern AI.
Open in Lab
Watch the sample space "zoom in" when you condition on event B. The remaining probabilities rescale to sum to 1.
The demo wakes as you arrive…

Independence: the multiplicative rule

Two events A and B are independent if knowing that one occurred tells you nothing about the other. Kolmogorov formalized this as a definition derived from his axioms: A and B are independent if and only if P(A∩B)=P(A)⋅P(B)P(A \cap B) = P(A) \cdot P(B).

Notice this is not an axiom — it's a definition. But it's an extraordinarily powerful one. Independence lets you multiply probabilities: the chance of flipping 10 heads in a row is (1/2)10(1/2)^{10}, because each flip is independent of the others. Without this concept, we could not model anything from coin sequences to .

Kolmogorov extended independence to arbitrary collections: events A1,A2,…A_1, A_2, \ldots are mutually independent if the probability of any subcollection's intersection equals the product of their individual probabilities. Pairwise independence is necessary but not sufficient — a subtle point that catches even experienced practitioners.

A⊥ ⁣ ⁣ ⁣⊥B  ⟺  P(A∩B)=P(A)⋅P(B)A \perp\!\!\!\perp B \iff P(A \cap B) = P(A) \cdot P(B)
Independence — a definition, not an axiom — Two events are independent precisely when their joint probability factors into the product of their marginals. This is the mathematical formalization of "no influence."

The law of large numbers: probability meets the real world

The law of large numbers is the bridge between abstract probability and physical observation. It says: if you repeat an experiment independently many times, the average of the results will converge to the expected value.

Kolmogorov proved the strong law: the average converges almost surely (with probability 1), not just "in probability" as the weak law states. This means for almost every possible infinite sequence of experiments, the running average will eventually settle at E[X] and stay there — not just get close on average.

This theorem is why statistics works. It's why polling can predict elections, why insurance companies can set premiums, and why a neural network on more data gives better results. The averages converge because the axioms guarantee it.

P ⁣(lim⁡n→∞1n∑i=1nXi=E[X])=1P\!\left(\lim_{n \to \infty} \frac{1}{n}\sum_{i=1}^{n} X_i = E[X]\right) = 1
Strong law of large numbers — For independent, identically distributed random variables with finite expectation, the sample average converges almost surely to the population mean.
Open in Lab
Run coin flips or dice rolls and watch the running average converge to the theoretical expectation. More trials = closer to the truth.
The demo wakes as you arrive…

The axioms in code

A probability space from scratch: axioms verifiedpython

Simplified to show the idea — not the real implementation.

import numpy as np

# === Build a probability space (Ω, F, P) ===

# Ω: sample space — outcomes of rolling a fair die
omega = set([1, 2, 3, 4, 5, 6])

# P: probability measure — each face equally likely
P = dict((outcome, 1/6) for outcome in omega)

def prob(event):
    """Compute P(event) by summing over outcomes."""
    return sum(P[o] for o in event if o in P)

# --- Verify the three axioms ---

# Axiom 1: Non-negativity
assert all(p >= 0 for p in P.values()), "Axiom 1 violated!"

# Axiom 2: Normalization
assert abs(prob(omega) - 1.0) < 1e-10, "Axiom 2 violated!"

# Axiom 3: Additivity (for disjoint events)
A = set([1, 2])        # rolling 1 or 2
B = set([3, 4, 5])     # rolling 3, 4, or 5
assert A & B == set(), "A and B must be disjoint"
assert abs(prob(A | B) - (prob(A) + prob(B))) < 1e-10

# --- Derived results (theorems, not axioms) ---

# P(∅) = 0
assert prob(set()) == 0

# P(Aᶜ) = 1 - P(A)
A_complement = omega - A
assert abs(prob(A_complement) - (1 - prob(A))) < 1e-10

# Conditional probability: P(A|B) = P(A∩B) / P(B)
C = set([2, 4, 6])     # even numbers
D = set([1, 2, 3])     # numbers ≤ 3
p_C_given_D = prob(C & D) / prob(D)
print("P(even | ≤3) = %.4f" % p_C_given_D)  # 0.3333

# Independence check: P(A∩B) =? P(A)·P(B)
independent = abs(prob(C & D) - prob(C) * prob(D)) < 1e-10
print("Even and ≤3 independent?", independent)  # True!

Why it matters — from 1933 to AI

  1. 1654

    Pascal & Fermat correspondence

    The birth of mathematical probability — two mathematicians exchange letters about gambling problems, laying the combinatorial foundations.

  2. 1812

    Laplace's "Théorie analytique des probabilités"

    The classical definition: probability as the ratio of favorable to total outcomes. Powerful but limited to equiprobable finite spaces.

  3. 1933

    Kolmogorov's Grundbegriffe

    The axiomatic foundations. Probability becomes measure theory, all paradoxes are resolved, and a universal mathematical language is born.

  4. 1946

    Monte Carlo methods

    Ulam and von Neumann use random sampling — grounded in Kolmogorov's framework — to simulate nuclear physics. Probability becomes a computational tool.

  5. 1948

    Shannon's information theory

    Entropy, mutual information, channel capacity — all built on Kolmogorov's probability spaces and his definition of expectation.

  6. 1970

    Kolmogorov complexity

    Kolmogorov himself defines algorithmic complexity — the shortest program that produces a string. A bridge between probability and computation.

  7. 2012

    Deep learning revolution

    Every loss function, every gradient descent step, every batch normalization layer is a theorem from Kolmogorov's framework applied at industrial scale.

Kolmogorov gave probability its mathematical passport. Before 1933, it was a guest in the house of mathematics; after 1933, it had its own room with the same structural guarantees as algebra, geometry, and analysis. Every probabilistic statement you encounter — in a textbook, a paper, or a machine learning library — rests on sixty pages written in German, ninety-one years ago.

CitationKolmogorov, A. N.. Grundbegriffe der Wahrscheinlichkeitsrechnung (Foundations of the Theory of Probability). Springer (Ergebnisse der Mathematik), 1933.

Terms in this paper