Mathematics1933foundational13 min read
Foundations of the Theory of Probability
أُسُس نظرية الاحتمالات
Kolmogorov, A. N. — Springer (Grundbegriffe der Wahrscheinlichkeitsrechnung)
The problem
Before 1933, was a patchwork: combinatorial counting for dice, frequency interpretations for experiments, and subjective degrees of belief — each with its own inconsistencies and paradoxes. There was no unified mathematical framework. Concepts like "" and "expected value" were used intuitively but never formally defined. Bertrand's paradox showed that even simple geometric probability questions could yield contradictory answers depending on unstated assumptions. Mathematics needed a foundation for probability as rigorous as Euclid's axioms were for geometry.
The contribution
Kolmogorov built probability theory on measure theory using three axioms: (1) probabilities are non-negative, (2) the certain event has probability 1, and (3) probabilities of mutually exclusive events add up. He introduced the probability space (Ω, F, P) — , σ-algebra, and probability measure — then derived random variables as measurable functions, expectation as the Lebesgue , via the Radon-Nikodym theorem, and proved the strong . The entire edifice stands on those three axioms plus one extension for infinite collections of events.
The impact
This 60-page monograph became the universal foundation of probability. Every branch — statistics, stochastic processes, information theory, , quantum mechanics, financial mathematics — rests on Kolmogorov's axioms. It resolved decades of foundational debates, answered Hilbert's sixth problem for probability, and gave mathematicians a shared language. Without it, concepts central to AI — , loss functions, probabilistic graphical models — would lack rigorous grounding.
Before Kolmogorov, probability was like a building with no blueprints: each room — gambling, insurance, physics — was built by different craftsmen using different rules. The rooms didn't always connect, and sometimes the floors disagreed about which way was up.
Kolmogorov drew the blueprint: three rules that every room must follow. The rooms can still look different inside — a casino and a physics lab use probability differently — but they all share the same foundation, the same walls, the same guarantee that the building won't collapse.
Those three rules are the axioms of probability. Everything else — Bayes' theorem, the law of large numbers, random variables — is built from them, brick by logical brick.
The crisis: probability without a foundation
By the early 20th century, probability had become indispensable — actuaries priced insurance, physicists modeled Brownian motion, geneticists predicted trait inheritance — yet the subject lacked any formal definition everyone agreed on. Three rival camps each claimed ownership:
-
Classical (Laplace): probability = favorable outcomes / total outcomes. Works for dice, but breaks immediately for continuous problems. How do you count "outcomes" when a dart can land at any real-valued point?
-
Frequentist (von Mises): probability = limiting frequency in repeated experiments. But not all events can be repeated, and defining "random sequence" precisely turned out to be circular.
-
Subjective (de Finetti): probability = rational degree of belief. Powerful for decision theory, but without formal constraints, different people assign different numbers to the same event, and there's no way to say who's right.
Bertrand's paradox (1889) crystallized the problem: "What is the probability that a random chord of a circle is longer than the side of an inscribed equilateral triangle?" Depending on how you define "random chord," the answer is 1/2, 1/3, or 1/4 — all with seemingly valid reasoning. Without agreed-upon rules, probability questions could have no definite answer.
The solution: the probability space (Ω, F, P)
Kolmogorov's insight was radical in its simplicity: don't define what probability is — define what it does. He borrowed the machinery of measure theory, developed by Lebesgue and Borel, and recast probability as a special kind of measurement.
The construction has three parts, often called the probabilist's trinity:
-
Ω (sample space) — the set of all possible outcomes of an experiment. For a coin flip, Ω = {H, T}. For a dart throw, Ω is the entire dartboard. Think of Ω as the complete catalog of everything that could happen.
-
F (σ-algebra) — the collection of events we're allowed to ask about. Not every subset of Ω needs to be measurable, especially in continuous spaces. F tells us which questions are well-posed. Think of it as the list of valid questions you can put to the experiment.
-
P (probability measure) — a function that assigns a number between 0 and 1 to each event in F, following the three axioms. Think of P as the oracle that answers each valid question with a number.
The three axioms: the entire foundation
Everything in probability theory flows from three axioms. Before presenting them formally, let's understand what each one demands and why it's needed.
Axiom 1 — Non-negativity: probabilities cannot be negative. You cannot have a −30% chance of rain. This anchors the scale: the minimum possible probability is zero.
Axiom 2 — Normalization: the probability that something happens is exactly 1. If you flip a coin, it will land on heads or tails (or the edge, if your sample space includes that). The certain event always gets probability 1.
Axiom 3 — Countable additivity: if events cannot happen simultaneously (they are mutually exclusive), then the probability that any one of them happens equals the sum of their individual probabilities. If there's a 30% chance of rain and a 20% chance of snow, and it can't rain and snow at the same time, there's a 50% chance of precipitation. This rule extends to countably infinite collections of disjoint events.
What the axioms give us for free
From just three axioms, a surprising wealth of results follows logically. No additional assumptions are needed — these are theorems, not postulates:
- P(∅) = 0 — the impossible event has zero probability.
- P(Aᶜ) = 1 − P(A) — the probability of "not A" is one minus the probability of A.
- P(A ∪ B) = P(A) + P(B) − P(A ∩ B) — the inclusion-exclusion principle.
- If A ⊆ B, then P(A) ≤ P(B) — a subset can't be more likely than its superset.
- 0 ≤ P(A) ≤ 1 — probabilities are always between 0 and 1.
These feel "obvious," but that's exactly the point. The axioms formalize our intuition so precisely that the obvious becomes provable and the non-obvious becomes discoverable.
Random variables: from outcomes to numbers
A sample space can be anything — coin sides, weather states, DNA sequences. But mathematics works with numbers. A random variable is a function that translates each outcome into a number. The key requirement: for any value , the set must be in the σ-algebra F. In other words, the question "is X less than or equal to x?" must be a valid question we can assign probability to. This requirement is called measurability.
Think of a random variable as a numerical reading from an experiment. A thermometer is a random variable: the outcome is the state of the atmosphere (an element of Ω), and the reading is a number on the scale. Different thermometers (different random variables) can extract different numbers from the same experiment.
Expectation: the Lebesgue integral in disguise
The expected value of a random variable is its "long-run average" — the value you'd converge to if you repeated the experiment infinitely many times and averaged the results. Kolmogorov defined it formally as the Lebesgue integral of X with respect to the probability measure P.
Why Lebesgue and not Riemann? Because the Riemann integral slices the x-axis into equal intervals, which doesn't work for arbitrary probability distributions. The Lebesgue integral slices the y-axis instead: it groups together all outcomes that give the same value, then weights by their total probability. This works for any measurable function on any probability space — discrete, continuous, or anything in between.
Conditional probability: updating belief with evidence
Conditional probability answers: "Given that B happened, what is the probability of A?" Kolmogorov defined it as a derived concept, not a primitive one. It falls out of the axioms naturally:
Think of it as zooming in on the sample space. Once you know B occurred, the universe of possibilities shrinks from Ω to B. The conditional probability rescales everything so that B becomes the new "certain event" — its probability becomes 1, and everything inside B is proportionally adjusted.
This definition works beautifully when P(B) > 0. For the case P(B) = 0 — which arises constantly in continuous probability, e.g., "given that the dart landed at exactly this point" — Kolmogorov introduced via the Radon-Nikodym theorem, one of the book's deepest contributions.
Bayes' theorem follows immediately from the definition of conditional probability by simple algebra. It reverses the direction of conditioning: if you know P(B|A) and want P(A|B), Bayes tells you how:
Independence: the multiplicative rule
Two events A and B are independent if knowing that one occurred tells you nothing about the other. Kolmogorov formalized this as a definition derived from his axioms: A and B are independent if and only if .
Notice this is not an axiom — it's a definition. But it's an extraordinarily powerful one. Independence lets you multiply probabilities: the chance of flipping 10 heads in a row is , because each flip is independent of the others. Without this concept, we could not model anything from coin sequences to .
Kolmogorov extended independence to arbitrary collections: events are mutually independent if the probability of any subcollection's intersection equals the product of their individual probabilities. Pairwise independence is necessary but not sufficient — a subtle point that catches even experienced practitioners.
The law of large numbers: probability meets the real world
The law of large numbers is the bridge between abstract probability and physical observation. It says: if you repeat an experiment independently many times, the average of the results will converge to the expected value.
Kolmogorov proved the strong law: the average converges almost surely (with probability 1), not just "in probability" as the weak law states. This means for almost every possible infinite sequence of experiments, the running average will eventually settle at E[X] and stay there — not just get close on average.
This theorem is why statistics works. It's why polling can predict elections, why insurance companies can set premiums, and why a neural network on more data gives better results. The averages converge because the axioms guarantee it.
The axioms in code
Simplified to show the idea — not the real implementation.
import numpy as np
# === Build a probability space (Ω, F, P) ===
# Ω: sample space — outcomes of rolling a fair die
omega = set([1, 2, 3, 4, 5, 6])
# P: probability measure — each face equally likely
P = dict((outcome, 1/6) for outcome in omega)
def prob(event):
"""Compute P(event) by summing over outcomes."""
return sum(P[o] for o in event if o in P)
# --- Verify the three axioms ---
# Axiom 1: Non-negativity
assert all(p >= 0 for p in P.values()), "Axiom 1 violated!"
# Axiom 2: Normalization
assert abs(prob(omega) - 1.0) < 1e-10, "Axiom 2 violated!"
# Axiom 3: Additivity (for disjoint events)
A = set([1, 2]) # rolling 1 or 2
B = set([3, 4, 5]) # rolling 3, 4, or 5
assert A & B == set(), "A and B must be disjoint"
assert abs(prob(A | B) - (prob(A) + prob(B))) < 1e-10
# --- Derived results (theorems, not axioms) ---
# P(∅) = 0
assert prob(set()) == 0
# P(Aᶜ) = 1 - P(A)
A_complement = omega - A
assert abs(prob(A_complement) - (1 - prob(A))) < 1e-10
# Conditional probability: P(A|B) = P(A∩B) / P(B)
C = set([2, 4, 6]) # even numbers
D = set([1, 2, 3]) # numbers ≤ 3
p_C_given_D = prob(C & D) / prob(D)
print("P(even | ≤3) = %.4f" % p_C_given_D) # 0.3333
# Independence check: P(A∩B) =? P(A)·P(B)
independent = abs(prob(C & D) - prob(C) * prob(D)) < 1e-10
print("Even and ≤3 independent?", independent) # True!Why it matters — from 1933 to AI
1654
Pascal & Fermat correspondence
The birth of mathematical probability — two mathematicians exchange letters about gambling problems, laying the combinatorial foundations.
1812
Laplace's "Théorie analytique des probabilités"
The classical definition: probability as the ratio of favorable to total outcomes. Powerful but limited to equiprobable finite spaces.
1933
Kolmogorov's Grundbegriffe
The axiomatic foundations. Probability becomes measure theory, all paradoxes are resolved, and a universal mathematical language is born.
1946
Monte Carlo methods
Ulam and von Neumann use random sampling — grounded in Kolmogorov's framework — to simulate nuclear physics. Probability becomes a computational tool.
1948
Shannon's information theory
Entropy, mutual information, channel capacity — all built on Kolmogorov's probability spaces and his definition of expectation.
1970
Kolmogorov complexity
Kolmogorov himself defines algorithmic complexity — the shortest program that produces a string. A bridge between probability and computation.
2012
Deep learning revolution
Every loss function, every gradient descent step, every batch normalization layer is a theorem from Kolmogorov's framework applied at industrial scale.
Kolmogorov gave probability its mathematical passport. Before 1933, it was a guest in the house of mathematics; after 1933, it had its own room with the same structural guarantees as algebra, geometry, and analysis. Every probabilistic statement you encounter — in a textbook, a paper, or a machine learning library — rests on sixty pages written in German, ninety-one years ago.
CitationKolmogorov, A. N.. Grundbegriffe der Wahrscheinlichkeitsrechnung (Foundations of the Theory of Probability). Springer (Ergebnisse der Mathematik), 1933.
Terms in this paper
- Probabilityالاحتمالية
- Distributionالتوزيع الإحصائي
- Conditional Probabilityالاحتمال الشرطي
- Random Variableمتغيّر عشوائي
- Sigma Algebraجبر سيغما
- Sample Spaceفضاء العيّنة
- Law of Large Numbersقانون الأعداد الكبيرة
- Bayes' Theoremمبرهنة بايز الاحتمالية