Mathematics1763foundational10 min read
An Essay Towards Solving a Problem in the Doctrine of Chances
مقالة نحو حلّ مسألة في نظرية الاحتمالات
Bayes, T. · Price, R. — Philosophical Transactions of the Royal Society of London
The problem
By the mid-1700s, probability theory could answer forward questions: if a coin lands heads 60% of the time, what is the chance of 7 heads in 10 flips? But the far more useful inverse question had no solution: given that I observed 7 heads in 10 flips, what can I say about the coin's true bias? No one had a principled method to reason backward from observations to the unknown probability that generated them.
The contribution
Bayes introduced the first framework for . He showed that if you assume a belief about an unknown probability (he used a — equal ignorance across all values), then after observing successes and failures in repeated trials, you can compute a distribution that concentrates around the true value as evidence accumulates. The mathematical core is that a uniform prior combined with binomial observations yields a Beta posterior — the first instance of Bayesian updating.
The impact
Bayes' essay is the seed of — the engine behind spam filters, medical diagnosis, A/B testing, search engines, and the training of every modern language model. Whenever a system updates its beliefs from data, it is doing what Bayes described in 1763. Laplace later generalized the framework, but the conceptual breakthrough — reasoning backward from evidence to cause — belongs to Bayes.
Imagine you are blindfolded in a dark room. Someone rolls a ball onto a long table — you have no idea where it stopped. Now they roll more balls, one at a time, and each time they only tell you: "left of the first ball" or "right of the first ball."
After hearing "left" 7 times and "right" 3 times out of 10 rolls, you start to form a picture: the first ball is probably sitting somewhere around the 70% mark of the table, but you are not certain — it could be at 60% or 80%.
Each new roll sharpens your picture. That is Bayesian inference: you never see the ball directly, but each piece of evidence narrows down where it must be.
The problem: reasoning backward from outcomes to causes
Before Bayes, probability worked in one direction only. If you know that a die is fair, you can calculate the chance of rolling three sixes in a row: . This is the forward problem — from a known cause to expected outcomes.
But real life almost always asks the opposite question. A doctor sees symptoms and asks: what disease most likely caused them? A scientist observes data and asks: which hypothesis best explains it? You flip a coin 100 times and see 73 heads — is this coin biased?
This is the inverse problem — from observed outcomes to the unknown cause. Before Bayes, no one had a systematic way to answer it.
The thought experiment: a ball on a table
Bayes designed a beautifully simple thought experiment. Picture a flat table that stretches from 0 to 1. A first ball is rolled onto it and stops at some unknown point — you cannot see where. Then you roll more balls, and for each one you only learn whether it landed left or right of the first ball.
Landing left of has probability , and landing right has probability . This is exactly a sequence of Bernoulli trials with unknown probability . If out of balls land left, you want to know: where is ?
Bayes' insight was to treat itself as a random variable with a prior distribution. Since the first ball was rolled uniformly (at random across the table), is equally likely to be anywhere from 0 to 1 — a uniform prior. The evidence from the subsequent balls then reshapes this flat prior into a peaked posterior concentrated near the observed proportion .
The core idea: prior × evidence = posterior
Bayes' contribution can be distilled into one sentence: your updated belief about an unknown quantity equals your prior belief, multiplied by how well each possible value explains the observed data, then normalized so the total probability is 1.
Think of it as a courtroom. Before the trial, the jury has a prior impression of guilt or innocence. Each piece of evidence (witness testimony, forensic results) either strengthens or weakens that impression. The verdict — the posterior — is the prior updated by all the evidence.
Mathematically, for a continuous unknown with observed data :
Read the formula as a recipe: start with your prior , weigh each possible value of by how well it explains what you saw — the — and then normalize. The result, the posterior , is your rational updated belief.
The denominator is the same for every value of , so it acts as a scaling constant. This is why you often see the theorem written in proportional form:
The uniform prior: encoding total ignorance
A crucial question in Bayesian reasoning is: what prior should you choose when you know nothing? Bayes' answer was the uniform distribution — assign equal probability density to every value from 0 to 1.
This is sometimes called the "principle of insufficient reason": if you have no reason to favor one value of over another, treat them all equally. In Bayes' table experiment, it arises naturally — the first ball was rolled at random, so its stopping point is uniformly distributed.
The beauty of this choice is that it lets the data speak. With a uniform prior, the posterior is shaped entirely by the evidence. As more evidence arrives, the prior's influence fades and the posterior converges on the true value — regardless of where the prior started.
The mathematical engine: Beta-Binomial conjugacy
Here is where Bayes' thought experiment yields a beautiful closed-form result. The setup is:
- Prior: , which is a special case of the distribution.
- Likelihood: successes in trials follows a distribution.
The intuition is: the prior says "any coin bias is equally plausible" and the data says "I saw heads in flips." Combining them through Bayes' theorem produces the posterior:
This result is remarkable for two reasons. First, the posterior has the same family as the prior — both are Beta distributions. When the prior and posterior belong to the same family, we say the prior is conjugate to the likelihood. means you can update your beliefs repeatedly without ever leaving the Beta family: each new batch of evidence just nudges the two parameters.
Second, the posterior mean is , which starts at (total ignorance) and moves toward the observed proportion as grows. The posterior concentrates: its variance shrinks as , so with enough data your uncertainty vanishes.
Bayesian updating in action
The power of Bayesian reasoning is that it is sequential. Yesterday's posterior becomes today's prior. Suppose you flip a coin 5 times and get 4 heads. Your posterior is Beta(5,2). Tomorrow you flip it 5 more times and get 3 heads. You don't need to restart — your new posterior is Beta(5+3, 2+2) = Beta(8,4), exactly as if you had observed 7 heads in 10 flips all at once.
This is why Bayesian methods are ideal for learning systems that see data incrementally — like recommendation engines, fraud detection, and language models during .
The same idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
from scipy.stats import beta
def bayesian_update(prior_a, prior_b, heads, tails):
"""Update Beta prior with observed coin flips.
prior_a, prior_b: Beta parameters (start at 1,1 for uniform)
heads, tails: observed counts
Returns: posterior Beta parameters"""
post_a = prior_a + heads
post_b = prior_b + tails
return post_a, post_b
# Start with total ignorance: Beta(1,1) = Uniform
a, b = 1, 1
# Observe 7 heads and 3 tails
a, b = bayesian_update(a, b, heads=7, tails=3)
# Posterior: Beta(8, 4)
# Posterior mean = a / (a+b) — our best estimate of the coin's bias
print(f"Posterior mean: {a/(a+b):.3f}") # 0.667
print(f"95% credible interval: "
f"{beta.ppf(0.025, a, b):.3f} – "
f"{beta.ppf(0.975, a, b):.3f}") # 0.399 – 0.881
# Observe 20 more flips: 14 heads, 6 tails
a, b = bayesian_update(a, b, heads=14, tails=6)
# Now posterior: Beta(22, 10)
print(f"\nAfter 30 flips:")
print(f"Posterior mean: {a/(a+b):.3f}") # 0.688
print(f"95% credible interval: "
f"{beta.ppf(0.025, a, b):.3f} – "
f"{beta.ppf(0.975, a, b):.3f}") # 0.521 – 0.830
# The interval shrank — more data = more certainty.
# That's Bayes: prior + evidence → sharper posterior.Richard Price: the editor who published the revolution
Bayes died in 1761 without publishing his essay. His friend Richard Price — a minister, philosopher, and mathematician — found the manuscript among Bayes' papers, recognized its significance, and spent two years editing and extending it. Price added a detailed appendix with numerical calculations showing the theorem in action, and read the paper to the Royal Society on December 23, 1763.
Price's contribution was more than editorial. He sharpened the practical implications: the theorem could be used to evaluate the probability that observed regularities (like the sun rising every day) reflect genuine laws of nature rather than coincidence. Price saw in Bayes' mathematics a philosophical tool for reasoning about evidence, a vision that anticipates modern Bayesian epistemology.
Why it mattered: from 1763 to modern AI
Bayes' essay lay relatively dormant until Pierre-Simon Laplace independently rediscovered and generalized the idea in the 1770s–1810s. But the core insight — reasoning backward from data to hypotheses — has become one of the most consequential ideas in science and engineering.
1763
Bayes' Essay published posthumously
Richard Price reads Bayes' essay to the Royal Society. The first formal framework for inverse probability — reasoning from observed data to unknown causes.
1774
Laplace generalizes the theorem
Laplace independently derives a more general form of Bayes' theorem and applies it to astronomy, population statistics, and legal reasoning.
1812
Laplace's Théorie Analytique des Probabilités
Laplace publishes his magnum opus codifying probability theory, with Bayesian methods at its core. The "Laplacian" formulation dominates for a century.
1950
Bayesian revival in statistics
Statisticians like Harold Jeffreys, L.J. Savage, and Dennis Lindley champion Bayesian methods against the dominant frequentist school, reigniting a foundational debate.
1990
Bayesian networks and spam filters
Judea Pearl formalizes Bayesian networks. Spam filters use Naive Bayes classifiers. Bayesian methods enter practical computing.
2012
Bayesian deep learning
Researchers apply Bayesian ideas to neural networks — modeling uncertainty, Bayesian optimization for hyperparameters, and posterior inference for model weights.
2024
Bayesian reasoning in LLMs
Modern language models implicitly perform Bayesian-like updating — each token of context refines the model's predictive distribution, echoing Bayes' sequential evidence accumulation.
CitationBayes, T. and Price, R.. An Essay Towards Solving a Problem in the Doctrine of Chances. Philosophical Transactions of the Royal Society of London, 1763.
Terms in this paper
- Inverse Probabilityالاحتمالية العكسية
- Bayesian Inferenceالاستدلال البايزي
- Conjugacyالترافق
- Beta Distributionتوزيع بيتا
- Uniform Priorالاحتمال القبلي المنتظم
- Uninformative Priorالقبلي غير الإعلامي
- Posterior Concentrationتركّز البعدي