Mathematics1763foundational10 min read

An Essay Towards Solving a Problem in the Doctrine of Chances

مقالة نحو حلّ مسألة في نظرية الاحتمالات

Bayes, T. · Price, R. — Philosophical Transactions of the Royal Society of London

The problem

By the mid-1700s, probability theory could answer forward questions: if a coin lands heads 60% of the time, what is the chance of 7 heads in 10 flips? But the far more useful inverse question had no solution: given that I observed 7 heads in 10 flips, what can I say about the coin's true bias? No one had a principled method to reason backward from observations to the unknown probability that generated them.

The contribution

Bayes introduced the first framework for . He showed that if you assume a belief about an unknown probability (he used a — equal ignorance across all values), then after observing successes and failures in repeated trials, you can compute a distribution that concentrates around the true value as evidence accumulates. The mathematical core is that a uniform prior combined with binomial observations yields a Beta posterior — the first instance of Bayesian updating.

The impact

Bayes' essay is the seed of — the engine behind spam filters, medical diagnosis, A/B testing, search engines, and the training of every modern language model. Whenever a system updates its beliefs from data, it is doing what Bayes described in 1763. Laplace later generalized the framework, but the conceptual breakthrough — reasoning backward from evidence to cause — belongs to Bayes.

Imagine you are blindfolded in a dark room. Someone rolls a ball onto a long table — you have no idea where it stopped. Now they roll more balls, one at a time, and each time they only tell you: "left of the first ball" or "right of the first ball."

After hearing "left" 7 times and "right" 3 times out of 10 rolls, you start to form a picture: the first ball is probably sitting somewhere around the 70% mark of the table, but you are not certain — it could be at 60% or 80%.

Each new roll sharpens your picture. That is Bayesian inference: you never see the ball directly, but each piece of evidence narrows down where it must be.

The problem: reasoning backward from outcomes to causes

Before Bayes, probability worked in one direction only. If you know that a die is fair, you can calculate the chance of rolling three sixes in a row: (1/6)3(1/6)^3. This is the forward problem — from a known cause to expected outcomes.

But real life almost always asks the opposite question. A doctor sees symptoms and asks: what disease most likely caused them? A scientist observes data and asks: which hypothesis best explains it? You flip a coin 100 times and see 73 heads — is this coin biased?

This is the inverse problem — from observed outcomes to the unknown cause. Before Bayes, no one had a systematic way to answer it.

Open in Lab
Compare forward vs. inverse probability: same math, opposite directions.
The demo wakes as you arrive…

The thought experiment: a ball on a table

Bayes designed a beautifully simple thought experiment. Picture a flat table that stretches from 0 to 1. A first ball is rolled onto it and stops at some unknown point theta\\theta — you cannot see where. Then you roll nn more balls, and for each one you only learn whether it landed left or right of the first ball.

Landing left of theta\\theta has probability theta\\theta, and landing right has probability 1−theta1-\\theta. This is exactly a sequence of Bernoulli trials with unknown probability theta\\theta. If kk out of nn balls land left, you want to know: where is theta\\theta?

Bayes' insight was to treat theta\\theta itself as a random variable with a prior distribution. Since the first ball was rolled uniformly (at random across the table), theta\\theta is equally likely to be anywhere from 0 to 1 — a uniform prior. The evidence from the subsequent balls then reshapes this flat prior into a peaked posterior concentrated near the observed proportion k/nk/n.

Open in Lab
Roll balls onto the table and watch your belief about θ sharpen.
The demo wakes as you arrive…

The core idea: prior × evidence = posterior

Bayes' contribution can be distilled into one sentence: your updated belief about an unknown quantity equals your prior belief, multiplied by how well each possible value explains the observed data, then normalized so the total probability is 1.

Think of it as a courtroom. Before the trial, the jury has a prior impression of guilt or innocence. Each piece of evidence (witness testimony, forensic results) either strengthens or weakens that impression. The verdict — the posterior — is the prior updated by all the evidence.

Mathematically, for a continuous unknown theta\\theta with observed data DD:

P(θ∣D)=P(D∣θ)⋅P(θ)P(D)P(\theta | D) = \frac{P(D | \theta) \cdot P(\theta)}{P(D)}
Bayes' Theorem — the engine of inverse probability — P(θ|D) = posterior · P(D|θ) = likelihood — how probable is the data if θ were true? · P(θ) = prior — your belief about θ before seeing data · P(D) = normalizer — ensures the posterior integrates to 1
Open in Lab
Step through the three ingredients of Bayes' theorem.
The demo wakes as you arrive…

Read the formula as a recipe: start with your prior P(theta)P(\\theta), weigh each possible value of theta\\theta by how well it explains what you saw — the P(D∣theta)P(D|\\theta) — and then normalize. The result, the posterior P(theta∣D)P(\\theta|D), is your rational updated belief.

The denominator P(D)P(D) is the same for every value of theta\\theta, so it acts as a scaling constant. This is why you often see the theorem written in proportional form:

P(θ∣D)∝P(D∣θ)⋅P(θ)P(\theta | D) \propto P(D | \theta) \cdot P(\theta)
The proportional form — posterior is proportional to likelihood × prior — "proportional to" means the shape is determined by likelihood × prior; the normalizer just makes it a proper probability distribution

The uniform prior: encoding total ignorance

A crucial question in Bayesian reasoning is: what prior should you choose when you know nothing? Bayes' answer was the uniform distribution — assign equal probability density to every value from 0 to 1.

This is sometimes called the "principle of insufficient reason": if you have no reason to favor one value of theta\\theta over another, treat them all equally. In Bayes' table experiment, it arises naturally — the first ball was rolled at random, so its stopping point is uniformly distributed.

The beauty of this choice is that it lets the data speak. With a uniform prior, the posterior is shaped entirely by the evidence. As more evidence arrives, the prior's influence fades and the posterior converges on the true value — regardless of where the prior started.

The mathematical engine: Beta-Binomial conjugacy

Here is where Bayes' thought experiment yields a beautiful closed-form result. The setup is:

  • Prior: thetasimtextUniform(0,1)\\theta \\sim \\text{Uniform}(0,1), which is a special case of the textBeta(1,1)\\text{Beta}(1,1) distribution.
  • Likelihood: kk successes in nn trials follows a textBinomial(n,theta)\\text{Binomial}(n,\\theta) distribution.

The intuition is: the prior says "any coin bias is equally plausible" and the data says "I saw kk heads in nn flips." Combining them through Bayes' theorem produces the posterior:

θ∣(k,n)∼Beta(k+1,  n−k+1)\theta | (k, n) \sim \text{Beta}(k+1,\; n-k+1)
The Beta posterior — Bayes' main result — Start with Beta(1,1) = Uniform · observe k successes and n−k failures · the posterior is Beta(k+1, n−k+1) — peaked near k/n and narrowing with more data

This result is remarkable for two reasons. First, the posterior has the same family as the prior — both are Beta distributions. When the prior and posterior belong to the same family, we say the prior is conjugate to the likelihood. means you can update your beliefs repeatedly without ever leaving the Beta family: each new batch of evidence just nudges the two parameters.

Second, the posterior mean is (k+1)/(n+2)(k+1)/(n+2), which starts at 1/21/2 (total ignorance) and moves toward the observed proportion k/nk/n as nn grows. The posterior concentrates: its variance shrinks as 1/n1/n, so with enough data your uncertainty vanishes.

Open in Lab
Drag the sliders to set k (successes) and n (trials). Watch the Beta posterior sharpen.
The demo wakes as you arrive…

Bayesian updating in action

The power of Bayesian reasoning is that it is sequential. Yesterday's posterior becomes today's prior. Suppose you flip a coin 5 times and get 4 heads. Your posterior is Beta(5,2). Tomorrow you flip it 5 more times and get 3 heads. You don't need to restart — your new posterior is Beta(5+3, 2+2) = Beta(8,4), exactly as if you had observed 7 heads in 10 flips all at once.

This is why Bayesian methods are ideal for learning systems that see data incrementally — like recommendation engines, fraud detection, and language models during .

Open in Lab
Click 'Flip' to add evidence coin-by-coin. Watch the posterior evolve in real time.
The demo wakes as you arrive…

The same idea in code

Bayesian coin estimation — completepython

Simplified to show the idea — not the real implementation.

import numpy as np
from scipy.stats import beta

def bayesian_update(prior_a, prior_b, heads, tails):
    """Update Beta prior with observed coin flips.
       prior_a, prior_b: Beta parameters (start at 1,1 for uniform)
       heads, tails: observed counts
       Returns: posterior Beta parameters"""
    post_a = prior_a + heads
    post_b = prior_b + tails
    return post_a, post_b

# Start with total ignorance: Beta(1,1) = Uniform
a, b = 1, 1

# Observe 7 heads and 3 tails
a, b = bayesian_update(a, b, heads=7, tails=3)
# Posterior: Beta(8, 4)

# Posterior mean = a / (a+b) — our best estimate of the coin's bias
print(f"Posterior mean: {a/(a+b):.3f}")        # 0.667
print(f"95% credible interval: "
      f"{beta.ppf(0.025, a, b):.3f} – "
      f"{beta.ppf(0.975, a, b):.3f}")          # 0.399 – 0.881

# Observe 20 more flips: 14 heads, 6 tails
a, b = bayesian_update(a, b, heads=14, tails=6)
# Now posterior: Beta(22, 10)

print(f"\nAfter 30 flips:")
print(f"Posterior mean: {a/(a+b):.3f}")        # 0.688
print(f"95% credible interval: "
      f"{beta.ppf(0.025, a, b):.3f} – "
      f"{beta.ppf(0.975, a, b):.3f}")          # 0.521 – 0.830

# The interval shrank — more data = more certainty.
# That's Bayes: prior + evidence → sharper posterior.

Richard Price: the editor who published the revolution

Bayes died in 1761 without publishing his essay. His friend Richard Price — a minister, philosopher, and mathematician — found the manuscript among Bayes' papers, recognized its significance, and spent two years editing and extending it. Price added a detailed appendix with numerical calculations showing the theorem in action, and read the paper to the Royal Society on December 23, 1763.

Price's contribution was more than editorial. He sharpened the practical implications: the theorem could be used to evaluate the probability that observed regularities (like the sun rising every day) reflect genuine laws of nature rather than coincidence. Price saw in Bayes' mathematics a philosophical tool for reasoning about evidence, a vision that anticipates modern Bayesian epistemology.

Why it mattered: from 1763 to modern AI

Bayes' essay lay relatively dormant until Pierre-Simon Laplace independently rediscovered and generalized the idea in the 1770s–1810s. But the core insight — reasoning backward from data to hypotheses — has become one of the most consequential ideas in science and engineering.

  1. 1763

    Bayes' Essay published posthumously

    Richard Price reads Bayes' essay to the Royal Society. The first formal framework for inverse probability — reasoning from observed data to unknown causes.

  2. 1774

    Laplace generalizes the theorem

    Laplace independently derives a more general form of Bayes' theorem and applies it to astronomy, population statistics, and legal reasoning.

  3. 1812

    Laplace's Théorie Analytique des Probabilités

    Laplace publishes his magnum opus codifying probability theory, with Bayesian methods at its core. The "Laplacian" formulation dominates for a century.

  4. 1950

    Bayesian revival in statistics

    Statisticians like Harold Jeffreys, L.J. Savage, and Dennis Lindley champion Bayesian methods against the dominant frequentist school, reigniting a foundational debate.

  5. 1990

    Bayesian networks and spam filters

    Judea Pearl formalizes Bayesian networks. Spam filters use Naive Bayes classifiers. Bayesian methods enter practical computing.

  6. 2012

    Bayesian deep learning

    Researchers apply Bayesian ideas to neural networks — modeling uncertainty, Bayesian optimization for hyperparameters, and posterior inference for model weights.

  7. 2024

    Bayesian reasoning in LLMs

    Modern language models implicitly perform Bayesian-like updating — each token of context refines the model's predictive distribution, echoing Bayes' sequential evidence accumulation.

CitationBayes, T. and Price, R.. An Essay Towards Solving a Problem in the Doctrine of Chances. Philosophical Transactions of the Royal Society of London, 1763.

Terms in this paper