Information Theory1948foundational10 min read

A Mathematical Theory of Communication

النظرية الرياضية للاتصال

Shannon, C. E. — Bell System Technical Journal

The problem

Before 1948, engineers treated communication as an analog signal problem — amplify louder, filter better. Nobody had a way to answer the most basic question: how much information can a channel carry? Without that answer, engineers couldn't know whether their systems were close to optimal or wasting most of their capacity. There was no theory of information itself — no unit, no measure, no limits.

The contribution

Shannon defined information mathematically using — a measure of surprise or uncertainty. He introduced the as the fundamental unit, proved that any source can be compressed down to its entropy (), and proved that reliable communication is possible over any noisy channel as long as the transmission rate stays below the (). He gave the formula C = B log₂(1 + S/N) for continuous channels.

The impact

Shannon's paper created an entire field — — and set the mathematical foundations for the digital age. CDs, Wi-Fi, 5G, JPEG, MP3, QR codes, deep-space communication, cryptography, and even machine learning all descend from the ideas in this single 1948 paper. The bit became the atom of the information era.

Imagine a post office that charges by surprise, not by weight. A letter saying "the sun rose this morning" costs almost nothing — everyone expected that. A letter saying "a new continent was discovered" costs a fortune — nobody saw it coming.

Shannon proved that every communication channel is exactly this post office. The entropy of a message source is its average surprise per symbol, measured in bits. And every channel has a maximum surprise-per-second it can deliver — its capacity. Stay under that limit, and you can communicate with essentially zero errors. Exceed it, and no amount of engineering can save you.

The blueprint: how every communication system works

Shannon's first insight was to abstract away the physical details. Telephone, telegraph, television, radio — they all reduce to the same five-block diagram:

Information Source → Transmitter () → Channel (+ ) → Receiver () → Destination

The source produces a message (text, speech, images). The encoder converts it into a signal suited for the channel. The channel carries the signal — but also adds noise. The decoder reconstructs the message from the noisy signal. The destination receives it.

This abstraction was revolutionary: once you have this diagram, you can reason about any communication system mathematically, regardless of whether the channel is a copper wire, a fiber optic cable, or the vacuum of space.

Open in Lab
Click on any block to learn its role in the communication chain.
The demo wakes as you arrive…

Measuring the unmeasurable: information as surprise

Before Shannon, "information" was a vague word. He gave it a precise definition: information is the resolution of uncertainty. A message that tells you something you already knew carries zero information. A message that tells you something completely unexpected carries maximum information.

The key question is: if a source generates symbols from an alphabet — say letters A through Z — each with some , how much "surprise" does the source produce on average per symbol?

Shannon called this average surprise the entropy of the source, borrowing the term from thermodynamics. A source that always emits the same letter has zero entropy — no surprise, no information. A source where every letter is equally likely has maximum entropy — maximum surprise, maximum information per symbol.

Think of it this way: in English, after "q" the letter "u" almost always follows — so "u" after "q" carries almost no information. But after "t", many letters could follow — so whichever one appears carries real information. Entropy measures this averaged over the entire source.

H(X)=−∑i=1np(xi)log⁡2p(xi)H(X) = -\sum_{i=1}^{n} p(x_i) \log_2 p(x_i)
Shannon entropy — the average surprise per symbol — Each symbol's surprise = -log₂(probability). Rare symbols carry more bits of surprise. H averages this over all symbols, weighted by how often each appears. The result: bits per symbol.
Open in Lab
Drag the probability sliders to see how entropy changes. Equal probabilities → maximum entropy.
The demo wakes as you arrive…

For the simplest case — a coin flip with probability pp for heads — entropy simplifies to the binary entropy function: H(p)=−plog⁡2p−(1−p)log⁡2(1−p)H(p) = -p \log_2 p - (1-p) \log_2 (1-p). This beautiful curve peaks at p=0.5p = 0.5 (1 bit — maximum uncertainty) and drops to zero at p=0p = 0 or p=1p = 1 (certainty — no surprise).

Open in Lab
The classic binary entropy curve — drag the point to explore H(p).
The demo wakes as you arrive…

Source coding: squeezing out the waste

Shannon's first theorem says: you can compress a message down to its entropy, but no further. If a source has entropy H = 2.5 bits/symbol, you need at least 2.5 bits per symbol on average to represent it faithfully — but you can get arbitrarily close to that limit.

The idea is simple: give short codes to common symbols and long codes to rare ones. If "e" appears 13% of the time in English, give it a short code. If "z" appears 0.07%, give it a long code. This is exactly what Morse code did intuitively — "e" is a single dot.

Shannon proved this mathematically: the optimal average code length equals the source entropy. Every compression algorithm since — Huffman coding, LZW (used in GIF), DEFLATE (used in ZIP), and modern algorithms — is chasing this bound.

Open in Lab
See how assigning short codes to frequent symbols compresses the message toward entropy.
The demo wakes as you arrive…

Channel capacity: the speed limit of communication

Shannon's second great result answers: how fast can you send data through a noisy channel without errors?

Every channel has a number called its capacity — the maximum rate at which you can transmit information with arbitrarily low error probability. Think of it as a highway speed limit: drive under it and traffic flows perfectly; try to exceed it and collisions become inevitable.

For a continuous channel with BB (Hz) and signal-to-noise ratio S/NS/N, the capacity is given by the famous :

C=Blog⁡2 ⁣(1+SN)C = B \log_2\!\left(1 + \frac{S}{N}\right)
Shannon-Hartley theorem — the channel speed limit — B = bandwidth in Hz · S/N = signal-to-noise ratio · C = maximum error-free data rate in bits/second. More bandwidth or less noise → higher capacity.

Notice the trade-off: you can compensate for noise by using more bandwidth, or compensate for limited bandwidth by increasing signal power. But there is always a ceiling — no real-world system can exceed Shannon capacity.

Open in Lab
Adjust bandwidth and SNR to see how channel capacity changes.
The demo wakes as you arrive…

The miracle: reliable communication over noisy channels

Here is perhaps the most surprising result in all of information theory — Shannon's noisy channel coding theorem:

As long as your transmission rate R is below the channel capacity C, there exist encoding schemes that make the error probability as small as you want — essentially zero.

Before Shannon, engineers assumed that noise always corrupts some data. Shannon proved that by adding carefully designed — extra bits that don't carry new information but help detect and correct errors — you can communicate reliably through any channel, no matter how noisy, provided R<CR < C.

The catch: Shannon proved these codes exist but didn't construct them. Finding practical codes that approach the Shannon limit became a decades-long quest — from Hamming codes (1950) to Reed-Solomon codes (used in CDs and QR codes) to turbo codes (1993) and LDPC codes (used in 5G and Wi-Fi 6) that nearly touch the limit.

Open in Lab
Send bits through a noisy channel — first without coding, then with error-correction. See the difference.
The demo wakes as you arrive…

Mutual information: how much does the output tell you about the input?

The channel capacity CC is defined as the maximum between the channel input XX and output YY: C=max⁡p(x)I(X;Y)C = \max_{p(x)} I(X;Y) Mutual information I(X;Y)=H(X)−H(X∣Y)I(X;Y) = H(X) - H(X|Y) measures how much knowing the output reduces your uncertainty about the input. In a noiseless channel, I(X;Y)=H(X)I(X;Y) = H(X) — the output tells you everything. In a completely noisy channel, I(X;Y)=0I(X;Y) = 0 — the output tells you nothing.

Picture two overlapping circles (a Venn diagram): one is H(X)H(X) (the sender's uncertainty), the other is H(Y)H(Y) (the receiver's uncertainty). The overlap is the mutual information — the shared knowledge. Channel capacity asks: how big can we make that overlap?

I(X;Y)=H(X)−H(X∣Y)=H(Y)−H(Y∣X)I(X;Y) = H(X) - H(X|Y) = H(Y) - H(Y|X)
Mutual information — the shared knowledge between sender and receiver — H(X) = sender's entropy · H(X|Y) = remaining uncertainty after observing the output · the difference = how much the channel actually conveyed

Redundancy: the hidden waste — and hidden protection — in language

Shannon estimated that English text has about 50% redundancy — roughly half the letters could be removed and you'd still reconstruct the message. This is why you can read "ths sntnc hs n vwls" — the redundancy of English lets your brain fill in gaps.

This redundancy has two sides:

  • For compression, it's waste. removes it to approach entropy.
  • For error correction, it's a gift. Natural redundancy helps humans detect typos. Engineered redundancy (error-correcting codes) lets machines do the same.

Shannon realized that the fundamental tension in communication is between these two forces: remove redundancy to compress, add redundancy to protect against noise. His theorems tell you exactly how far you can push each direction.

The core ideas in code

Shannon entropy and channel capacity, from scratchpython

Simplified to show the idea — not the real implementation.

import numpy as np

def entropy(probs):
    """Shannon entropy: average surprise per symbol, in bits."""
    probs = np.array(probs)
    probs = probs[probs > 0]                # 0 log 0 = 0 by convention
    return -np.sum(probs * np.log2(probs))

# Example: a 4-symbol source
p = [0.5, 0.25, 0.125, 0.125]
H = entropy(p)
print(f"Entropy H = {H:.3f} bits/symbol")   # 1.750 — that's the compression limit

# Binary entropy: the special case of a coin flip
def binary_entropy(p):
    if p == 0 or p == 1: return 0
    return -p * np.log2(p) - (1-p) * np.log2(1-p)

print(f"H(0.5) = {binary_entropy(0.5):.3f}")  # 1.0 — max uncertainty
print(f"H(0.9) = {binary_entropy(0.9):.3f}")  # 0.469 — biased coin → less surprise

# Shannon-Hartley channel capacity
def channel_capacity(bandwidth_hz, snr_linear):
    """C = B * log2(1 + S/N) in bits per second."""
    return bandwidth_hz * np.log2(1 + snr_linear)

# A typical Wi-Fi channel: 20 MHz bandwidth, 30 dB SNR
B = 20e6                        # 20 MHz
snr_db = 30
snr = 10 ** (snr_db / 10)       # convert dB to linear
C = channel_capacity(B, snr)
print(f"Wi-Fi capacity ≈ {C/1e6:.0f} Mbit/s")  # ~199 Mbit/s — theoretical max

Why it mattered

  1. 1948

    Shannon's paper

    Published in the Bell System Technical Journal. Created information theory as a field and introduced the bit, entropy, channel capacity, and the two coding theorems.

  2. 1950

    Hamming codes

    Richard Hamming published the first practical error-correcting codes, proving Shannon's theorem was achievable in practice.

  3. 1951

    Huffman coding

    David Huffman created an optimal prefix-free compression algorithm — the first code to approach Shannon's source coding limit.

  4. 1960

    Reed-Solomon codes

    Powerful error-correcting codes that would later make CDs, DVDs, QR codes, and deep-space communication possible.

  5. 1993

    Turbo codes

    Came within a fraction of a dB of the Shannon limit — a half-century after the theorem was proved. Called "the most exciting development in coding theory in decades."

  6. 2009

    LDPC codes adopted in standards

    Low-density parity-check codes (actually invented by Gallager in 1960, rediscovered in the 1990s) were adopted in Wi-Fi, 5G, and satellite standards — essentially touching the Shannon limit.

Every time you stream a video, make a phone call, scan a QR code, or store data on an SSD, Shannon's theorems are working invisibly beneath the surface. The language models reading this text — GPT, Claude, Gemini — use cross-entropy loss, Shannon's entropy applied to next-token prediction. The entire field of machine learning optimizes information-theoretic quantities that Shannon defined in 1948.

CitationShannon, C. E.. A Mathematical Theory of Communication. Bell System Technical Journal, 1948.

Terms in this paper