Information Theory1948foundational10 min read
A Mathematical Theory of Communication
النظرية الرياضية للاتصال
Shannon, C. E. — Bell System Technical Journal
The problem
Before 1948, engineers treated communication as an analog signal problem — amplify louder, filter better. Nobody had a way to answer the most basic question: how much information can a channel carry? Without that answer, engineers couldn't know whether their systems were close to optimal or wasting most of their capacity. There was no theory of information itself — no unit, no measure, no limits.
The contribution
Shannon defined information mathematically using — a measure of surprise or uncertainty. He introduced the as the fundamental unit, proved that any source can be compressed down to its entropy (), and proved that reliable communication is possible over any noisy channel as long as the transmission rate stays below the (). He gave the formula C = B log₂(1 + S/N) for continuous channels.
The impact
Shannon's paper created an entire field — — and set the mathematical foundations for the digital age. CDs, Wi-Fi, 5G, JPEG, MP3, QR codes, deep-space communication, cryptography, and even machine learning all descend from the ideas in this single 1948 paper. The bit became the atom of the information era.
Imagine a post office that charges by surprise, not by weight. A letter saying "the sun rose this morning" costs almost nothing — everyone expected that. A letter saying "a new continent was discovered" costs a fortune — nobody saw it coming.
Shannon proved that every communication channel is exactly this post office. The entropy of a message source is its average surprise per symbol, measured in bits. And every channel has a maximum surprise-per-second it can deliver — its capacity. Stay under that limit, and you can communicate with essentially zero errors. Exceed it, and no amount of engineering can save you.
The blueprint: how every communication system works
Shannon's first insight was to abstract away the physical details. Telephone, telegraph, television, radio — they all reduce to the same five-block diagram:
Information Source → Transmitter () → Channel (+ ) → Receiver () → Destination
The source produces a message (text, speech, images). The encoder converts it into a signal suited for the channel. The channel carries the signal — but also adds noise. The decoder reconstructs the message from the noisy signal. The destination receives it.
This abstraction was revolutionary: once you have this diagram, you can reason about any communication system mathematically, regardless of whether the channel is a copper wire, a fiber optic cable, or the vacuum of space.
Measuring the unmeasurable: information as surprise
Before Shannon, "information" was a vague word. He gave it a precise definition: information is the resolution of uncertainty. A message that tells you something you already knew carries zero information. A message that tells you something completely unexpected carries maximum information.
The key question is: if a source generates symbols from an alphabet — say letters A through Z — each with some , how much "surprise" does the source produce on average per symbol?
Shannon called this average surprise the entropy of the source, borrowing the term from thermodynamics. A source that always emits the same letter has zero entropy — no surprise, no information. A source where every letter is equally likely has maximum entropy — maximum surprise, maximum information per symbol.
Think of it this way: in English, after "q" the letter "u" almost always follows — so "u" after "q" carries almost no information. But after "t", many letters could follow — so whichever one appears carries real information. Entropy measures this averaged over the entire source.
For the simplest case — a coin flip with probability for heads — entropy simplifies to the binary entropy function: . This beautiful curve peaks at (1 bit — maximum uncertainty) and drops to zero at or (certainty — no surprise).
Source coding: squeezing out the waste
Shannon's first theorem says: you can compress a message down to its entropy, but no further. If a source has entropy H = 2.5 bits/symbol, you need at least 2.5 bits per symbol on average to represent it faithfully — but you can get arbitrarily close to that limit.
The idea is simple: give short codes to common symbols and long codes to rare ones. If "e" appears 13% of the time in English, give it a short code. If "z" appears 0.07%, give it a long code. This is exactly what Morse code did intuitively — "e" is a single dot.
Shannon proved this mathematically: the optimal average code length equals the source entropy. Every compression algorithm since — Huffman coding, LZW (used in GIF), DEFLATE (used in ZIP), and modern algorithms — is chasing this bound.
Channel capacity: the speed limit of communication
Shannon's second great result answers: how fast can you send data through a noisy channel without errors?
Every channel has a number called its capacity — the maximum rate at which you can transmit information with arbitrarily low error probability. Think of it as a highway speed limit: drive under it and traffic flows perfectly; try to exceed it and collisions become inevitable.
For a continuous channel with (Hz) and signal-to-noise ratio , the capacity is given by the famous :
Notice the trade-off: you can compensate for noise by using more bandwidth, or compensate for limited bandwidth by increasing signal power. But there is always a ceiling — no real-world system can exceed Shannon capacity.
The miracle: reliable communication over noisy channels
Here is perhaps the most surprising result in all of information theory — Shannon's noisy channel coding theorem:
As long as your transmission rate R is below the channel capacity C, there exist encoding schemes that make the error probability as small as you want — essentially zero.
Before Shannon, engineers assumed that noise always corrupts some data. Shannon proved that by adding carefully designed — extra bits that don't carry new information but help detect and correct errors — you can communicate reliably through any channel, no matter how noisy, provided .
The catch: Shannon proved these codes exist but didn't construct them. Finding practical codes that approach the Shannon limit became a decades-long quest — from Hamming codes (1950) to Reed-Solomon codes (used in CDs and QR codes) to turbo codes (1993) and LDPC codes (used in 5G and Wi-Fi 6) that nearly touch the limit.
Mutual information: how much does the output tell you about the input?
The channel capacity is defined as the maximum between the channel input and output : Mutual information measures how much knowing the output reduces your uncertainty about the input. In a noiseless channel, — the output tells you everything. In a completely noisy channel, — the output tells you nothing.
Picture two overlapping circles (a Venn diagram): one is (the sender's uncertainty), the other is (the receiver's uncertainty). The overlap is the mutual information — the shared knowledge. Channel capacity asks: how big can we make that overlap?
Redundancy: the hidden waste — and hidden protection — in language
Shannon estimated that English text has about 50% redundancy — roughly half the letters could be removed and you'd still reconstruct the message. This is why you can read "ths sntnc hs n vwls" — the redundancy of English lets your brain fill in gaps.
This redundancy has two sides:
- For compression, it's waste. removes it to approach entropy.
- For error correction, it's a gift. Natural redundancy helps humans detect typos. Engineered redundancy (error-correcting codes) lets machines do the same.
Shannon realized that the fundamental tension in communication is between these two forces: remove redundancy to compress, add redundancy to protect against noise. His theorems tell you exactly how far you can push each direction.
The core ideas in code
Simplified to show the idea — not the real implementation.
import numpy as np
def entropy(probs):
"""Shannon entropy: average surprise per symbol, in bits."""
probs = np.array(probs)
probs = probs[probs > 0] # 0 log 0 = 0 by convention
return -np.sum(probs * np.log2(probs))
# Example: a 4-symbol source
p = [0.5, 0.25, 0.125, 0.125]
H = entropy(p)
print(f"Entropy H = {H:.3f} bits/symbol") # 1.750 — that's the compression limit
# Binary entropy: the special case of a coin flip
def binary_entropy(p):
if p == 0 or p == 1: return 0
return -p * np.log2(p) - (1-p) * np.log2(1-p)
print(f"H(0.5) = {binary_entropy(0.5):.3f}") # 1.0 — max uncertainty
print(f"H(0.9) = {binary_entropy(0.9):.3f}") # 0.469 — biased coin → less surprise
# Shannon-Hartley channel capacity
def channel_capacity(bandwidth_hz, snr_linear):
"""C = B * log2(1 + S/N) in bits per second."""
return bandwidth_hz * np.log2(1 + snr_linear)
# A typical Wi-Fi channel: 20 MHz bandwidth, 30 dB SNR
B = 20e6 # 20 MHz
snr_db = 30
snr = 10 ** (snr_db / 10) # convert dB to linear
C = channel_capacity(B, snr)
print(f"Wi-Fi capacity ≈ {C/1e6:.0f} Mbit/s") # ~199 Mbit/s — theoretical maxWhy it mattered
1948
Shannon's paper
Published in the Bell System Technical Journal. Created information theory as a field and introduced the bit, entropy, channel capacity, and the two coding theorems.
1950
Hamming codes
Richard Hamming published the first practical error-correcting codes, proving Shannon's theorem was achievable in practice.
1951
Huffman coding
David Huffman created an optimal prefix-free compression algorithm — the first code to approach Shannon's source coding limit.
1960
Reed-Solomon codes
Powerful error-correcting codes that would later make CDs, DVDs, QR codes, and deep-space communication possible.
1993
Turbo codes
Came within a fraction of a dB of the Shannon limit — a half-century after the theorem was proved. Called "the most exciting development in coding theory in decades."
2009
LDPC codes adopted in standards
Low-density parity-check codes (actually invented by Gallager in 1960, rediscovered in the 1990s) were adopted in Wi-Fi, 5G, and satellite standards — essentially touching the Shannon limit.
Every time you stream a video, make a phone call, scan a QR code, or store data on an SSD, Shannon's theorems are working invisibly beneath the surface. The language models reading this text — GPT, Claude, Gemini — use cross-entropy loss, Shannon's entropy applied to next-token prediction. The entire field of machine learning optimizes information-theoretic quantities that Shannon defined in 1948.
CitationShannon, C. E.. A Mathematical Theory of Communication. Bell System Technical Journal, 1948.
Terms in this paper
- Entropyالعشوائية الدلالية
- Bitالبِتّ
- Channel Capacityسعة القناة
- Source Coding Theoremمبرهنة ترميز المصدر
- Noisy Channel Coding Theoremمبرهنة ترميز القناة المشوَّشة
- Mutual Informationالمعلومات المتبادلة
- Redundancyالتكرار اللغوي
- Conditional Entropyالعشوائية الشرطية
- Shannon-Hartley Theoremمبرهنة شانون-هارتلي
- Information Theoryنظرية المعلومات