Neural Networks1986foundational8 min read

Learning Representations by Back-Propagating Errors

تعلُّم التمثيلات عبر الانتشار العكسي للأخطاء

Rumelhart, D. E. · Hinton, G. E. · Williams, R. J. — Nature

The problem

Single-layer perceptrons can only learn linearly separable patterns — they famously fail on XOR. Everyone knew that adding hidden layers would solve this, but nobody had a practical algorithm to train them. How do you tell a hidden unit, which has no direct contact with the desired output, how to change its weights?

The contribution

A practical, general-purpose learning algorithm for multi-layer networks. By applying the of calculus systematically, the error at the output is decomposed into a for every in every layer — even hidden layers with no direct . Each weight learns exactly its share of the blame, and updates them all. The hidden layers, without being told what to represent, discover useful internal features automatically.

The impact

made possible. Without it, neural networks were limited to one layer and could not learn complex patterns. It is the foundational algorithm behind every modern — from the ConvNets that won ImageNet to the Transformers that power GPT and Claude. It turned neural networks from a theoretical curiosity into a practical engineering tool.

Imagine a factory assembly line with three stations. A defective product comes out at the end. The factory inspector doesn't just yell at the last station — she traces the defect backward: the paint was fine (station 3), but the part was bent (station 2) because the raw material was cut wrong (station 1).

That's backpropagation. The "defect" is the error between the network's prediction and the truth. The chain rule traces it backward through every layer, and each weight learns exactly how much of the defect it caused — then adjusts itself to cause less next time.

The problem: hidden layers couldn't learn

By the mid-1980s, neural networks were stuck. Single-layer perceptrons could only solve problems where a straight line separates the classes — and Minsky & Papert had proved in 1969 that this excluded even simple functions like XOR. The fix was obvious: add hidden layers that could learn internal features. The obstacle was equally obvious: how do you compute a learning signal for a hidden unit that never sees the desired answer?

  • For the , the error is simply ŷ − y. You can compute ∂L/∂w directly.

  • For a , there is no target. The hidden unit's "error" must be inferred from how it affected the output — but that influence passes through every subsequent layer and nonlinearity.

Rumelhart, Hinton, and Williams showed that the chain rule of calculus solves this exactly: decompose the output error into a product of local derivatives, one per layer, and multiply them backward. Each weight gets a gradient that tells it how to change.

Open in Lab
Toggle between a single perceptron and a network with a hidden layer. No line can solve XOR alone.
The demo wakes as you arrive…

The idea: trace the error backward with the chain rule

The algorithm has two phases that alternate on every training example:

— data flows from input through hidden layers to output, computing a prediction ŷ and a loss L = ½(ŷ − y)².

— the gradient of the loss flows in reverse, layer by layer. At each layer, the chain rule says: multiply the incoming gradient by the local derivative of that layer's operation. The result is ∂L/∂w for every weight — the direction that reduces the error.

Then update: w ← w − η · ∂L/∂w, where η is the — a small step size.

Open in Lab
Press Play to watch data flow forward, then error flow backward through the network.
The demo wakes as you arrive…
∂L∂w1=∂L∂y^⋅∂y^∂h⋅∂h∂w1\frac{\partial L}{\partial w_1} = \frac{\partial L}{\partial \hat{y}} \cdot \frac{\partial \hat{y}}{\partial h} \cdot \frac{\partial h}{\partial w_1}
The chain rule — backpropagation's engine — Each factor is a local derivative: how much does the loss change per unit change in ŷ? How much does ŷ change per unit change in h? How much does h change per unit change in w₁? Multiply them and you know how the loss changes with w₁ — even though w₁ never sees the loss directly.

Think of it as a relay of blame. The loss says: "the prediction was off by this much." The output layer says: "that's because the hidden layer gave me these values." The hidden layer says: "that's because the input weights gave me those values." Each layer passes its share of blame backward — and the multiplication of local derivatives is the chain rule.

Open in Lab
Step through forward and backward passes on a tiny network. Watch the chain of local derivatives multiply.
The demo wakes as you arrive…

Gradient descent: following the slope downhill

Once backpropagation gives you ∂L/∂w for every weight, the question is: how do you use it? The gradient points in the direction of steepest ascent of the loss. To minimize the loss, move in the opposite direction:

w ← w − η · ∂L/∂w

η is the . Too large and you overshoot the minimum — the loss bounces around. Too small and you crawl toward it, potentially getting stuck in a . The 1986 paper also introduced — a running average of past gradients — to accelerate and smooth out oscillations.

Open in Lab
Adjust the learning rate and press Go. Watch the ball descend — or overshoot if η is too high.
The demo wakes as you arrive…

The breakthrough: hidden layers discover features

The deepest insight in the paper is not the algorithm itself — it's what the algorithm produces. No one tells the hidden units what features to detect. Yet after training, they develop that make the problem solvable.

In XOR, the hidden layer warps the input space so that points with the same label cluster together — and a straight line can now separate them. The network invented a coordinate system that makes the problem easy. This is : the hidden layers don't just transform data, they discover the right way to look at it.

This is the foundation of all deep learning. Every filter, every attention head, every word — they exist because backpropagation makes it possible for hidden layers to discover useful features without being told what to look for.

Open in Lab
Left: XOR in input space — no line separates the classes. Right: the same 4 points after the hidden layer — now linearly separable.
The demo wakes as you arrive…

Putting it all together

The full backpropagation training loop, as described in the 1986 paper, stacks these pieces: an , one or more hidden layers with activation, an output layer, and a squared-error loss. The forward pass computes predictions; the backward pass computes gradients; gradient descent updates weights.

Click any component below to see what it does:

Open in Lab
The demo wakes as you arrive…

The ceiling: vanishing gradients in deep networks

The chain rule is a product — and products can shrink. The maximum of sigmoid'(z) is 0.25, so each layer multiplies the backward gradient by at most 0.25. Through 10 layers, the gradient reaching the first layer is at most 0.25¹⁰ ≈ 0.000001 — essentially zero. The early layers stop learning. This is the problem, which Bengio et al. would formally analyze in 1994.

The fix would take decades: activation (gradient = 1 when active), residual connections (skip the multiplication), , and careful weight initialization (Xavier, He). But the algorithm itself — backpropagation — remained unchanged. Every fix works with the chain rule, not around it.

Open in Lab
Add layers and watch early-layer gradients shrink toward zero. This is why deep sigmoid networks couldn't learn.
The demo wakes as you arrive…

The same idea in code

Backpropagation for a 2-layer network, completepython

Simplified to show the idea — not the real implementation.

import numpy as np

def sigmoid(z):
    return 1 / (1 + np.exp(-z))

def sigmoid_deriv(z):
    s = sigmoid(z)
    return s * (1 - s)          # max = 0.25 — this is why gradients vanish

# ── Forward pass ──────────────────────────────────────────────────
def forward(X, W1, b1, W2, b2):
    z1 = X @ W1 + b1            # hidden pre-activation
    h  = sigmoid(z1)            # hidden activation
    z2 = h @ W2 + b2            # output pre-activation
    y_hat = sigmoid(z2)         # prediction
    return z1, h, z2, y_hat

# ── Backward pass (the chain rule, layer by layer) ────────────────
def backward(X, y, z1, h, z2, y_hat, W2):
    m = X.shape[0]
    dL_dyhat = (y_hat - y)                         # ∂L/∂ŷ
    dL_dz2   = dL_dyhat * sigmoid_deriv(z2)        # ∂L/∂z₂
    dL_dW2   = h.T @ dL_dz2 / m                   # ∂L/∂W₂
    dL_db2   = dL_dz2.mean(axis=0)                 # ∂L/∂b₂

    dL_dh    = dL_dz2 @ W2.T                       # chain through W₂
    dL_dz1   = dL_dh * sigmoid_deriv(z1)           # chain through σ
    dL_dW1   = X.T @ dL_dz1 / m                   # ∂L/∂W₁
    dL_db1   = dL_dz1.mean(axis=0)                 # ∂L/∂b₁
    return dL_dW1, dL_db1, dL_dW2, dL_db2

# ── Training loop ────────────────────────────────────────────────
# XOR dataset
X = np.array([[0,0],[0,1],[1,0],[1,1]])
y = np.array([[0],[1],[1],[0]])

np.random.seed(42)
W1 = np.random.randn(2, 4) * 0.5
b1 = np.zeros(4)
W2 = np.random.randn(4, 1) * 0.5
b2 = np.zeros(1)
lr = 2.0   # learning rate

for epoch in range(5000):
    z1, h, z2, y_hat = forward(X, W1, b1, W2, b2)
    dW1, db1, dW2, db2 = backward(X, y, z1, h, z2, y_hat, W2)
    W1 -= lr * dW1;  b1 -= lr * db1   # gradient descent
    W2 -= lr * dW2;  b2 -= lr * db2

print(np.round(y_hat, 2))   # [[0.02], [0.98], [0.98], [0.02]] — XOR solved!

Why it mattered

  1. 1989

    LeCun — Backprop meets ConvNets

    LeCun applied backpropagation to convolutional networks for handwritten digit recognition — the first practical demonstration that backprop could train non-trivial architectures end-to-end.

  2. 1994

    Bengio — Vanishing gradients formalized

    Bengio, Simard & Frasconi showed that sigmoid-based backpropagation suffers exponential gradient decay in deep networks, explaining why deep training had failed for a decade.

  3. 1997

    Hochreiter & Schmidhuber — LSTM

    Long Short-Term Memory introduced an additive cell state whose gradient derivative is exactly 1.0, allowing backpropagation to flow unchanged across hundreds of time steps.

  4. 2006

    Hinton — Deep belief nets revive depth

    Hinton, Osindero & Teh showed that deep networks could be pre-trained layer-by-layer with restricted Boltzmann machines, then fine-tuned with backpropagation — reigniting the field's interest in depth.

  5. 2010

    Xavier init & ReLU — Vanishing gradients solved at the source

    Glorot & Bengio's Xavier initialization and Nair & Hinton's ReLU activation (derivative = 1 when active) together eliminated the vanishing gradient problem, making deep training practical without pre-training.

  6. 2012

    AlexNet — Backprop wins ImageNet

    AlexNet won ImageNet by a 10-point margin — a deep ConvNet trained end-to-end with backpropagation on GPUs. The moment deep learning became the dominant paradigm.

  7. 2017

    The Transformer — backprop trains self-attention

    The Transformer replaced recurrence with self-attention, but its training engine is still backpropagation. Every attention weight is learned via the chain rule — exactly as the 1986 paper described.

  8. 2026

    Every network alive

    GPT, Claude, Gemini, AlphaFold, Whisper — every neural network trains with backpropagation. The algorithm from 1986 runs trillions of times a day, unchanged in its mathematical core.

The 1986 paper didn't just introduce an algorithm. It proved that the right learning procedure, applied to the right architecture, produces representations that nobody designed — representations that turned out to be more powerful than anything humans could have engineered by hand.

CitationRumelhart, Hinton, Williams. Learning Representations by Back-Propagating Errors. Nature, 1986.

Terms in this paper