Neural Networks1986foundational8 min read
Learning Representations by Back-Propagating Errors
تعلُّم التمثيلات عبر الانتشار العكسي للأخطاء
Rumelhart, D. E. · Hinton, G. E. · Williams, R. J. — Nature
The problem
Single-layer perceptrons can only learn linearly separable patterns — they famously fail on XOR. Everyone knew that adding hidden layers would solve this, but nobody had a practical algorithm to train them. How do you tell a hidden unit, which has no direct contact with the desired output, how to change its weights?
The contribution
A practical, general-purpose learning algorithm for multi-layer networks. By applying the of calculus systematically, the error at the output is decomposed into a for every in every layer — even hidden layers with no direct . Each weight learns exactly its share of the blame, and updates them all. The hidden layers, without being told what to represent, discover useful internal features automatically.
The impact
made possible. Without it, neural networks were limited to one layer and could not learn complex patterns. It is the foundational algorithm behind every modern — from the ConvNets that won ImageNet to the Transformers that power GPT and Claude. It turned neural networks from a theoretical curiosity into a practical engineering tool.
Imagine a factory assembly line with three stations. A defective product comes out at the end. The factory inspector doesn't just yell at the last station — she traces the defect backward: the paint was fine (station 3), but the part was bent (station 2) because the raw material was cut wrong (station 1).
That's backpropagation. The "defect" is the error between the network's prediction and the truth. The chain rule traces it backward through every layer, and each weight learns exactly how much of the defect it caused — then adjusts itself to cause less next time.
The problem: hidden layers couldn't learn
By the mid-1980s, neural networks were stuck. Single-layer perceptrons could only solve problems where a straight line separates the classes — and Minsky & Papert had proved in 1969 that this excluded even simple functions like XOR. The fix was obvious: add hidden layers that could learn internal features. The obstacle was equally obvious: how do you compute a learning signal for a hidden unit that never sees the desired answer?
-
For the , the error is simply ŷ − y. You can compute ∂L/∂w directly.
-
For a , there is no target. The hidden unit's "error" must be inferred from how it affected the output — but that influence passes through every subsequent layer and nonlinearity.
Rumelhart, Hinton, and Williams showed that the chain rule of calculus solves this exactly: decompose the output error into a product of local derivatives, one per layer, and multiply them backward. Each weight gets a gradient that tells it how to change.
The idea: trace the error backward with the chain rule
The algorithm has two phases that alternate on every training example:
— data flows from input through hidden layers to output, computing a prediction ŷ and a loss L = ½(ŷ − y)².
— the gradient of the loss flows in reverse, layer by layer. At each layer, the chain rule says: multiply the incoming gradient by the local derivative of that layer's operation. The result is ∂L/∂w for every weight — the direction that reduces the error.
Then update: w ← w − η · ∂L/∂w, where η is the — a small step size.
Think of it as a relay of blame. The loss says: "the prediction was off by this much." The output layer says: "that's because the hidden layer gave me these values." The hidden layer says: "that's because the input weights gave me those values." Each layer passes its share of blame backward — and the multiplication of local derivatives is the chain rule.
Gradient descent: following the slope downhill
Once backpropagation gives you ∂L/∂w for every weight, the question is: how do you use it? The gradient points in the direction of steepest ascent of the loss. To minimize the loss, move in the opposite direction:
w ← w − η · ∂L/∂w
η is the . Too large and you overshoot the minimum — the loss bounces around. Too small and you crawl toward it, potentially getting stuck in a . The 1986 paper also introduced — a running average of past gradients — to accelerate and smooth out oscillations.
The breakthrough: hidden layers discover features
The deepest insight in the paper is not the algorithm itself — it's what the algorithm produces. No one tells the hidden units what features to detect. Yet after training, they develop that make the problem solvable.
In XOR, the hidden layer warps the input space so that points with the same label cluster together — and a straight line can now separate them. The network invented a coordinate system that makes the problem easy. This is : the hidden layers don't just transform data, they discover the right way to look at it.
This is the foundation of all deep learning. Every filter, every attention head, every word — they exist because backpropagation makes it possible for hidden layers to discover useful features without being told what to look for.
Putting it all together
The full backpropagation training loop, as described in the 1986 paper, stacks these pieces: an , one or more hidden layers with activation, an output layer, and a squared-error loss. The forward pass computes predictions; the backward pass computes gradients; gradient descent updates weights.
Click any component below to see what it does:
The ceiling: vanishing gradients in deep networks
The chain rule is a product — and products can shrink. The maximum of sigmoid'(z) is 0.25, so each layer multiplies the backward gradient by at most 0.25. Through 10 layers, the gradient reaching the first layer is at most 0.25¹⁰ ≈ 0.000001 — essentially zero. The early layers stop learning. This is the problem, which Bengio et al. would formally analyze in 1994.
The fix would take decades: activation (gradient = 1 when active), residual connections (skip the multiplication), , and careful weight initialization (Xavier, He). But the algorithm itself — backpropagation — remained unchanged. Every fix works with the chain rule, not around it.
The same idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
def sigmoid(z):
return 1 / (1 + np.exp(-z))
def sigmoid_deriv(z):
s = sigmoid(z)
return s * (1 - s) # max = 0.25 — this is why gradients vanish
# ── Forward pass ──────────────────────────────────────────────────
def forward(X, W1, b1, W2, b2):
z1 = X @ W1 + b1 # hidden pre-activation
h = sigmoid(z1) # hidden activation
z2 = h @ W2 + b2 # output pre-activation
y_hat = sigmoid(z2) # prediction
return z1, h, z2, y_hat
# ── Backward pass (the chain rule, layer by layer) ────────────────
def backward(X, y, z1, h, z2, y_hat, W2):
m = X.shape[0]
dL_dyhat = (y_hat - y) # ∂L/∂ŷ
dL_dz2 = dL_dyhat * sigmoid_deriv(z2) # ∂L/∂z₂
dL_dW2 = h.T @ dL_dz2 / m # ∂L/∂W₂
dL_db2 = dL_dz2.mean(axis=0) # ∂L/∂b₂
dL_dh = dL_dz2 @ W2.T # chain through W₂
dL_dz1 = dL_dh * sigmoid_deriv(z1) # chain through σ
dL_dW1 = X.T @ dL_dz1 / m # ∂L/∂W₁
dL_db1 = dL_dz1.mean(axis=0) # ∂L/∂b₁
return dL_dW1, dL_db1, dL_dW2, dL_db2
# ── Training loop ────────────────────────────────────────────────
# XOR dataset
X = np.array([[0,0],[0,1],[1,0],[1,1]])
y = np.array([[0],[1],[1],[0]])
np.random.seed(42)
W1 = np.random.randn(2, 4) * 0.5
b1 = np.zeros(4)
W2 = np.random.randn(4, 1) * 0.5
b2 = np.zeros(1)
lr = 2.0 # learning rate
for epoch in range(5000):
z1, h, z2, y_hat = forward(X, W1, b1, W2, b2)
dW1, db1, dW2, db2 = backward(X, y, z1, h, z2, y_hat, W2)
W1 -= lr * dW1; b1 -= lr * db1 # gradient descent
W2 -= lr * dW2; b2 -= lr * db2
print(np.round(y_hat, 2)) # [[0.02], [0.98], [0.98], [0.02]] — XOR solved!Why it mattered
1989
LeCun — Backprop meets ConvNets
LeCun applied backpropagation to convolutional networks for handwritten digit recognition — the first practical demonstration that backprop could train non-trivial architectures end-to-end.
1994
Bengio — Vanishing gradients formalized
Bengio, Simard & Frasconi showed that sigmoid-based backpropagation suffers exponential gradient decay in deep networks, explaining why deep training had failed for a decade.
1997
Hochreiter & Schmidhuber — LSTM
Long Short-Term Memory introduced an additive cell state whose gradient derivative is exactly 1.0, allowing backpropagation to flow unchanged across hundreds of time steps.
2006
Hinton — Deep belief nets revive depth
Hinton, Osindero & Teh showed that deep networks could be pre-trained layer-by-layer with restricted Boltzmann machines, then fine-tuned with backpropagation — reigniting the field's interest in depth.
2010
Xavier init & ReLU — Vanishing gradients solved at the source
Glorot & Bengio's Xavier initialization and Nair & Hinton's ReLU activation (derivative = 1 when active) together eliminated the vanishing gradient problem, making deep training practical without pre-training.
2012
AlexNet — Backprop wins ImageNet
AlexNet won ImageNet by a 10-point margin — a deep ConvNet trained end-to-end with backpropagation on GPUs. The moment deep learning became the dominant paradigm.
2017
The Transformer — backprop trains self-attention
The Transformer replaced recurrence with self-attention, but its training engine is still backpropagation. Every attention weight is learned via the chain rule — exactly as the 1986 paper described.
2026
Every network alive
GPT, Claude, Gemini, AlphaFold, Whisper — every neural network trains with backpropagation. The algorithm from 1986 runs trillions of times a day, unchanged in its mathematical core.
The 1986 paper didn't just introduce an algorithm. It proved that the right learning procedure, applied to the right architecture, produces representations that nobody designed — representations that turned out to be more powerful than anything humans could have engineered by hand.
CitationRumelhart, Hinton, Williams. Learning Representations by Back-Propagating Errors. Nature, 1986.
Terms in this paper
- Backpropagationالتحديث التراجعي
- Chain ruleقاعدة السلسلة
- Forward Passالتمرير الأمامي
- Backward Passالتمرير الخلفي
- Error signalإشارة الخطأ الرياضية
- Vanishing Gradientاضمحلال متجهات الميل
- Internal representationsالتمثيلات الداخلية
- Gradient Descentالانحدار التدريجي