Deep Learning2010intermediate11 min read
Rectified Linear Units Improve Restricted Boltzmann Machines
الوحدات الخطية المُصحَّحة تُحسِّن آلات بولتزمان المقيَّدة
Nair, V. · Hinton, G.E. — ICML
The problem
Restricted Boltzmann Machines (RBMs) — the building blocks of deep belief networks — used binary stochastic hidden units: each unit is either "on" (1) or "off" (0). This binary bottleneck crushed information about relative signal intensities. When you stack multiple layers to form a deep network, the binary discards how strongly each was detected, keeping only whether it was present. Training deep belief networks was slow, features were noisy, and image tasks with large intensity variations (lighting, contrast) were especially hard.
The contribution
Nair & Hinton replaced binary hidden units with Noisy Rectified Linear Units (NReLUs). The derivation is elegant: take a binary unit and make infinite copies with progressively shifted biases — a "" whose total activity is a smoothed integer. Approximate that with max(0, x + N(0,σ²)) and you get a continuous, sparse, fast unit. NReLUs beat binary units on NORB object recognition and LFW face verification, and preserved intensity information across layers — a property called that binary units fundamentally lack.
The impact
This paper launched into the deep learning mainstream. Within two years, AlexNet (2012) used ReLU to train the first deep CNN that crushed ImageNet — and every major architecture since (VGG, ResNet, Transformers, GPT) adopted some variant of rectified activation. ReLU's simplicity (no exp, no division) and (half the units off) became foundational design principles. It also motivated He initialization (2015) and influenced GELU (2016). The most widely used in AI history traces back to this paper.
Imagine a row of water pipes, each with a valve. The old valves () are sticky: they never fully close, always letting a trickle through, and they cap the maximum flow so even a firehose comes out as a gentle stream.
ReLU replaces every sticky valve with a simple flap valve: when water pressure is negative (pushing the wrong way), the flap slams shut — zero flow. When pressure is positive, the flap opens wide and lets exactly as much water through as you push.
The result: pipes that are truly off contribute no noise, pipes that are on preserve the full signal, and you can read the exact water pressure downstream — not a squished version of it.
The problem: binary hidden units crush information
Restricted Boltzmann Machines (RBMs) are the building blocks of deep belief networks. Each has a visible (the input) and a (the learned features), with symmetric connections between them but no connections within the same layer. Training uses : the model reconstructs the input from the hidden features and adjusts weights to minimize the .
The catch: each hidden unit was binary — it could only say "yes, I detect this feature" (1) or "no" (0). If one image of a face is twice as bright as another, both activate the same hidden unit the same way. The intensity information is destroyed. This is fine for simple binary patterns, but catastrophic for real-world images where brightness, contrast, and relative intensities carry meaning.
From binary to stepped to rectified: the derivation
The key insight comes in three steps, like building a staircase:
Step 1 — Binomial units: Take a binary unit and make N copies sharing the same weights and . The total activation is the count of "on" copies — an integer between 0 and N. More copies means more information, but N must be fixed in advance.
Step 2 — Stepped Sigmoid Units: Now take infinitely many copies, but give each copy a progressively more negative bias offset: −0.5, −1.5, −2.5, … The total activity of all copies has a closed form that looks like a smoothed staircase. The learning rules stay identical to binary RBMs — this is mathematically elegant because the decomposes into independent per-copy gradients.
Step 3 — Noisy ReLU (): Computing infinite sigmoid copies is expensive. But the staircase is well approximated by a simple function: — a rectified linear unit with Gaussian noise. The noise makes sampling easy; the rectification (clipping at zero) creates natural sparsity.
The ReLU activation function
Strip away the noise and the RBM context, and the core idea is beautifully simple. Given a total input to a (the weighted sum of inputs plus bias), ReLU computes:
Compare this with the sigmoid function , which squashes everything into the range (0, 1). With sigmoid, a neuron receiving an input of 10 gives an output of 0.99995 — almost identical to an input of 5 (output 0.9933). ReLU outputs 10 and 5 respectively, preserving the ratio.
This matters enormously when you stack layers. Binary or sigmoid representations progressively destroy how strongly each feature fires, keeping only a compressed echo. ReLU representations preserve exact magnitudes, so a feature that fires twice as hard in one image stays twice as hard in every subsequent layer.
Intensity equivariance: why ReLU sees brightness
One of the paper's most elegant observations is intensity equivariance. If you scale all pixel intensities in an image by a factor , then every neuron's input also scales by . For a zero-bias ReLU unit:
- If the unit was "off" (input < 0), scaling by keeps it off.
- If the unit was "on" (input > 0), its output scales by exactly .
So the entire activation pattern scales uniformly — the network's internal representation varies in the same way as the input intensity. This is intensity equivariance: the representation is a faithful, proportional map of the input.
Binary units lack this property entirely. Whether a pixel's intensity is 0.5 or 5.0, a binary unit either fires or doesn't — the ratio is lost. For tasks like face verification, where the same face might appear under very different lighting, intensity equivariance means ReLU features can be compared using the cosine of the angle between activation vectors, which is inherently intensity-invariant.
Sparsity: the gift of silence
When a ReLU unit receives a negative total input, its output is exactly zero — not "close to zero" like a sigmoid, but exactly zero. In a typical trained network, around 50% of ReLU units are off for any given input. This property is called sparsity.
Why is sparsity valuable? Think of it as a room full of experts. When a sigmoid network processes an image, every expert mumbles an opinion (even the irrelevant ones). When a ReLU network processes the same image, half the experts stay silent — only the relevant ones speak. The result is a cleaner, more interpretable, and more efficient representation.
Sparsity also helps with computation: multiplying by zero is free. And it introduces a form of natural — the network can't rely on all features simultaneously, which reduces .
Why ReLU trains faster
ReLU's gradient is wonderfully simple:
- For : the gradient is exactly 1. The signal passes backward through the network without shrinking or growing. No .
- For : the gradient is 0. The unit is off and ignores the update.
Contrast this with sigmoid, whose maximum gradient is 0.25 (at x = 0) and shrinks exponentially as |x| grows. When you chain many sigmoid layers, each one multiplies the gradient by at most 0.25 — after 10 layers, the gradient is smaller by a factor of . This is the vanishing gradient problem.
ReLU doesn't multiply — it either passes the gradient unchanged (1) or blocks it (0). There is no "in between" that slowly strangles the learning signal. This is why ReLU networks can be trained with simple at depths that would cause sigmoid networks to stall.
How RBMs learn with ReLU units
An RBM defines an over visible units and hidden units . Lower energy means higher probability — the model learns by pushing the energy of training examples down and the energy of everything else up.
For binary RBMs, Contrastive Divergence samples binary hidden states given the input, reconstructs the visible layer, then re-samples the hidden layer. The update is the difference between two correlations: the "positive" phase (data-driven) and the "negative" phase (reconstruction-driven).
The beautiful thing is that this learning rule works unchanged for NReLU units. The Stepped Sigmoid derivation guarantees it: since each NReLU is a limit of infinitely many tied binary units, the gradient decomposes identically. You simply replace the binary sampling step with: sample where is the unit's mean total input. The rest of the algorithm stays the same.
The same idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
def relu(x):
"""Rectified Linear Unit: zero if negative, identity if positive."""
return np.maximum(0, x)
def relu_gradient(x):
"""Gradient is 1 where x > 0, 0 elsewhere."""
return (x > 0).astype(float)
def noisy_relu(x, training=True):
"""NReLU as in Nair & Hinton 2010: add Gaussian noise before rectifying."""
if training:
# Noise variance = sigmoid(x), as the unit's "soft" probability
sigma = np.sqrt(sigmoid(x))
return np.maximum(0, x + np.random.normal(0, sigma))
return relu(x) # At test time, just use standard ReLU
def sigmoid(x):
"""The old activation: squashes everything to (0, 1)."""
return 1.0 / (1.0 + np.exp(-x))
# Compare: sigmoid(10) ≈ 0.99995, sigmoid(5) ≈ 0.993 — ratio lost!
# Compare: relu(10) = 10, relu(5) = 5 — ratio preserved.
print(f"sigmoid(10)={sigmoid(10):.5f}, sigmoid(5)={sigmoid(5):.5f}")
print(f"relu(10)={relu(10)}, relu(5)={relu(5)}")Experiments: NORB and Labeled Faces in the Wild
Nair & Hinton tested NReLU on two challenging vision benchmarks:
Jittered-Cluttered NORB — synthetic 3D object recognition with 5 classes of toys photographed in stereo under varying viewpoint, lighting, and clutter. The objects were randomly jittered in position, scale, and brightness. NReLU consistently outperformed binary units: 5.2% lower error without pre-training, 2.2% lower error with pre-training. Remarkably, NReLU without pre-training was better than binary units with pre-training.
Labeled Faces in the Wild (LFW) — real-world face verification: given two face photos, predict "same person" or "different person." Faces vary wildly in lighting, pose, expression, and background. NReLU models showed better accuracy, and crucially used between feature vectors — directly exploiting intensity equivariance to handle lighting variation.
Why it mattered
2010
ReLU — the activation that launched an era
Nair & Hinton showed that noisy rectified linear units outperform binary hidden units in RBMs for object recognition and face verification.
2012
AlexNet — ReLU goes mainstream
Krizhevsky, Sutskever & Hinton used ReLU in a deep CNN that crushed ImageNet by a wide margin. ReLU's fast training was critical — 6× faster than tanh.
2013
Leaky ReLU & PReLU variants
Researchers addressed the "dying ReLU" problem (units permanently stuck at zero) by allowing a small negative slope, keeping the gradient alive.
2015
He initialization — designed for ReLU
He et al. derived the correct weight initialization for ReLU networks, accounting for the fact that half the units are off. This enabled training of very deep networks like ResNet.
2016
GELU — a smooth ReLU for Transformers
Hendrycks & Gimpel proposed the Gaussian Error Linear Unit, a smooth approximation of ReLU that weights inputs by their probability under a Gaussian. GELU became the default in BERT and GPT.
2017
Swish / SiLU
Ramachandran et al. discovered x·sigmoid(x) via automated search — another smooth ReLU variant adopted in EfficientNet and modern architectures.
Every major neural network architecture today uses a rectified activation function — whether standard ReLU, GELU, SiLU, or a variant. The simple idea of "clip negatives to zero" is now as fundamental to deep learning as backpropagation itself. What Nair and Hinton showed was not just a better activation function, but a design principle: simplicity, sparsity, and preserving signal magnitude matter more than biological plausibility.
CitationNair, Hinton. Rectified Linear Units Improve Restricted Boltzmann Machines. ICML, 2010.
Terms in this paper
- ReLUدالة الوحدة الخطية المصححة
- Noisy Rectified Linear Unit (NReLU)الوحدة الخطية المصحَّحة المُضافٌ إليها ضجيج
- Stepped Sigmoid Unitوحدة سيجمويد المُدرَّجة
- Intensity Equivarianceالتكافؤ الشدّي
- Restricted Boltzmann Machine (RBM)آلة بولتزمان المقيَّدة
- Binomial Unitالوحدة ذات التوزيع الحدّين
- Contrastive Divergence (CD)التباعد التبايني
- Deep Belief Networkشبكة الاحتمالات العميقة
- Sparsityالتناثر البنيوي للمصفوفات