Deep Learning2010intermediate11 min read

Rectified Linear Units Improve Restricted Boltzmann Machines

الوحدات الخطية المُصحَّحة تُحسِّن آلات بولتزمان المقيَّدة

Nair, V. · Hinton, G.E. — ICML

The problem

Restricted Boltzmann Machines (RBMs) — the building blocks of deep belief networks — used binary stochastic hidden units: each unit is either "on" (1) or "off" (0). This binary bottleneck crushed information about relative signal intensities. When you stack multiple layers to form a deep network, the binary discards how strongly each was detected, keeping only whether it was present. Training deep belief networks was slow, features were noisy, and image tasks with large intensity variations (lighting, contrast) were especially hard.

The contribution

Nair & Hinton replaced binary hidden units with Noisy Rectified Linear Units (NReLUs). The derivation is elegant: take a binary unit and make infinite copies with progressively shifted biases — a "" whose total activity is a smoothed integer. Approximate that with max(0, x + N(0,σ²)) and you get a continuous, sparse, fast unit. NReLUs beat binary units on NORB object recognition and LFW face verification, and preserved intensity information across layers — a property called that binary units fundamentally lack.

The impact

This paper launched into the deep learning mainstream. Within two years, AlexNet (2012) used ReLU to train the first deep CNN that crushed ImageNet — and every major architecture since (VGG, ResNet, Transformers, GPT) adopted some variant of rectified activation. ReLU's simplicity (no exp, no division) and (half the units off) became foundational design principles. It also motivated He initialization (2015) and influenced GELU (2016). The most widely used in AI history traces back to this paper.

Imagine a row of water pipes, each with a valve. The old valves () are sticky: they never fully close, always letting a trickle through, and they cap the maximum flow so even a firehose comes out as a gentle stream.

ReLU replaces every sticky valve with a simple flap valve: when water pressure is negative (pushing the wrong way), the flap slams shut — zero flow. When pressure is positive, the flap opens wide and lets exactly as much water through as you push.

The result: pipes that are truly off contribute no noise, pipes that are on preserve the full signal, and you can read the exact water pressure downstream — not a squished version of it.

The problem: binary hidden units crush information

Restricted Boltzmann Machines (RBMs) are the building blocks of deep belief networks. Each has a visible (the input) and a (the learned features), with symmetric connections between them but no connections within the same layer. Training uses : the model reconstructs the input from the hidden features and adjusts weights to minimize the .

The catch: each hidden unit was binary — it could only say "yes, I detect this feature" (1) or "no" (0). If one image of a face is twice as bright as another, both activate the same hidden unit the same way. The intensity information is destroyed. This is fine for simple binary patterns, but catastrophic for real-world images where brightness, contrast, and relative intensities carry meaning.

Open in Lab
Drag the input slider to see how sigmoid squashes everything to [0,1], while ReLU preserves the full positive range.
The demo wakes as you arrive…

From binary to stepped to rectified: the derivation

The key insight comes in three steps, like building a staircase:

Step 1 — Binomial units: Take a binary unit and make N copies sharing the same weights and . The total activation is the count of "on" copies — an integer between 0 and N. More copies means more information, but N must be fixed in advance.

Step 2 — Stepped Sigmoid Units: Now take infinitely many copies, but give each copy a progressively more negative bias offset: −0.5, −1.5, −2.5, … The total activity of all copies has a closed form that looks like a smoothed staircase. The learning rules stay identical to binary RBMs — this is mathematically elegant because the decomposes into independent per-copy gradients.

Step 3 — Noisy ReLU (): Computing infinite sigmoid copies is expensive. But the staircase is well approximated by a simple function: max⁡(0,  x+N(0,σ2))\max(0,\; x + \mathcal{N}(0, \sigma^2)) — a rectified linear unit with Gaussian noise. The noise makes sampling easy; the rectification (clipping at zero) creates natural sparsity.

Open in Lab
Watch the three-step derivation: binary copies → stepped sigmoid staircase → smooth ReLU approximation.
The demo wakes as you arrive…

The ReLU activation function

Strip away the noise and the RBM context, and the core idea is beautifully simple. Given a total input xx to a (the weighted sum of inputs plus bias), ReLU computes:

f(x)=max⁡(0,  x)f(x) = \max(0,\; x)
The ReLU activation function — If x is negative, output zero (the unit is "off"). If x is positive, output x unchanged (the unit is "on" and preserves the signal magnitude). No exponents, no divisions — just a comparison and a pass-through.

Compare this with the sigmoid function σ(x)=1/(1+e−x)\sigma(x) = 1 / (1 + e^{-x}), which squashes everything into the range (0, 1). With sigmoid, a neuron receiving an input of 10 gives an output of 0.99995 — almost identical to an input of 5 (output 0.9933). ReLU outputs 10 and 5 respectively, preserving the ratio.

This matters enormously when you stack layers. Binary or sigmoid representations progressively destroy how strongly each feature fires, keeping only a compressed echo. ReLU representations preserve exact magnitudes, so a feature that fires twice as hard in one image stays twice as hard in every subsequent layer.

Intensity equivariance: why ReLU sees brightness

One of the paper's most elegant observations is intensity equivariance. If you scale all pixel intensities in an image by a factor α>0\alpha > 0, then every neuron's input also scales by α\alpha. For a zero-bias ReLU unit:

  • If the unit was "off" (input < 0), scaling by α\alpha keeps it off.
  • If the unit was "on" (input > 0), its output scales by exactly α\alpha.

So the entire activation pattern scales uniformly — the network's internal representation varies in the same way as the input intensity. This is intensity equivariance: the representation is a faithful, proportional map of the input.

Binary units lack this property entirely. Whether a pixel's intensity is 0.5 or 5.0, a binary unit either fires or doesn't — the ratio is lost. For tasks like face verification, where the same face might appear under very different lighting, intensity equivariance means ReLU features can be compared using the cosine of the angle between activation vectors, which is inherently intensity-invariant.

Open in Lab
Scale image brightness and watch: ReLU activations scale proportionally while sigmoid activations saturate and lose the ratio.
The demo wakes as you arrive…

Sparsity: the gift of silence

When a ReLU unit receives a negative total input, its output is exactly zero — not "close to zero" like a sigmoid, but exactly zero. In a typical trained network, around 50% of ReLU units are off for any given input. This property is called sparsity.

Why is sparsity valuable? Think of it as a room full of experts. When a sigmoid network processes an image, every expert mumbles an opinion (even the irrelevant ones). When a ReLU network processes the same image, half the experts stay silent — only the relevant ones speak. The result is a cleaner, more interpretable, and more efficient representation.

Sparsity also helps with computation: multiplying by zero is free. And it introduces a form of natural — the network can't rely on all features simultaneously, which reduces .

Open in Lab
Toggle between sigmoid and ReLU activations on sample inputs. Count the truly "off" units — sigmoid never reaches zero.
The demo wakes as you arrive…

Why ReLU trains faster

ReLU's gradient is wonderfully simple:

  • For x>0x > 0: the gradient is exactly 1. The signal passes backward through the network without shrinking or growing. No .
  • For x≤0x \leq 0: the gradient is 0. The unit is off and ignores the update.

Contrast this with sigmoid, whose maximum gradient is 0.25 (at x = 0) and shrinks exponentially as |x| grows. When you chain many sigmoid layers, each one multiplies the gradient by at most 0.25 — after 10 layers, the gradient is smaller by a factor of 0.2510≈10−60.25^{10} \approx 10^{-6}. This is the vanishing gradient problem.

ReLU doesn't multiply — it either passes the gradient unchanged (1) or blocks it (0). There is no "in between" that slowly strangles the learning signal. This is why ReLU networks can be trained with simple at depths that would cause sigmoid networks to stall.

f′(x)={1if x>00if x≤0f'(x) = \begin{cases} 1 & \text{if } x > 0 \\ 0 & \text{if } x \leq 0 \end{cases}
Gradient of ReLU — A constant gradient of 1 for positive inputs means no signal decay through layers. A gradient of 0 for negative inputs creates sparsity in the backward pass too.

How RBMs learn with ReLU units

An RBM defines an over visible units v\mathbf{v} and hidden units h\mathbf{h}. Lower energy means higher probability — the model learns by pushing the energy of training examples down and the energy of everything else up.

For binary RBMs, Contrastive Divergence samples binary hidden states given the input, reconstructs the visible layer, then re-samples the hidden layer. The update is the difference between two correlations: the "positive" phase (data-driven) and the "negative" phase (reconstruction-driven).

The beautiful thing is that this learning rule works unchanged for NReLU units. The Stepped Sigmoid derivation guarantees it: since each NReLU is a limit of infinitely many tied binary units, the gradient decomposes identically. You simply replace the binary sampling step with: sample hj=max⁡(0,μj+N(0,σj2))h_j = \max(0, \mu_j + \mathcal{N}(0, \sigma_j^2)) where μj\mu_j is the unit's mean total input. The rest of the algorithm stays the same.

Open in Lab
Step through one round of Contrastive Divergence: data → hidden → reconstruction → weight update. Toggle between binary and NReLU units.
The demo wakes as you arrive…

The same idea in code

ReLU and Noisy ReLU — the full activation functionpython

Simplified to show the idea — not the real implementation.

import numpy as np

def relu(x):
    """Rectified Linear Unit: zero if negative, identity if positive."""
    return np.maximum(0, x)

def relu_gradient(x):
    """Gradient is 1 where x > 0, 0 elsewhere."""
    return (x > 0).astype(float)

def noisy_relu(x, training=True):
    """NReLU as in Nair & Hinton 2010: add Gaussian noise before rectifying."""
    if training:
        # Noise variance = sigmoid(x), as the unit's "soft" probability
        sigma = np.sqrt(sigmoid(x))
        return np.maximum(0, x + np.random.normal(0, sigma))
    return relu(x)  # At test time, just use standard ReLU

def sigmoid(x):
    """The old activation: squashes everything to (0, 1)."""
    return 1.0 / (1.0 + np.exp(-x))

# Compare: sigmoid(10) ≈ 0.99995, sigmoid(5) ≈ 0.993 — ratio lost!
# Compare: relu(10) = 10, relu(5) = 5 — ratio preserved.
print(f"sigmoid(10)={sigmoid(10):.5f}, sigmoid(5)={sigmoid(5):.5f}")
print(f"relu(10)={relu(10)}, relu(5)={relu(5)}")

Experiments: NORB and Labeled Faces in the Wild

Nair & Hinton tested NReLU on two challenging vision benchmarks:

Jittered-Cluttered NORB — synthetic 3D object recognition with 5 classes of toys photographed in stereo under varying viewpoint, lighting, and clutter. The objects were randomly jittered in position, scale, and brightness. NReLU consistently outperformed binary units: 5.2% lower error without pre-training, 2.2% lower error with pre-training. Remarkably, NReLU without pre-training was better than binary units with pre-training.

Labeled Faces in the Wild (LFW) — real-world face verification: given two face photos, predict "same person" or "different person." Faces vary wildly in lighting, pose, expression, and background. NReLU models showed better accuracy, and crucially used between feature vectors — directly exploiting intensity equivariance to handle lighting variation.

Open in Lab
Compare error rates: binary units vs NReLU, with and without pre-training, on NORB.
The demo wakes as you arrive…

Why it mattered

  1. 2010

    ReLU — the activation that launched an era

    Nair & Hinton showed that noisy rectified linear units outperform binary hidden units in RBMs for object recognition and face verification.

  2. 2012

    AlexNet — ReLU goes mainstream

    Krizhevsky, Sutskever & Hinton used ReLU in a deep CNN that crushed ImageNet by a wide margin. ReLU's fast training was critical — 6× faster than tanh.

  3. 2013

    Leaky ReLU & PReLU variants

    Researchers addressed the "dying ReLU" problem (units permanently stuck at zero) by allowing a small negative slope, keeping the gradient alive.

  4. 2015

    He initialization — designed for ReLU

    He et al. derived the correct weight initialization for ReLU networks, accounting for the fact that half the units are off. This enabled training of very deep networks like ResNet.

  5. 2016

    GELU — a smooth ReLU for Transformers

    Hendrycks & Gimpel proposed the Gaussian Error Linear Unit, a smooth approximation of ReLU that weights inputs by their probability under a Gaussian. GELU became the default in BERT and GPT.

  6. 2017

    Swish / SiLU

    Ramachandran et al. discovered x·sigmoid(x) via automated search — another smooth ReLU variant adopted in EfficientNet and modern architectures.

Every major neural network architecture today uses a rectified activation function — whether standard ReLU, GELU, SiLU, or a variant. The simple idea of "clip negatives to zero" is now as fundamental to deep learning as backpropagation itself. What Nair and Hinton showed was not just a better activation function, but a design principle: simplicity, sparsity, and preserving signal magnitude matter more than biological plausibility.

CitationNair, Hinton. Rectified Linear Units Improve Restricted Boltzmann Machines. ICML, 2010.

Terms in this paper