Deep Learning2016beginner10 min read

Gaussian Error Linear Units (GELUs)

وحدات الخطأ الغاوسي الخطية (GELUs)

Hendrycks, D. · Gimpel, K. — arXiv

The problem

Activation functions and stochastic regularizers (like ) have always been treated as separate design choices. makes a hard binary decision — keep or kill — based only on the sign of the input, producing a non-smooth kink at zero. This hard gating ignores how confident the network is about an input's importance. Meanwhile, dropout randomly zeroes neurons without considering their value at all. No unified these two ideas into a single, smooth, probabilistically motivated operation.

The contribution

The activation function: GELU(x) = x · Φ(x), where Φ(x) is the standard Gaussian . Instead of gating by sign (like ReLU) or gating randomly (like dropout), GELU weights each input by how likely it is to be greater than other inputs under a Gaussian assumption. This produces a smooth, non-monotonic curve that acts as the expected value of a stochastic regularizer — unifying nonlinearity and in one formula. GELU outperformed ReLU and ELU across computer vision, NLP, and speech tasks.

The impact

GELU became the default activation function in Transformers. BERT, GPT-2, GPT-3, and nearly every major language adopted GELU, making it one of the most widely deployed activation functions in modern . The paper also introduced the Linear Unit (SiLU), later independently rediscovered as "Swish" by Google Brain. GELU demonstrated that principled probabilistic thinking about activations leads to practical performance gains.

Think of a nightclub bouncer. ReLU is a bouncer who checks your ID: positive? You're in. Negative? Absolutely not. There's no middle ground — even someone at −0.001 is turned away cold, while +0.001 walks right through.

GELU is a smarter bouncer with a bell-curve intuition. Instead of a hard cutoff, this bouncer asks: "How likely is it that you belong here?" Very positive inputs stroll in with full confidence. Very negative inputs are turned away. But the borderline cases — the values near zero — get a proportional welcome: the bouncer lets through a fraction of you, scaled by how plausible your presence is. No sudden cliff, just a smooth judgment call.

The problem: hard gates and disconnected regularizers

By 2016, the ReLU activation had become the workhorse of . Its formula is disarmingly simple: ReLU(x)=max⁡(0,x)\text{ReLU}(x) = \max(0, x). If the input is positive, pass it through unchanged. If negative, output zero. This simplicity made it fast and effective — far better than the sigmoid functions that came before, which suffered from vanishing gradients in deep networks.

But ReLU has a sharp edge. At x=0x = 0, the function makes a sudden, discontinuous jump in its — from 0 to 1. This hard kink means the function is not smooth: its flips instantaneously. For values just below zero, the is completely dead. For values just above, it is fully alive. There is no nuance near the boundary.

Meanwhile, stochastic regularizers like dropout were applied separately. Dropout randomly zeroes neurons during to prevent — but it does so blindly, without considering whether a neuron's signal is large or small. The activation decision and the regularization decision were two independent mechanisms, never talking to each other.

Open in Lab
Drag the input value and compare how ReLU and GELU respond — notice the smooth curve vs the sharp kink.
The demo wakes as you arrive…

The idea: let the input decide its own fate, probabilistically

The key insight behind GELU is beautifully simple: merge the activation function and the regularizer into one operation.

Start with a thought experiment. Imagine that for each neuron, we flip a biased coin to decide whether to keep or drop the input — but the coin's bias depends on the input itself. Large positive inputs are almost certainly kept. Large negative inputs are almost certainly dropped. Values near zero have roughly a 50/50 chance.

Formally: multiply the input xx by a mask m∼Bernoulli(Φ(x))m \sim \text{Bernoulli}(\Phi(x)), where Φ(x)=P(X≤x)\Phi(x) = P(X \leq x) is the cumulative function (CDF) of the standard normal distribution. This means: the of keeping xx equals the probability that a random standard normal variable is less than or equal to xx.

This is a stochastic regularizer — like dropout, but input-aware. Now take its expected value to get a deterministic function:

E[x⋅m]=x⋅Φ(x)+0⋅(1−Φ(x))=x⋅Φ(x)\mathbb{E}[x \cdot m] = x \cdot \Phi(x) + 0 \cdot (1 - \Phi(x)) = x \cdot \Phi(x)

That is the GELU. No randomness at time — just a smooth, deterministic curve that encodes the probabilistic intuition.

The GELU formula

Now that we understand the intuition — the probabilistic coin flip that keeps or drops each input based on its value — let us write down the exact formula. The expected value of that stochastic process gives us a clean, deterministic activation function.

GELU(x)=x⋅Φ(x)=x⋅12[1+erf(x2)]GELU(x) = x \cdot \Phi(x) = x \cdot \frac{1}{2}\left[1 + \text{erf}\left(\frac{x}{\sqrt{2}}\right)\right]
The GELU activation — the full exact form — x = the neuron's input · Φ(x) = CDF of the standard normal distribution · erf = the Gaussian error function · the result: each input is scaled by the probability that a standard normal variable is ≤ x

Computing the exact error function can be slow in some frameworks, so the paper provides two fast approximations:

approximation (more accurate): GELU(x)≈0.5x(1+tanh⁡[2/π(x+0.044715x3)])\text{GELU}(x) \approx 0.5x\left(1 + \tanh\left[\sqrt{2/\pi}\left(x + 0.044715x^3\right)\right]\right)

Sigmoid approximation (faster): GELU(x)≈x⋅σ(1.702x)\text{GELU}(x) \approx x \cdot \sigma(1.702x)

Today, most frameworks implement the exact version efficiently, but these approximations played a key role in early adoption — especially the tanh form, which was used in BERT's original implementation.

Open in Lab
Toggle between the exact GELU and its two approximations — see how close they are.
The demo wakes as you arrive…

Why GELU works better

Three properties set GELU apart from ReLU and ELU:

Smoothness. GELU is infinitely differentiable everywhere. There is no kink, no sudden change in gradient. This makes optimization landscapes smoother, helping gradient-based optimizers find better paths. ReLU's hard corner at zero can cause oscillations in training.

Non-monotonicity. GELU has a small dip below zero near x≈−0.17x \approx -0.17. This means it can output slightly negative values — a counterintuitive property that allows it to encode subtle "anti-features." ReLU and ELU are strictly monotonic: they never decrease as the input increases.

Curvature everywhere. ReLU is perfectly linear for x>0x > 0 — there is zero curvature in the positive domain. GELU curves gently even for positive inputs, gradually approaching the identity line y=xy = x only as x→∞x \to \infty. This curvature at all points may help the network approximate complex functions more easily.

Open in Lab
Click each property tab to highlight it on the GELU curve compared to ReLU.
The demo wakes as you arrive…

The activation family tree: from ReLU to GELU

GELU sits in a clear lineage. Understanding where it came from reveals why it works:

ReLU — max⁡(0,x)=x⋅1x>0\max(0, x) = x \cdot \mathbf{1}_{x > 0} — gates by sign. The indicator function 1x>0\mathbf{1}_{x > 0} is a hard step: exactly 0 or 1. This is actually GELU with σ→0\sigma \to 0: as the Gaussian CDF sharpens to a step function, Φ(x)→1x>0\Phi(x) \to \mathbf{1}_{x > 0}, and GELU collapses to ReLU. So ReLU is a special case of GELU.

ELU — smooths ReLU's negative side with an exponential curve: α(ex−1)\alpha(e^x - 1) for x<0x < 0, but remains linear for x>0x > 0. It handles the dead-neuron problem but has no probabilistic motivation.

SiLU (Sigmoid Linear Unit) — x⋅σ(x)x \cdot \sigma(x), where σ\sigma is the logistic sigmoid. Also introduced in this paper, the SiLU uses a logistic CDF instead of a Gaussian CDF. It was later independently proposed as "Swish" by Google Brain. Similar in spirit to GELU but slightly less accurate because the logistic CDF is a looser approximation of the Gaussian.

Open in Lab
Slide σ toward zero and watch GELU collapse into ReLU — they are the same function in the limit.
The demo wakes as you arrive…

GELU in code

Three ways to compute GELUpython

Simplified to show the idea — not the real implementation.

import numpy as np
from scipy.special import erf

def gelu_exact(x):
    """Exact GELU using the error function."""
    return x * 0.5 * (1.0 + erf(x / np.sqrt(2.0)))

def gelu_tanh_approx(x):
    """Tanh approximation — used in BERT's original code."""
    return 0.5 * x * (1.0 + np.tanh(
        np.sqrt(2.0 / np.pi) * (x + 0.044715 * x**3)
    ))

def gelu_sigmoid_approx(x):
    """Sigmoid approximation — fastest, slightly less accurate."""
    return x * 1.0 / (1.0 + np.exp(-1.702 * x))

# In PyTorch: torch.nn.functional.gelu(x)
# In TensorFlow: tf.nn.gelu(x)
# Both use the exact form by default.

Experimental results

The paper tested GELU against ReLU and ELU across five domains:

  • MNIST — 8- feedforward network: GELU achieved the lowest training and validation both with and without dropout. It also showed better robustness to noised inputs.
  • MNIST autoencoding — deep (1000→500→250→30→250→500→1000): GELU significantly outperformed both competitors at multiple learning rates.
  • Twitter POS tagging — 2-layer network with pretrained word vectors: GELU achieved 12.57% error vs 12.67% for ReLU and 12.91% for ELU.
  • TIMIT speech recognition — 5-layer, 2048-neuron classifier: GELU achieved 29.3% error vs 29.5% ReLU and 29.6% ELU.
  • CIFAR-10/100 — convolutional networks: GELU achieved 7.89% error on CIFAR-10 vs 8.16% ReLU; on CIFAR-100 with wide residual networks, GELU achieved 20.74% vs 21.77% ReLU.

The consistent pattern: GELU matched or outperformed on every single benchmark.

Open in Lab
Hover over each task to see the error rates — GELU wins across the board.
The demo wakes as you arrive…

GELU in the Transformer era

Where did GELU end up? Inside the .

The original Transformer used ReLU in its feed-forward network () layers. When BERT was published in 2018, it swapped ReLU for GELU in the FFN — and achieved state-of-the-art results on 11 NLP benchmarks. GPT-2 followed suit. By GPT-3, GELU was the unchallenged default.

Why was GELU such a good fit for Transformers? The feed-forward network in each Transformer block is where the model stores and retrieves factual knowledge. GELU's smooth, value-weighted gating gives the FFN more expressive power than a hard ReLU gate. The smoothness also helps with the very deep stacks (12–96 layers) that Transformers use, because smooth gradients flow more reliably through many layers of residual connections.

Today, GELU is to Transformers what ReLU was to convolutional networks: the activation you use unless you have a specific reason not to.

  1. 2016

    GELU proposed

    Hendrycks & Gimpel introduced GELU as a probabilistically motivated activation, showing gains over ReLU and ELU on vision, NLP, and speech tasks.

  2. 2017

    Transformer paper

    "Attention Is All You Need" used ReLU in the feed-forward layers. GELU was not yet the default.

  3. 2018

    BERT adopts GELU

    BERT replaced ReLU with GELU in its FFN layers and set new records on 11 NLP benchmarks. This cemented GELU as the Transformer activation.

  4. 2019

    GPT-2 follows

    GPT-2 used GELU throughout. From this point, nearly every major language model adopted GELU by default.

  5. 2020

    GPT-3 — 175B parameters with GELU

    The largest language model of its era ran on GELU activations. Its success validated GELU at unprecedented scale.

CitationHendrycks, Gimpel. Gaussian Error Linear Units (GELUs). arXiv, 2016.

Terms in this paper