Deep Learning2016beginner10 min read
Gaussian Error Linear Units (GELUs)
وحدات الخطأ الغاوسي الخطية (GELUs)
Hendrycks, D. · Gimpel, K. — arXiv
The problem
Activation functions and stochastic regularizers (like ) have always been treated as separate design choices. makes a hard binary decision — keep or kill — based only on the sign of the input, producing a non-smooth kink at zero. This hard gating ignores how confident the network is about an input's importance. Meanwhile, dropout randomly zeroes neurons without considering their value at all. No unified these two ideas into a single, smooth, probabilistically motivated operation.
The contribution
The activation function: GELU(x) = x · Φ(x), where Φ(x) is the standard Gaussian . Instead of gating by sign (like ReLU) or gating randomly (like dropout), GELU weights each input by how likely it is to be greater than other inputs under a Gaussian assumption. This produces a smooth, non-monotonic curve that acts as the expected value of a stochastic regularizer — unifying nonlinearity and in one formula. GELU outperformed ReLU and ELU across computer vision, NLP, and speech tasks.
The impact
GELU became the default activation function in Transformers. BERT, GPT-2, GPT-3, and nearly every major language adopted GELU, making it one of the most widely deployed activation functions in modern . The paper also introduced the Linear Unit (SiLU), later independently rediscovered as "Swish" by Google Brain. GELU demonstrated that principled probabilistic thinking about activations leads to practical performance gains.
Think of a nightclub bouncer. ReLU is a bouncer who checks your ID: positive? You're in. Negative? Absolutely not. There's no middle ground — even someone at −0.001 is turned away cold, while +0.001 walks right through.
GELU is a smarter bouncer with a bell-curve intuition. Instead of a hard cutoff, this bouncer asks: "How likely is it that you belong here?" Very positive inputs stroll in with full confidence. Very negative inputs are turned away. But the borderline cases — the values near zero — get a proportional welcome: the bouncer lets through a fraction of you, scaled by how plausible your presence is. No sudden cliff, just a smooth judgment call.
The problem: hard gates and disconnected regularizers
By 2016, the ReLU activation had become the workhorse of . Its formula is disarmingly simple: . If the input is positive, pass it through unchanged. If negative, output zero. This simplicity made it fast and effective — far better than the sigmoid functions that came before, which suffered from vanishing gradients in deep networks.
But ReLU has a sharp edge. At , the function makes a sudden, discontinuous jump in its — from 0 to 1. This hard kink means the function is not smooth: its flips instantaneously. For values just below zero, the is completely dead. For values just above, it is fully alive. There is no nuance near the boundary.
Meanwhile, stochastic regularizers like dropout were applied separately. Dropout randomly zeroes neurons during to prevent — but it does so blindly, without considering whether a neuron's signal is large or small. The activation decision and the regularization decision were two independent mechanisms, never talking to each other.
The idea: let the input decide its own fate, probabilistically
The key insight behind GELU is beautifully simple: merge the activation function and the regularizer into one operation.
Start with a thought experiment. Imagine that for each neuron, we flip a biased coin to decide whether to keep or drop the input — but the coin's bias depends on the input itself. Large positive inputs are almost certainly kept. Large negative inputs are almost certainly dropped. Values near zero have roughly a 50/50 chance.
Formally: multiply the input by a mask , where is the cumulative function (CDF) of the standard normal distribution. This means: the of keeping equals the probability that a random standard normal variable is less than or equal to .
This is a stochastic regularizer — like dropout, but input-aware. Now take its expected value to get a deterministic function:
That is the GELU. No randomness at time — just a smooth, deterministic curve that encodes the probabilistic intuition.
The GELU formula
Now that we understand the intuition — the probabilistic coin flip that keeps or drops each input based on its value — let us write down the exact formula. The expected value of that stochastic process gives us a clean, deterministic activation function.
Computing the exact error function can be slow in some frameworks, so the paper provides two fast approximations:
approximation (more accurate):
Sigmoid approximation (faster):
Today, most frameworks implement the exact version efficiently, but these approximations played a key role in early adoption — especially the tanh form, which was used in BERT's original implementation.
Why GELU works better
Three properties set GELU apart from ReLU and ELU:
Smoothness. GELU is infinitely differentiable everywhere. There is no kink, no sudden change in gradient. This makes optimization landscapes smoother, helping gradient-based optimizers find better paths. ReLU's hard corner at zero can cause oscillations in training.
Non-monotonicity. GELU has a small dip below zero near . This means it can output slightly negative values — a counterintuitive property that allows it to encode subtle "anti-features." ReLU and ELU are strictly monotonic: they never decrease as the input increases.
Curvature everywhere. ReLU is perfectly linear for — there is zero curvature in the positive domain. GELU curves gently even for positive inputs, gradually approaching the identity line only as . This curvature at all points may help the network approximate complex functions more easily.
The activation family tree: from ReLU to GELU
GELU sits in a clear lineage. Understanding where it came from reveals why it works:
ReLU — — gates by sign. The indicator function is a hard step: exactly 0 or 1. This is actually GELU with : as the Gaussian CDF sharpens to a step function, , and GELU collapses to ReLU. So ReLU is a special case of GELU.
ELU — smooths ReLU's negative side with an exponential curve: for , but remains linear for . It handles the dead-neuron problem but has no probabilistic motivation.
SiLU (Sigmoid Linear Unit) — , where is the logistic sigmoid. Also introduced in this paper, the SiLU uses a logistic CDF instead of a Gaussian CDF. It was later independently proposed as "Swish" by Google Brain. Similar in spirit to GELU but slightly less accurate because the logistic CDF is a looser approximation of the Gaussian.
GELU in code
Simplified to show the idea — not the real implementation.
import numpy as np
from scipy.special import erf
def gelu_exact(x):
"""Exact GELU using the error function."""
return x * 0.5 * (1.0 + erf(x / np.sqrt(2.0)))
def gelu_tanh_approx(x):
"""Tanh approximation — used in BERT's original code."""
return 0.5 * x * (1.0 + np.tanh(
np.sqrt(2.0 / np.pi) * (x + 0.044715 * x**3)
))
def gelu_sigmoid_approx(x):
"""Sigmoid approximation — fastest, slightly less accurate."""
return x * 1.0 / (1.0 + np.exp(-1.702 * x))
# In PyTorch: torch.nn.functional.gelu(x)
# In TensorFlow: tf.nn.gelu(x)
# Both use the exact form by default.Experimental results
The paper tested GELU against ReLU and ELU across five domains:
- MNIST — 8- feedforward network: GELU achieved the lowest training and validation both with and without dropout. It also showed better robustness to noised inputs.
- MNIST autoencoding — deep (1000→500→250→30→250→500→1000): GELU significantly outperformed both competitors at multiple learning rates.
- Twitter POS tagging — 2-layer network with pretrained word vectors: GELU achieved 12.57% error vs 12.67% for ReLU and 12.91% for ELU.
- TIMIT speech recognition — 5-layer, 2048-neuron classifier: GELU achieved 29.3% error vs 29.5% ReLU and 29.6% ELU.
- CIFAR-10/100 — convolutional networks: GELU achieved 7.89% error on CIFAR-10 vs 8.16% ReLU; on CIFAR-100 with wide residual networks, GELU achieved 20.74% vs 21.77% ReLU.
The consistent pattern: GELU matched or outperformed on every single benchmark.
GELU in the Transformer era
Where did GELU end up? Inside the .
The original Transformer used ReLU in its feed-forward network () layers. When BERT was published in 2018, it swapped ReLU for GELU in the FFN — and achieved state-of-the-art results on 11 NLP benchmarks. GPT-2 followed suit. By GPT-3, GELU was the unchallenged default.
Why was GELU such a good fit for Transformers? The feed-forward network in each Transformer block is where the model stores and retrieves factual knowledge. GELU's smooth, value-weighted gating gives the FFN more expressive power than a hard ReLU gate. The smoothness also helps with the very deep stacks (12–96 layers) that Transformers use, because smooth gradients flow more reliably through many layers of residual connections.
Today, GELU is to Transformers what ReLU was to convolutional networks: the activation you use unless you have a specific reason not to.
2016
GELU proposed
Hendrycks & Gimpel introduced GELU as a probabilistically motivated activation, showing gains over ReLU and ELU on vision, NLP, and speech tasks.
2017
Transformer paper
"Attention Is All You Need" used ReLU in the feed-forward layers. GELU was not yet the default.
2018
BERT adopts GELU
BERT replaced ReLU with GELU in its FFN layers and set new records on 11 NLP benchmarks. This cemented GELU as the Transformer activation.
2019
GPT-2 follows
GPT-2 used GELU throughout. From this point, nearly every major language model adopted GELU by default.
2020
GPT-3 — 175B parameters with GELU
The largest language model of its era ran on GELU activations. Its success validated GELU at unprecedented scale.
CitationHendrycks, Gimpel. Gaussian Error Linear Units (GELUs). arXiv, 2016.
Terms in this paper
- Activation Functionدالة التنشيط
- ReLUدالة الوحدة الخطية المصححة
- GELUوحدة الخطأ الخطية الغاوسية (GELU)
- Dropoutالإسقاط العشوائي للعصبونات
- Stochastic Regularizationالتنظيم العشوائي
- Softplusالدالة اللينة الموجبة
- Sigmoidدالة سيجمويد
- Cumulative Distribution Functionدالة التوزيع التراكمي