Optimization2014intermediate9 min read
Adam: A Method for Stochastic Optimization
آدَم: خوارزمية للأمثَلَة العشوائية
Kingma, D. P. · Ba, J. — ICLR
The problem
By 2014, deep neural networks required choosing an — but no single method worked well everywhere. Plain needed careful learning-rate tuning and zigzagged through narrow valleys. -based SGD added inertia but still used one global rate. AdaGrad adapted rates per- but shrank them so aggressively that learning stalled on non-convex problems. fixed the shrinking but lacked momentum and had no . Practitioners needed an optimizer that combined adaptive per-parameter rates with momentum, worked out of the box, and scaled to millions of parameters.
The contribution
(Adaptive Moment Estimation): an optimizer that maintains two exponential moving averages per parameter — the (mean of gradients, like momentum) and the (mean of squared gradients, like RMSProp) — then divides the momentum step by the root of the second moment to get an . A crucial innovation is bias correction: since the moving averages start at zero, early estimates are biased toward zero, so Adam divides by (1 − βᵗ) to compensate. The result is an optimizer with only 3 hyperparameters (α, β₁, β₂) whose defaults (0.001, 0.9, 0.999) work well across most tasks. The paper also introduces , a variant using the infinity norm.
The impact
Adam became the default optimizer for deep learning. It is used to train GPT, BERT, Vision Transformers, diffusion models, and virtually every major architecture since 2015. Its per-parameter adaptivity eliminated most learning-rate tuning, letting researchers focus on architecture design. Variants like AdamW (decoupled weight decay), LAMB, and AdaFactor descend directly from Adam. With over 200,000 citations, it is one of the most cited papers in all of computer science.
Imagine descending a foggy mountain. Plain SGD takes equal-sized steps in every direction — fine on smooth slopes, but it zigzags wildly through narrow valleys.
Momentum is like a heavy ball rolling downhill: it builds speed in consistent directions and dampens the zigzag. But it rolls at the same speed whether the terrain is flat or steep.
Adam is a smart hiker with two instruments: a compass (momentum — which direction have I been going?) and a slope meter (second moment — how steep is it here?). On flat plateaus the slope meter reads low, so Adam takes big confident strides. On steep, noisy ridges it reads high, so Adam takes careful small steps. The compass keeps the hiker from zigzagging, and the slope meter keeps the pace appropriate for the terrain.
The problem: one learning rate doesn't fit all
By 2014, stochastic descent was the workhorse of deep learning, but it had a fundamental tension: each parameter in a network lives in a different part of the landscape. Some parameters sit on steep cliffs where even small updates cause oscillation. Others sit on flat plateaus where the gradient is tiny and progress crawls. A single global cannot serve both.
Three prior ideas each solved part of the puzzle:
-
Momentum (Polyak, 1964) — accumulate a running average of past gradients. This smooths out noise and builds velocity through consistent directions, like a ball gaining speed downhill. But it uses one rate for all parameters.
-
AdaGrad (Duchi et al., 2011) — divide each parameter's update by the square root of the sum of all its past squared gradients. Frequently-updated parameters get smaller steps; rare ones get bigger steps. Perfect for sparse data, but the denominator only grows, so learning eventually dies.
-
RMSProp (Hinton, 2012) — fix AdaGrad's decay by using an of squared gradients instead of a sum. The denominator can shrink, so learning stays alive. But RMSProp has no momentum term and no bias correction.
What if you could combine momentum's direction-smoothing with RMSProp's per-parameter scaling — and fix the initialization bias that both moving averages suffer from?
The idea: two running averages and a bias fix
Adam maintains two quantities for every single parameter in the network:
First moment estimate (m) — an exponential moving average of the gradient. This is momentum: it remembers which direction the gradient has been pointing.
Second moment estimate (v) — an exponential moving average of the squared gradient. This tracks how large the gradient has been recently — the "energy" of the gradient signal for this parameter.
The update then divides momentum by the root of the second moment. Where the hat denotes bias-corrected estimates. This division is the key: it gives each parameter its own effective learning rate.
This is the golden key: when v is small (flat plateau), the denominator is small so the step is large — advance quickly where terrain is safe. When v is large (steep ridge), the denominator is large so the step is small — slow down where terrain is dangerous.
Why bias correction matters
Both m and v are initialized to zero. In the first few steps, the exponential moving average hasn't accumulated enough history, so the estimates are biased toward zero. This is not a minor technicality — without correction, the effective learning rate in the first steps can be 10× larger than intended, causing wild initial updates.
At step 1 with β₁ = 0.9, the denominator is 0.1, so m̂ is 10× larger than m — exactly compensating the bias. As t grows, β₁ᵗ → 0 and the correction vanishes. This elegant trick is one of Adam's key contributions over RMSProp.
Adaptive learning rates: why each parameter gets its own pace
The division m̂ / √v̂ has an elegant interpretation. Consider three parameters:
-
Parameter A has a consistently large gradient pointing the same direction → high m, high v. The ratio m/√v stays moderate, so it moves steadily.
-
Parameter B has a tiny gradient → low m, low v. The ratio m/√v is also moderate, so it moves at a similar effective pace despite the smaller raw gradient.
-
Parameter C has a noisy gradient that flips sign → m ≈ 0 (signs cancel), but v is large (squared values don't cancel). The ratio m/√v ≈ 0, so Adam correctly avoids updating a parameter with no clear direction.
This per-parameter scaling is what makes Adam robust across architectures: layers with sparse gradients, layers with tiny gradients, and output layers with large gradients all get appropriately-sized updates without manual tuning.
The complete algorithm
Here is Adam in its entirety — just 10 lines of math that train most of modern AI:
Simplified to show the idea — not the real implementation.
import numpy as np
def adam(params, grads, m, v, t, lr=0.001, beta1=0.9, beta2=0.999, eps=1e-8):
"""One step of Adam. Updates params in place, returns updated m, v, t."""
t += 1
for i in range(len(params)):
# Update biased first moment (momentum)
m[i] = beta1 * m[i] + (1 - beta1) * grads[i]
# Update biased second moment (squared gradients)
v[i] = beta2 * v[i] + (1 - beta2) * grads[i]**2
# Bias correction
m_hat = m[i] / (1 - beta1**t)
v_hat = v[i] / (1 - beta2**t)
# Update parameters
params[i] -= lr * m_hat / (np.sqrt(v_hat) + eps)
return m, v, t
# That's it. PyTorch's torch.optim.Adam does exactly this
# (with a few engineering optimizations) for every parameter
# in networks with billions of weights.Adam vs. the rest: a family tree of optimizers
Adam sits at the intersection of two optimization traditions:
| Optimizer | Momentum? | Adaptive rate? | Bias correction? |
|---|---|---|---|
| SGD | ✗ | ✗ | — |
| SGD + Momentum | ✓ | ✗ | — |
| AdaGrad | ✗ | ✓ (sum) | — |
| RMSProp | ✗ | ✓ (EMA) | ✗ |
| Adam | ✓ | ✓ (EMA) | ✓ |
Adam essentially combines momentum's first moment with RMSProp's second moment, then adds the bias correction that neither had. This is why its name stands for Adaptive Moment estimation.
AdaMax: the infinity-norm variant
The paper also introduces AdaMax, which replaces the second moment's L² norm with an L∞ norm (the maximum absolute value). Instead of tracking the average squared gradient, AdaMax tracks the maximum recent gradient magnitude:
This has two practical benefits: no square root is needed (uₜ is already a magnitude), and the denominator is more stable because max is less sensitive to outlier gradients than mean-of-squares. AdaMax is sometimes preferred when gradient magnitudes vary wildly.
The hyperparameters: why the defaults work
Adam has three hyperparameters, each with a clear role:
-
α = 0.001 (learning rate) — the overall step size. Adam adapts per-parameter, so this global rate needs less tuning than SGD's.
-
β₁ = 0.9 (first moment decay) — how much momentum to keep. 0.9 means "remember 90% of the previous direction." Lower values respond faster to gradient changes; higher values smooth more aggressively.
-
β₂ = 0.999 (second moment decay) — how much of the recent gradient magnitude to remember. 0.999 means a long memory of ~1000 steps. This keeps the denominator stable.
-
ε = 10⁻⁸ — prevents division by zero. Rarely needs changing.
The paper shows these defaults work well on convolutional networks, recurrent networks, and logistic regression — an unusually broad claim that held up in practice and is a major reason Adam became the default choice.
Why it mattered
The legacy: Adam's descendants
2014
Adam (this paper)
Combined momentum and adaptive rates with bias correction. Became the default optimizer for deep learning.
2017
AMSGrad
Fixed a theoretical convergence issue by keeping the maximum of past v values, ensuring the denominator never shrinks.
2017
AdamW (Decoupled Weight Decay)
Loshchilov & Hutter showed that L2 regularization in Adam is not equivalent to weight decay. AdamW applies weight decay directly to parameters, improving generalization. Now the standard for training Transformers.
2019
LAMB
Layer-wise Adaptive Moments for large batch training. Enabled training BERT in 76 minutes by scaling Adam to batch sizes of 64K.
2020
AdaFactor
Memory-efficient variant that factorizes the second moment matrix, reducing optimizer memory from O(mn) to O(m+n). Used in T5 and PaLM.
2024
Adam remains dominant
GPT-4, Claude, Gemini, Llama, and DeepSeek all use Adam variants. A decade later, no fundamentally different optimizer has displaced it for large-scale training.
CitationKingma, D. P. and Ba, J.. Adam: A Method for Stochastic Optimization. ICLR, 2015.
Terms in this paper
- Adamخوارزمية آدام
- Adaptive Learning Rateمعدل التعلم التكيُّفي
- First Momentالعزم الأول
- Second Momentالعزم الثاني
- Bias Correctionتصحيح الانحياز
- AdaMaxأداماكس