Optimization2014intermediate9 min read

Adam: A Method for Stochastic Optimization

آدَم: خوارزمية للأمثَلَة العشوائية

Kingma, D. P. · Ba, J. — ICLR

The problem

By 2014, deep neural networks required choosing an — but no single method worked well everywhere. Plain needed careful learning-rate tuning and zigzagged through narrow valleys. -based SGD added inertia but still used one global rate. AdaGrad adapted rates per- but shrank them so aggressively that learning stalled on non-convex problems. fixed the shrinking but lacked momentum and had no . Practitioners needed an optimizer that combined adaptive per-parameter rates with momentum, worked out of the box, and scaled to millions of parameters.

The contribution

(Adaptive Moment Estimation): an optimizer that maintains two exponential moving averages per parameter — the (mean of gradients, like momentum) and the (mean of squared gradients, like RMSProp) — then divides the momentum step by the root of the second moment to get an . A crucial innovation is bias correction: since the moving averages start at zero, early estimates are biased toward zero, so Adam divides by (1 − βᵗ) to compensate. The result is an optimizer with only 3 hyperparameters (α, β₁, β₂) whose defaults (0.001, 0.9, 0.999) work well across most tasks. The paper also introduces , a variant using the infinity norm.

The impact

Adam became the default optimizer for deep learning. It is used to train GPT, BERT, Vision Transformers, diffusion models, and virtually every major architecture since 2015. Its per-parameter adaptivity eliminated most learning-rate tuning, letting researchers focus on architecture design. Variants like AdamW (decoupled weight decay), LAMB, and AdaFactor descend directly from Adam. With over 200,000 citations, it is one of the most cited papers in all of computer science.

Imagine descending a foggy mountain. Plain SGD takes equal-sized steps in every direction — fine on smooth slopes, but it zigzags wildly through narrow valleys.

Momentum is like a heavy ball rolling downhill: it builds speed in consistent directions and dampens the zigzag. But it rolls at the same speed whether the terrain is flat or steep.

Adam is a smart hiker with two instruments: a compass (momentum — which direction have I been going?) and a slope meter (second moment — how steep is it here?). On flat plateaus the slope meter reads low, so Adam takes big confident strides. On steep, noisy ridges it reads high, so Adam takes careful small steps. The compass keeps the hiker from zigzagging, and the slope meter keeps the pace appropriate for the terrain.

The problem: one learning rate doesn't fit all

By 2014, stochastic descent was the workhorse of deep learning, but it had a fundamental tension: each parameter in a network lives in a different part of the landscape. Some parameters sit on steep cliffs where even small updates cause oscillation. Others sit on flat plateaus where the gradient is tiny and progress crawls. A single global cannot serve both.

Three prior ideas each solved part of the puzzle:

  • Momentum (Polyak, 1964) — accumulate a running average of past gradients. This smooths out noise and builds velocity through consistent directions, like a ball gaining speed downhill. But it uses one rate for all parameters.

  • AdaGrad (Duchi et al., 2011) — divide each parameter's update by the square root of the sum of all its past squared gradients. Frequently-updated parameters get smaller steps; rare ones get bigger steps. Perfect for sparse data, but the denominator only grows, so learning eventually dies.

  • RMSProp (Hinton, 2012) — fix AdaGrad's decay by using an of squared gradients instead of a sum. The denominator can shrink, so learning stays alive. But RMSProp has no momentum term and no bias correction.

What if you could combine momentum's direction-smoothing with RMSProp's per-parameter scaling — and fix the initialization bias that both moving averages suffer from?

Open in Lab
Watch four optimizers descend the same loss surface. Notice how Adam combines momentum's smooth path with adaptive step sizes.
The demo wakes as you arrive…

The idea: two running averages and a bias fix

Adam maintains two quantities for every single parameter in the network:

First moment estimate (m) — an exponential moving average of the gradient. This is momentum: it remembers which direction the gradient has been pointing.

mt=β1⋅mt−1+(1−β1)⋅gtm_t = \beta_1 \cdot m_{t-1} + (1 - \beta_1) \cdot g_t
First moment update — the compass (momentum) — Mix 90% of the previous direction with 10% of the new gradient

Second moment estimate (v) — an exponential moving average of the squared gradient. This tracks how large the gradient has been recently — the "energy" of the gradient signal for this parameter.

vt=β2⋅vt−1+(1−β2)⋅gt2v_t = \beta_2 \cdot v_{t-1} + (1 - \beta_2) \cdot g_t^2
Second moment update — the slope meter (gradient magnitude) — Mix 99.9% of the previous magnitude estimate with 0.1% of the new squared gradient

The update then divides momentum by the root of the second moment. Where the hat denotes bias-corrected estimates. This division is the key: it gives each parameter its own effective learning rate.

θt=θt−1−α⋅m^tv^t+ϵ\theta_t = \theta_{t-1} - \alpha \cdot \frac{\hat{m}_t}{\sqrt{\hat{v}_t} + \epsilon}
The Adam update rule — compass direction ÷ slope reading — Each parameter gets its own effective step size via the m̂/√v̂ ratio

This is the golden key: when v is small (flat plateau), the denominator is small so the step is large — advance quickly where terrain is safe. When v is large (steep ridge), the denominator is large so the step is small — slow down where terrain is dangerous.

Open in Lab
Walk through Adam's update step by step. Watch how m (momentum) and v (squared gradients) build up, and how bias correction rescales them.
The demo wakes as you arrive…

Why bias correction matters

Both m and v are initialized to zero. In the first few steps, the exponential moving average hasn't accumulated enough history, so the estimates are biased toward zero. This is not a minor technicality — without correction, the effective learning rate in the first steps can be 10× larger than intended, causing wild initial updates.

m^t=mt1−β1t  ,v^t=vt1−β2t\hat{m}_t = \frac{m_t}{1 - \beta_1^t} \;,\quad \hat{v}_t = \frac{v_t}{1 - \beta_2^t}
Bias correction — compensating for zero initialization — At step 1, dividing by (1−0.9¹)=0.1 amplifies m by 10×, exactly compensating the zero-init bias

At step 1 with β₁ = 0.9, the denominator is 0.1, so m̂ is 10× larger than m — exactly compensating the bias. As t grows, β₁ᵗ → 0 and the correction vanishes. This elegant trick is one of Adam's key contributions over RMSProp.

Open in Lab
Drag the step slider and watch the bias-corrected estimate (green) track the true mean far better than the raw moving average (gray) in early steps.
The demo wakes as you arrive…

Adaptive learning rates: why each parameter gets its own pace

The division m̂ / √v̂ has an elegant interpretation. Consider three parameters:

  • Parameter A has a consistently large gradient pointing the same direction → high m, high v. The ratio m/√v stays moderate, so it moves steadily.

  • Parameter B has a tiny gradient → low m, low v. The ratio m/√v is also moderate, so it moves at a similar effective pace despite the smaller raw gradient.

  • Parameter C has a noisy gradient that flips sign → m ≈ 0 (signs cancel), but v is large (squared values don't cancel). The ratio m/√v ≈ 0, so Adam correctly avoids updating a parameter with no clear direction.

This per-parameter scaling is what makes Adam robust across architectures: layers with sparse gradients, layers with tiny gradients, and output layers with large gradients all get appropriately-sized updates without manual tuning.

Open in Lab
Three parameters with different gradient profiles. Watch how Adam automatically adjusts the effective learning rate for each one.
The demo wakes as you arrive…

The complete algorithm

Here is Adam in its entirety — just 10 lines of math that train most of modern AI:

gt=∇θft(θt−1)mt=β1mt−1+(1−β1)gtvt=β2vt−1+(1−β2)gt2m^t=mt/(1−β1t)v^t=vt/(1−β2t)θt=θt−1−α⋅m^t/(v^t+ϵ)\begin{aligned} &g_t = \nabla_\theta f_t(\theta_{t-1}) \\ &m_t = \beta_1 m_{t-1} + (1-\beta_1) g_t \\ &v_t = \beta_2 v_{t-1} + (1-\beta_2) g_t^2 \\ &\hat{m}_t = m_t / (1 - \beta_1^t) \\ &\hat{v}_t = v_t / (1 - \beta_2^t) \\ &\theta_t = \theta_{t-1} - \alpha \cdot \hat{m}_t / (\sqrt{\hat{v}_t} + \epsilon) \end{aligned}
The Adam update rule — the engine behind most modern AI training — g = gradient · m = momentum (first moment) · v = squared gradient average (second moment) · hat = bias-corrected · α = learning rate · β₁ = momentum decay (default 0.9) · β₂ = second moment decay (default 0.999) · ε = tiny constant to prevent division by zero (default 10⁻⁸)
Adam optimizer in pure NumPypython

Simplified to show the idea — not the real implementation.

import numpy as np

def adam(params, grads, m, v, t, lr=0.001, beta1=0.9, beta2=0.999, eps=1e-8):
    """One step of Adam. Updates params in place, returns updated m, v, t."""
    t += 1
    for i in range(len(params)):
        # Update biased first moment (momentum)
        m[i] = beta1 * m[i] + (1 - beta1) * grads[i]
        # Update biased second moment (squared gradients)
        v[i] = beta2 * v[i] + (1 - beta2) * grads[i]**2
        # Bias correction
        m_hat = m[i] / (1 - beta1**t)
        v_hat = v[i] / (1 - beta2**t)
        # Update parameters
        params[i] -= lr * m_hat / (np.sqrt(v_hat) + eps)
    return m, v, t

# That's it. PyTorch's torch.optim.Adam does exactly this
# (with a few engineering optimizations) for every parameter
# in networks with billions of weights.

Adam vs. the rest: a family tree of optimizers

Adam sits at the intersection of two optimization traditions:

OptimizerMomentum?Adaptive rate?Bias correction?
SGD✗✗—
SGD + Momentum✓✗—
AdaGrad✗✓ (sum)—
RMSProp✗✓ (EMA)✗
Adam✓✓ (EMA)✓

Adam essentially combines momentum's first moment with RMSProp's second moment, then adds the bias correction that neither had. This is why its name stands for Adaptive Moment estimation.

Open in Lab
Toggle between first-moment-only (momentum), second-moment-only (RMSProp), and both combined (Adam) to see what each component contributes.
The demo wakes as you arrive…

AdaMax: the infinity-norm variant

The paper also introduces AdaMax, which replaces the second moment's L² norm with an L∞ norm (the maximum absolute value). Instead of tracking the average squared gradient, AdaMax tracks the maximum recent gradient magnitude:

ut=max⁡(β2⋅ut−1,  ∣gt∣)u_t = \max(\beta_2 \cdot u_{t-1},\; |g_t|)
AdaMax update — track the maximum gradient magnitude — Instead of mean-of-squares (Adam), use the maximum absolute gradient seen recently
θt=θt−1−α⋅m^tut\theta_t = \theta_{t-1} - \alpha \cdot \frac{\hat{m}_t}{u_t}
AdaMax update rule — momentum ÷ max gradient — No square root needed — uₜ is already a magnitude, not a squared quantity

This has two practical benefits: no square root is needed (uₜ is already a magnitude), and the denominator is more stable because max is less sensitive to outlier gradients than mean-of-squares. AdaMax is sometimes preferred when gradient magnitudes vary wildly.

The hyperparameters: why the defaults work

Adam has three hyperparameters, each with a clear role:

  • α = 0.001 (learning rate) — the overall step size. Adam adapts per-parameter, so this global rate needs less tuning than SGD's.

  • β₁ = 0.9 (first moment decay) — how much momentum to keep. 0.9 means "remember 90% of the previous direction." Lower values respond faster to gradient changes; higher values smooth more aggressively.

  • β₂ = 0.999 (second moment decay) — how much of the recent gradient magnitude to remember. 0.999 means a long memory of ~1000 steps. This keeps the denominator stable.

  • ε = 10⁻⁸ — prevents division by zero. Rarely needs changing.

The paper shows these defaults work well on convolutional networks, recurrent networks, and logistic regression — an unusually broad claim that held up in practice and is a major reason Adam became the default choice.

Open in Lab
Adjust β₁ and β₂ to see how they change Adam's behavior. Low β₁ = less smoothing, high β₂ = longer memory of gradient magnitudes.
The demo wakes as you arrive…

Why it mattered

The legacy: Adam's descendants

  1. 2014

    Adam (this paper)

    Combined momentum and adaptive rates with bias correction. Became the default optimizer for deep learning.

  2. 2017

    AMSGrad

    Fixed a theoretical convergence issue by keeping the maximum of past v values, ensuring the denominator never shrinks.

  3. 2017

    AdamW (Decoupled Weight Decay)

    Loshchilov & Hutter showed that L2 regularization in Adam is not equivalent to weight decay. AdamW applies weight decay directly to parameters, improving generalization. Now the standard for training Transformers.

  4. 2019

    LAMB

    Layer-wise Adaptive Moments for large batch training. Enabled training BERT in 76 minutes by scaling Adam to batch sizes of 64K.

  5. 2020

    AdaFactor

    Memory-efficient variant that factorizes the second moment matrix, reducing optimizer memory from O(mn) to O(m+n). Used in T5 and PaLM.

  6. 2024

    Adam remains dominant

    GPT-4, Claude, Gemini, Llama, and DeepSeek all use Adam variants. A decade later, no fundamentally different optimizer has displaced it for large-scale training.

CitationKingma, D. P. and Ba, J.. Adam: A Method for Stochastic Optimization. ICLR, 2015.

Terms in this paper