Optimization2019intermediate11 min read
Decoupled Weight Decay Regularization
فصل اضمحلال الأوزان عن التحديث التكيُّفي
Loshchilov, I. · Hutter, F. — ICLR
The problem
Adam was known to generalize worse than with on image classification benchmarks. Practitioners had to keep switching between the two optimizers depending on the task. The root cause turned out to be a subtle but critical bug: libraries implemented and called it "," but for adaptive methods the two are not equivalent. L2 gets scaled by Adam's per-parameter adaptive learning rates, causing large-gradient weights to be regularized less — the opposite of what you want.
The contribution
A one-line fix with deep consequences: decouple the weight decay step from the adaptive gradient update. In , weights shrink by a fixed fraction λ at every step — independently of the gradient magnitude — then the adaptive update is applied separately. This makes (1) weight decay equally effective across all parameters, (2) the optimal and weight decay factor largely independent, easing tuning, and (3) Adam competitive with SGD+momentum on tasks where it previously fell short, with 15% relative improvement in test error on CIFAR-10.
The impact
AdamW became the default for Transformers and large language models. GPT-2, GPT-3, BERT, Vision Transformers, and virtually every modern foundation model uses AdamW. The insight that weight decay and L2 regularization diverge under adaptive methods reshaped how the community thinks about regularization. PyTorch's torch.optim.AdamW is now the most-used optimizer in deep learning research.
Imagine a gym with a personal trainer. The trainer is supposed to apply the same weight-loss plan to every muscle group: shrink all weights equally. But the gym has an automatic adjustment system (adaptive learning rates) that notices some muscle groups work harder, and secretly reduces their diet plan — thinking it's being helpful.
The result? The muscles that need the most restraint get the least. Your training program looks right on paper, but the automatic system has quietly rewritten it.
AdamW fires the automatic system from the diet department. It says: the weight-loss plan (weight decay) runs on a separate track — no interference from the adaptive system. The adaptive system can still adjust training intensity (learning rates), but it has no say in the diet (regularization).
The confusion: L2 regularization ≠ weight decay
For decades, two regularization techniques were treated as interchangeable:
L2 regularization adds a penalty proportional to the squared magnitude of the weights to the : . The gradient of this penalty is , which gets added to the gradient of the loss.
Weight decay directly shrinks the weights by a fixed fraction at each step: .
For standard SGD, these are mathematically equivalent (with ). But this equivalence breaks when you use an adaptive optimizer like Adam — and that breakdown was the hidden cause of Adam's poor .
Why they diverge in Adam
The core issue is simple. Adam divides each gradient by the square root of its running average of squared gradients — this is the . When L2 regularization is used, the regularization gradient also passes through this adaptive scaling.
Consider a weight with historically large gradients: Adam's denominator is large for this weight, so the effective learning rate is small. That same large denominator also shrinks the regularization signal — meaning weights that have large gradients get regularized less. The adaptive mechanism, meant to help optimization, silently undermines regularization.
Conversely, a weight with small gradients gets divided by a small denominator, amplifying both the update and the regularization — meaning quiet weights get punished more. This is exactly backwards from the intended uniform regularization.
The fix: one line changes everything
AdamW's fix is almost embarrassingly simple. Instead of adding the regularization term to the loss function (where it gets entangled with adaptive rates), apply weight decay directly to the weights as a separate step after the adaptive update:
- Compute the gradient of the loss only (no L2 term)
- Update the adaptive moments and using this clean gradient
- Compute the bias-corrected adaptive step:
- Separately shrink the weights:
The weight decay term never touches the adaptive machinery. Every weight decays at the same rate , regardless of its gradient history.
The AdamW algorithm in full
The complete AdamW algorithm is nearly identical to Adam — every line is the same except line 6 and line 12. Line 6 computes the gradient of the loss without adding the regularization term. Line 12 subtracts the weight decay term outside the adaptive ratio, keeping it independent. Click each step below to understand its role:
Why it matters: separable hyperparameters
One of the most practical benefits of is that it makes the optimal learning rate and weight decay factor largely independent. In standard Adam with L2 regularization, these two hyperparameters are tightly coupled: changing one requires retuning the other. The best settings lie on a diagonal in the grid — you can't optimize one while holding the other fixed.
With AdamW, the best settings form a rectangular region: you can find a good and a good separately, and their combination will still be near-optimal. This dramatically reduces the cost of hyperparameter search from a 2D grid to two independent 1D sweeps.
Cosine annealing and warm restarts
The paper also demonstrates that Adam benefits substantially from a scheduled learning rate — despite being adaptive. The authors propose AdamWR: AdamW combined with and from their earlier SGDR work. The learning rate multiplier follows:
The learning rate starts high, smoothly decays to near zero following a cosine curve, then "restarts" — jumps back to the top and repeats. Each restart period can grow by a factor (e.g., 100 → 200 → 400 epochs). This gives the optimizer a chance to escape local minima and explore new regions of the loss landscape.
The key insight is that adaptive learning rates (per-parameter) and global learning rate schedules serve different purposes and compose well. The per-parameter rates handle direction, while the global schedule handles exploration vs. exploitation.
Seeing weight decay in action
To build intuition for what weight decay does, picture each weight as a spring attached to zero. At every training step, the gradient pulls the weight toward the loss minimum, but the spring tugs it back toward zero. The balance between these forces determines the final weight magnitude.
With L2 regularization in Adam, the spring strength varies per weight (because the adaptive rates scale it). With decoupled weight decay, the spring strength is uniform — every weight feels the same pull toward zero. This uniform pull encourages smaller, more distributed weight values, which tend to generalize better.
The same idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
def adam_step(theta, grad, m, v, t, lr=0.001, beta1=0.9, beta2=0.999, eps=1e-8, wd=0.01):
"""Standard Adam with L2 regularization (weight decay mixed into gradient)."""
grad = grad + wd * theta # ← L2: regularizer gradient ADDED to loss gradient
m = beta1 * m + (1 - beta1) * grad
v = beta2 * v + (1 - beta2) * grad ** 2
m_hat = m / (1 - beta1 ** t)
v_hat = v / (1 - beta2 ** t)
theta = theta - lr * m_hat / (np.sqrt(v_hat) + eps)
return theta, m, v # wd was scaled by 1/sqrt(v_hat) — distorted!
def adamw_step(theta, grad, m, v, t, lr=0.001, beta1=0.9, beta2=0.999, eps=1e-8, wd=0.01):
"""AdamW: weight decay is DECOUPLED from the gradient update."""
# grad is clean — no L2 term added
m = beta1 * m + (1 - beta1) * grad
v = beta2 * v + (1 - beta2) * grad ** 2
m_hat = m / (1 - beta1 ** t)
v_hat = v / (1 - beta2 ** t)
theta = theta - lr * (m_hat / (np.sqrt(v_hat) + eps) + wd * theta) # ← decay SEPARATE
return theta, m, v # wd applied uniformly — as intended!
# That's it. One line moved. GPT, BERT, and every modern Transformer uses this fix.What the experiments showed
The authors tested AdamW against standard Adam (with L2) across multiple configurations on CIFAR-10 and ImageNet32x32. The results were consistent:
- AdamW achieved a 15% relative improvement in test error on CIFAR-10 compared to Adam with L2 regularization.
- Cosine annealing with warm restarts (AdamWR) further sped up by up to 10× in early training.
- The optimal hyperparameters for AdamW were far more separable: learning rate and weight decay could be tuned independently, slashing tuning time.
- For comparable training loss values, AdamW consistently showed better test error — evidence of genuinely improved generalization, not just faster convergence.
- AdamW closed most of the gap between Adam and SGD with momentum, making optimizer selection less of a guessing game.
Normalized weight decay: one setting for all budgets
The authors noticed that the optimal weight decay value depends on the total training budget — longer training needs smaller . They proposed to reduce this dependence: , where is , is dataset size, and is total epochs. This way, remains roughly constant across different training lengths and even across different datasets.
This normalization is not a deep theoretical result — the square-root scaling was determined empirically. But it substantially reduces the hyperparameter tuning burden: a good found on short runs transfers well to long ones.
Why it mattered
2014
Adam optimizer introduced
Kingma & Ba proposed Adam, combining momentum and adaptive learning rates. It quickly became the default optimizer but showed weaker generalization than SGD on some tasks.
2017
AdamW proposed (arXiv preprint)
Loshchilov & Hutter identified the L2/weight-decay inequivalence in adaptive optimizers and proposed decoupled weight decay. The preprint title was "Fixing Weight Decay Regularization in Adam."
2018
GPT-1 adopts AdamW
Radford et al. used AdamW to train the first GPT Transformer, establishing it as the go-to optimizer for language model pre-training.
2019
AdamW published at ICLR
The paper was formally published at ICLR 2019. PyTorch and TensorFlow added official AdamW implementations. BERT had already been using AdamW since 2018.
2020
LAMB — AdamW for large batches
You et al. extended AdamW with layer-wise adaptive rate scaling for training BERT with batch sizes up to 64K, reducing training time from 3 days to 76 minutes.
2023
Lion — evolved beyond AdamW
Chen et al. used program search to discover Lion, a simpler optimizer that uses only sign operations. Still uses decoupled weight decay — the principle persists even as the base algorithm evolves.
The original paper's title on arXiv was "Fixing Weight Decay Regularization in Adam" — refreshingly honest. The rename to the more academic "Decoupled Weight Decay Regularization" obscures the spirit: this was a bug fix. The most consequential bug fix in deep learning optimization history.
CitationLoshchilov, Hutter. Decoupled Weight Decay Regularization. ICLR, 2019.
Terms in this paper
- AdamWآدَم مع اضمحلال أوزان مفصول
- Decoupled Weight Decayاضمحلال الأوزان المفصول
- Weight Decayاضمحلال الأوزان
- L2 Regularizationالضبط الهيكلي L2
- Normalized Weight Decayاضمحلال الأوزان المُقيَّس
- Hyperparameter Decouplingفصل المعاملات الفائقة
- Cosine Annealingالتلدين الجيبي
- Warm Restartsإعادة التشغيل الدافئة