Optimization2017intermediate10 min read

SGDR: Stochastic Gradient Descent with Warm Restarts

SGDR: النزول الاشتقاقي العشوائي مع إعادة التشغيل الدافئة

Loshchilov, I. · Hutter, F. — ICLR

The problem

deep networks with requires a schedule — but the standard practice of dropping the learning rate by a fixed factor at predetermined epochs is rigid and fragile. Pick the drop points wrong and the model either converges too slowly or gets stuck in a suboptimal minimum. There is no built-in mechanism to escape a bad basin once the learning rate has decayed to a small value.

The contribution

SGDR replaces the rigid step-decay schedule with plus periodic . Within each cycle the learning rate follows a cosine curve from η_max down to η_min. Then it jumps back to η_max — a "warm restart" because the model keeps its current weights rather than reinitializing. Cycle lengths can double after each restart (T_mult=2) for progressively finer exploration. SGDR achieves comparable or better results 2–4× faster than standard schedules on CIFAR-10/100, and snapshot ensembles from different restart points provide further gains for free.

The impact

Cosine annealing became the default learning rate schedule across — used in training GPT, BERT, Vision Transformers, and virtually every modern foundation model. The same authors later combined warm restarts with to create , the behind most large language models today. SGDR also popularized snapshot ensembles, showing that a single training run can produce multiple diverse models.

Imagine hiking down a mountain in fog, searching for the lowest valley. The standard approach slowly shortens your stride as you descend — eventually you're taking tiny steps, trapped in whatever dip you found, even if a deeper valley lies just over the next ridge.

SGDR works like a helicopter rescue: every so often it lifts you back up above the fog and drops you again with full-length strides. You keep your map (the learned weights) but regain the freedom to cross ridges. Each drop uses a cosine glide path — fast descent at first, gentle touchdown — and each successive flight covers a wider area so you explore more thoroughly.

The problem: rigid schedules trap the optimizer

When training deep networks with SGD, the learning rate controls how big each parameter update is. Too high and training oscillates; too low and it crawls. The standard 2016 recipe was : train at a fixed learning rate, then divide it by 5 or 10 at hand-picked epochs (e.g., epochs 60, 120, 160 out of 200).

This has two problems:

  • Fragile timing. The drop epochs are hyperparameters you must tune for every architecture and dataset. Drop too early and the model hasn't explored enough; drop too late and you waste compute.

  • No escape. Once the learning rate is small, the optimizer is locked into its current basin. If that basin is a mediocre , there's no way out — the model will fine-tune forever in the wrong neighborhood.

Open in Lab
Compare step-decay (blue) and cosine annealing with warm restarts (green). Notice how SGDR periodically jumps back up, giving the optimizer fresh momentum.
The demo wakes as you arrive…

The idea: cosine annealing with periodic restarts

SGDR replaces step decay with a smooth, periodic schedule. Within each cycle of length TiT_i epochs, the learning rate follows a cosine curve from its maximum ηmax\eta_{max} down to its minimum ηmin\eta_{min} (typically 0). When the cycle ends, the learning rate jumps back to ηmax\eta_{max} — this is the warm restart.

Why cosine? Unlike a linear ramp, the cosine shape spends more time at high learning rates (good for exploration) and more time near zero (good for ), with a smooth transition between the two. Think of it as braking gently rather than slamming the brakes.

The "warm" in warm restart is crucial: the model keeps all its learned weights. Only the learning rate resets. This means each restart starts from a region the model already considers promising, rather than from random initialization.

ηt=ηmin⁡i+12(ηmax⁡i−ηmin⁡i)(1+cos⁡(TcurTiπ))\eta_t = \eta_{\min}^{i} + \frac{1}{2}\left(\eta_{\max}^{i} - \eta_{\min}^{i}\right)\left(1 + \cos\left(\frac{T_{cur}}{T_i}\pi\right)\right)
Cosine annealing schedule — the heartbeat of SGDR — η_min and η_max define the learning rate range · T_cur counts epochs since the last restart · T_i is the current cycle length · When T_cur = 0 the cosine outputs 1 so η = η_max · When T_cur = T_i the cosine outputs −1 so η = η_min

Read the formula as a dimmer switch on a lamp: at the start of each cycle the lamp is at full brightness (high learning rate = big exploratory steps). The cosine smoothly dims it to zero (tiny refinement steps). Then the restart flips it back to full brightness and the next cycle begins.

Open in Lab
Drag the sliders to change T₀ (initial cycle length) and T_mult (cycle growth factor). Watch how doubling cycles explore progressively finer regions.
The demo wakes as you arrive…

Why restarts help: escaping local minima

The of a deep network is full of local minima and saddle points. A monotonically decreasing learning rate commits the optimizer to whichever basin it happened to fall into. Warm restarts give it a second (and third, and fourth) chance:

When the learning rate jumps back up, the large steps can push the model's parameters over nearby barriers — ridges in the loss surface that were impassable at a small learning rate. The model doesn't forget what it learned (the weights stay), but it regains the ability to make bold moves.

Loshchilov and Hutter observed that SGDR showed very mild compared to standard schedules. The periodic restarts act as implicit : the model never settles too deeply into one narrow minimum, which tends to generalize better.

Open in Lab
Watch the optimizer navigate the loss landscape. Step-decay gets stuck in basin A; SGDR's restart kicks it over the ridge into the deeper basin B.
The demo wakes as you arrive…

Doubling cycle lengths: from broad sweeps to fine polish

SGDR's second trick is multiplicative cycle growth. Instead of fixed-length cycles, each new cycle is TmultT_{mult} times longer than the previous one. With T0=1T_0 = 1 and Tmult=2T_{mult} = 2, the cycles are 1, 2, 4, 8, 16, 32, … epochs long.

The intuition: early short cycles perform rapid, coarse exploration — quickly scanning many basins. Later long cycles allow deep, fine-grained convergence within the best basin found. This is analogous to how a photographer first takes wide shots to frame the scene, then zooms in for the detail work.

This doubling strategy achieves good : the model produces a reasonable result early (after just a few short cycles) and keeps improving as longer cycles fine-tune the solution. You can stop training at any restart point and have a competitive model.

Snapshot ensembles: multiple models from one training run

Each time the learning rate hits ηmin\eta_{min} (just before a restart), the model sits at the bottom of a different basin. Loshchilov and Hutter (building on work by Huang et al.) showed that saving model snapshots at these points and averaging their predictions produces a powerful — for free.

Why does this work? The snapshots come from different basins of the loss landscape, so their errors are largely uncorrelated. When one model gets an example wrong, the others are likely to get it right. On CIFAR-10, a single WRN-28-10 scored 4.03% error; combining 3 snapshots from one run dropped it to 3.51%; and 16 runs × 3 snapshots reached 3.14% — the best result at the time.

Traditional ensembles require training N separate models (N× the compute). Snapshot ensembles give you M models from a single training run, making them essentially free.

Open in Lab
Each numbered marker is a snapshot taken at the bottom of a cosine cycle. Click "Ensemble" to see how combining them reduces error.
The demo wakes as you arrive…

The same idea in code

Cosine annealing with warm restarts, from scratchpython

Simplified to show the idea — not the real implementation.

import math

def sgdr_lr(epoch, T_0, T_mult=1, eta_max=0.05, eta_min=0):
    """Return the learning rate at a given epoch under SGDR."""
    # Find which cycle we're in and how far through it
    if T_mult == 1:
        cycle = epoch // T_0
        T_cur = epoch - cycle * T_0
        T_i = T_0
    else:
        # Geometric series: cycle lengths are T_0, T_0*T_mult, T_0*T_mult^2, ...
        cycle = 0
        t_acc = T_0
        while t_acc <= epoch:
            cycle += 1
            t_acc += T_0 * (T_mult ** cycle)
        T_i = T_0 * (T_mult ** cycle)
        T_cur = epoch - (t_acc - T_i)

    # The cosine annealing formula (eq. 5 in the paper)
    lr = eta_min + 0.5 * (eta_max - eta_min) * (1 + math.cos(math.pi * T_cur / T_i))
    return lr

# Example: T_0=10, T_mult=2 → cycles of length 10, 20, 40, 80, ...
for epoch in range(160):
    lr = sgdr_lr(epoch, T_0=10, T_mult=2, eta_max=0.05)
    # Use lr in your optimizer: optimizer.param_groups[0]['lr'] = lr

Results — faster and better

Loshchilov and Hutter trained Wide Residual Networks (WRN-28-10, 36.5M parameters) on CIFAR-10 and CIFAR-100. SGDR with T0=10,Tmult=2T_0 = 10, T_{mult} = 2 achieved around 4.03% error on CIFAR-10 and 19.58% on CIFAR-100, matching or beating the default schedule while reaching good error rates 2–4× faster.

Wider networks (WRN-28-20, 145.8M parameters) improved further to 3.74% on CIFAR-10 and 18.70% on CIFAR-100 with SGDR. The aggressive schedule let researchers test bigger networks in the same wall-clock time, since a good result appeared within the first few short cycles.

Beyond image classification, SGDR also improved results on EEG brain recordings and a downsampled ImageNet, demonstrating generality across domains.

Open in Lab
Race: step-decay vs SGDR on a simulated loss surface. SGDR reaches a low error region much faster.
The demo wakes as you arrive…

The legacy: from SGDR to AdamW and modern training

The same authors, Loshchilov and Hutter, discovered that combining warm restarts with revealed a problem: and are not equivalent in adaptive optimizers like Adam. This insight led them to propose AdamW (2019) — Adam with decoupled decay — which is now the default optimizer for training large language models, vision transformers, and diffusion models.

Cosine annealing itself became ubiquitous. Nearly every modern training recipe uses it: GPT-3, BERT, , Stable Diffusion, LLaMA — all use cosine learning rate schedules, often with a linear warmup phase at the start. The "cosine schedule" in these works traces directly back to SGDR's equation (5).

  1. 2015

    Cyclical Learning Rates (Smith)

    First proposal to cycle the learning rate during training. Used triangular and triangular2 windows. Close in spirit to SGDR but without the cosine shape or the warm restart framing.

  2. 2017

    SGDR (this paper)

    Cosine annealing with warm restarts and multiplicative cycle growth. Achieved state-of-the-art on CIFAR and demonstrated snapshot ensembles.

  3. 2017

    Snapshot Ensembles (Huang et al.)

    Formalized using SGDR's restart points to build ensembles. "Train 1, get M for free" — multiple models from one training run.

  4. 2019

    AdamW (Loshchilov & Hutter)

    The same authors discovered that warm restarts with Adam exposed a flaw in L2 regularization for adaptive methods, leading to decoupled weight decay — the optimizer behind GPT, BERT, and most modern LLMs.

  5. 2020

    Cosine schedule adopted universally

    GPT-3, ViT, and other landmark models all use cosine annealing schedules. The cosine shape from SGDR became the industry default.

The journey from SGDR to AdamW is a story of one insight building on another. Cosine annealing solved the fragile scheduling problem. Warm restarts revealed the weight decay issue. Decoupled weight decay became the foundation of modern optimization. All from a simple idea: what if we reset the learning rate?

CitationLoshchilov, Hutter. SGDR: Stochastic Gradient Descent with Warm Restarts. ICLR, 2017.

Terms in this paper