Optimization2017beginner9 min read
Cyclical Learning Rates for Training Neural Networks
معدّلات التعلّم الدورية لتدريب الشبكات العصبية
Smith, L. N. — WACV
The problem
The is the single most important in deep neural networks, yet choosing it requires tedious trial and error. Conventional wisdom says to pick a fixed value and monotonically decrease it during training. Too high and the network diverges; too low and training crawls. Practitioners waste days running grid searches, and the "optimal" schedule for one architecture rarely transfers to another.
The contribution
A family of cyclical learning rate (CLR) policies that let the learning rate oscillate between a minimum and maximum bound during training instead of only decreasing. Smith also introduces the LR range test — a single short run where the learning rate increases linearly — to automatically find good bounds. The combination eliminates manual tuning, often converges faster, and sometimes reaches higher accuracy than the best fixed schedule.
The impact
CLR became a go-to practical tool adopted by fast.ai's courses and countless practitioners. The LR range test in particular is now standard practice before any training run. The paper's core insight — that temporarily increasing the learning rate is beneficial — directly inspired SGDR (warm restarts) and the one-cycle policy, which became the default in PyTorch's scheduler library. It shifted the field's mindset from "always decay" to "oscillate intelligently."
Training a neural network with a fixed, ever-decreasing learning rate is like driving across a mountain range with your foot slowly lifting off the gas — you'll crawl safely through valleys but stall on every plateau.
Cyclical learning rates are a driver who guns the engine on plateaus to blast through them, then eases off on downhill stretches to land precisely in the valley floor. The periodic bursts of speed are the key: they look reckless for a moment but get you to the destination faster and often to a better valley.
The problem: the learning rate guessing game
Every step of updates the network's weights by: , where is the learning rate. This single number controls how far each step goes:
- Too large: steps overshoot the minimum and the explodes — the network diverges.
- Too small: steps inch forward and training takes forever, often getting stuck in shallow local minima or on the flat plateaus of saddle points.
- Just right: the sweet spot changes during training. Early on, large steps make fast progress; later, smaller steps refine the solution.
The conventional approach is a step-decay schedule: start at a reasonable value (say 0.1), then divide by 10 every epochs. Finding the right starting value, division factor, and timing requires expensive grid searches — and what works for ResNet on CIFAR-10 rarely transfers to GoogLeNet on ImageNet.
The idea: let the learning rate breathe
Smith's key observation is counterintuitive: temporarily increasing the learning rate can hurt short-term performance but help long-term training. Why?
Think about the loss landscape as a hilly terrain. A fixed learning rate settles into the nearest valley — but that valley might be shallow. A burst of higher learning rate lets the jump over low ridges and explore neighboring valleys that may be deeper (lower loss) and wider (better ).
More formally, Dauphin et al. showed that the real difficulty in isn't poor local minima — it's saddle points: flat regions where the is near zero and progress stalls. A higher learning rate provides the to blast through these plateaus.
This leads to a simple recipe: instead of only decreasing , let it cycle between a minimum bound and a maximum bound . The simplest version — the triangular policy — linearly increases from to over a half-cycle, then linearly decreases back.
The three CLR policies
Smith proposes three variants, all sharing the same idea of oscillation but differing in how the amplitude changes over time:
- triangular: the learning rate linearly rises then linearly falls between fixed bounds, repeating forever. The simplest and most robust option.
- triangular2: same triangle, but the amplitude (the difference between max and min) is halved at the end of each cycle. This gradually narrows the oscillation, combining exploration early with refinement later.
- exp_range: the bounds themselves decay exponentially by a factor of . Useful when you want the envelope to shrink smoothly rather than in steps.
All three produce the same key behavior: the learning rate spends most of its time somewhere between the bounds, so if the true optimal rate lies in that range, it gets visited regularly.
The math: computing the cyclic rate
Before seeing the formula, here is the mental picture. Imagine a timeline split into equal cycles, each cycle split into two halves (rise and fall). At any iteration, we ask: where am I in the current cycle? The answer is a number between 0 and 1 that says how far along the rise or fall I am. Multiply that by the amplitude and add the base — done.
For triangular2, simply divide the amplitude by so it halves each cycle. For exp_range, multiply the amplitude by , where is close to 1 (e.g. 0.99994).
Simplified to show the idea — not the real implementation.
import math
def triangular_clr(iteration, base_lr=0.001, max_lr=0.006, stepsize=2000):
"""Return learning rate at a given iteration."""
cycle = math.floor(1 + iteration / (2 * stepsize))
x = abs(iteration / stepsize - 2 * cycle + 1)
return base_lr + (max_lr - base_lr) * max(0, 1 - x)
# Example: stepsize=2000 means one full cycle = 4000 iterations.
# At iter 0 → lr=0.001; at iter 2000 → lr=0.006; at iter 4000 → lr=0.001.The LR range test: finding the bounds automatically
CLR needs two numbers: and . How do you pick them? Smith introduces a brilliant trick called the LR range test:
- Start with a very small learning rate (e.g. ).
- Train for a few epochs, linearly increasing the learning rate after each until it reaches a large value (e.g. 1.0).
- Plot the loss (or accuracy) against the learning rate.
The plot tells you everything you need. Where the loss starts to drop, that's a good . Where the loss stops dropping and starts to climb or oscillate wildly, that's a good . It takes one short run — typically 1–3 epochs — and replaces days of grid search.
Choosing the cycle length
The third is stepsize — the number of iterations in one half-cycle. Smith recommends setting it to 2–10 times the number of iterations in one . For example, CIFAR-10 has 50,000 images with 100, so one epoch = 500 iterations. A stepsize of 2,000–5,000 works well.
Two practical guidelines emerge from the experiments:
- Run at least 3 full cycles for the learning rate to explore the landscape thoroughly. 4 or more cycles yield even better results.
- Stop training at the end of a cycle — that's when the learning rate is at its minimum and the accuracy peaks. Stopping mid-cycle wastes the high-LR exploration phase.
CLR works with adaptive optimizers too
One might wonder: do cyclical learning rates only help plain SGD? Smith tested CLR on top of Nesterov momentum, , RMSProp, AdaGrad, and AdaDelta on CIFAR-10. The results show:
- With Nesterov+CLR, training reaches 81.3% in 25,000 iterations — the same as Nesterov alone takes 70,000 to reach.
- With Adam, CLR provides modest gains: the adaptive learning rate already handles much of the problem.
- With AdaGrad+CLR, accuracy actually improved from 74.6% to 76.0% in one-third the iterations.
The takeaway: CLR is complementary. Adaptive methods adjust per-parameter rates; CLR adjusts the global rate schedule. They solve different problems and can coexist.
Results across architectures
Smith demonstrated CLR on a deliberate variety of architectures and datasets to prove it isn't a narrow trick:
- CIFAR-10 (Caffe baseline): triangular2 matched the fixed-LR accuracy of 81.4% in 25,000 iterations instead of 70,000 — a 2.8× speedup.
- ResNet-56 on CIFAR-10/100: CLR boosted accuracy from 92.8% to 93.6% (CIFAR-10) and from 71.2% to 72.5% (CIFAR-100).
- DenseNet on CIFAR-10/100: CLR pushed accuracy from 94.5% to 94.9% on CIFAR-10 and from 75.3% to 75.9% on CIFAR-100.
- AlexNet on ImageNet: triangular2 improved accuracy from 58.0% to 58.4%.
- GoogLeNet on ImageNet: triangular2 improved from 63.0% to 64.4% — a 1.4% gain, which is significant at ImageNet scale.
The pattern is consistent: CLR either matches or beats the best fixed/decaying schedule, typically with fewer iterations.
Why it changed practice
2015
CLR paper (arXiv)
Smith posts the first version, introducing the triangular policy and the LR range test. The idea is simple, practical, and immediately useful.
2017
Published at WACV
The final version adds experiments on ResNets, DenseNets, Stochastic Depth, and ImageNet. CLR is validated across diverse architectures.
2017
SGDR — warm restarts
Loshchilov & Hutter replace the triangle with cosine annealing and add restarts. The spirit is identical to CLR: oscillate, don't just decay.
2018
Super-convergence & one-cycle policy
Smith & Topin discover that with the right CLR range, networks can train 10× faster. The one-cycle policy becomes the default in fast.ai's library.
2019
PyTorch CyclicLR & OneCycleLR
PyTorch ships built-in CyclicLR and OneCycleLR schedulers, making CLR a one-line addition to any training loop.
The LR range test alone would have been enough to make this paper essential. Combined with cyclical policies, it gave practitioners a complete, principled workflow: test once, set bounds, cycle, done. No guesswork, no grid search, no wasted GPU hours.
CitationSmith, L. N.. Cyclical Learning Rates for Training Neural Networks. WACV, 2017.
Terms in this paper
- Learning Rateمعدل التعلم
- Learning Rate Scheduleجدول معدل التعلم
- Convergenceالتقارب الحسابي
- Saddle Pointالنقطة السرجية
- Hyperparameterالمعلمة الفائقة
- Stochastic Gradient Descent (SGD)الانحدار التدريجي العشوائي
- Momentumالزخم
- Batch Sizeحجم الدفعة الحسابية
- Epochالدورة التدريبية الشاملة
- Adamخوارزمية آدام
- Lossالفقد
- Optimizerالـمُحسِّن