Optimization2012beginner9 min read

RMSProp: Divide the Gradient by a Running Average of Its Recent Magnitude

RMSProp: قسمة التدرّج على المتوسط المتحرك لمقداره الحديث

Tieleman, T. · Hinton, G. — COURSERA: Neural Networks for Machine Learning

The problem

By 2012 deep networks with vanilla was painful. The 's magnitude varies wildly across parameters — some weights see huge gradients, others see tiny ones — and this landscape shifts during training. A single global is either too large for the steep dimensions (causing oscillation) or too small for the flat ones (causing slow crawling). fixed this by accumulating all past squared gradients per , but its denominator grew forever, eventually shrinking the learning rate to near zero and stalling training. rprop used only the gradient sign, which works for full-batch but breaks with mini-batches because it ignores gradient magnitude averaging.

The contribution

keeps a per-parameter (EMA) of recent squared gradients, then divides each gradient by the square root of that average before updating. This gives every parameter its own that responds to recent curvature: steep directions get scaled down, flat directions get scaled up. Unlike AdaGrad, the moving average forgets the distant past, so the learning rate can recover — the adapts to the current landscape rather than being trapped by ancient history. It is, in essence, rprop made compatible with training.

The impact

Despite never being published in a formal paper — it appeared only in Geoffrey Hinton's Coursera lecture slides — RMSProp became one of the most widely used optimizers in . It was the default optimizer for training early deep reinforcement learning agents (DQN) and many early deep networks. Its core idea — adaptive per-parameter learning rates via a moving average of squared gradients — was directly incorporated into (2014), the dominant optimizer today, which combines RMSProp's second-moment estimate with 's first-moment estimate.

Training a with a single learning rate is like driving through a city where every road has a different speed limit — but your car is stuck in one gear. On the highway you crawl; in school zones you fly.

AdaGrad tries to fix this by remembering every speed bump it ever hit. After a long drive the brakes are so worn that the car barely moves at all.

RMSProp keeps a short-term memory: it tracks recent road conditions through a moving average and adjusts speed now based on what the road looks like lately — not what it looked like an hour ago. The car adapts to the current terrain, not the entire trip history.

The problem: one learning rate does not fit all

In , every parameter gets updated by the same step size. But the landscape is not uniform: some parameters sit in steep ravines where the gradient is huge, and others sit on gentle plateaus where the gradient is tiny.

A large learning rate causes steep dimensions to oscillate wildly; a small one makes flat dimensions crawl. You cannot win with a single number.

This is why researchers developed adaptive learning rate methods: give each parameter its own step size, calibrated to the geometry it actually sees.

Open in Lab
Toggle between SGD, AdaGrad, and RMSProp to see how each optimizer navigates an elongated loss valley. Notice how SGD oscillates, AdaGrad slows down, and RMSProp adapts.
The demo wakes as you arrive…

From rprop to AdaGrad: adaptive steps and their cost

rprop (resilient propagation) was an early adaptive method for full-batch learning. Its idea was simple: ignore the gradient's magnitude, use only its sign. Every parameter gets a fixed-size step — just in the direction the gradient points. This escapes plateaus quickly because even a tiny gradient produces a full step.

But rprop breaks with mini-batches. Consider a that receives gradients of +0.1 on nine mini-batches and −0.9 on the tenth. SGD would average these and keep the weight roughly in place (the mean is nearly zero). rprop, using only the sign, would step up nine times and down once — a net movement of 8 steps in the wrong direction.

AdaGrad took a different approach: divide each gradient by the square root of the sum of all past squared gradients for that parameter. Parameters with historically large gradients get smaller updates; parameters with historically small gradients get larger updates. This is excellent for sparse features.

But AdaGrad has a fatal flaw: its denominator only grows. Over long training, every parameter's effective learning rate shrinks toward zero — and it never recovers. For deep networks with non-convex loss surfaces that require long training runs, AdaGrad eventually stalls.

Open in Lab
Watch AdaGrad's effective learning rate decay to near-zero over 200 steps while RMSProp's rate stays responsive.
The demo wakes as you arrive…

The RMSProp fix: forget the distant past

Hinton's insight was elegant: instead of accumulating all past squared gradients (like AdaGrad), keep an exponential moving average of them. Recent gradients matter more than ancient ones, so the denominator reflects current curvature, not the entire training history.

The mechanism works in two steps per parameter:

Step 1 — Track recent gradient magnitude. Maintain a running estimate of how large the gradients have been lately. At each step, blend the old estimate with the new squared gradient using a decay rate ρ\rho (typically 0.9). This is like a leaky bucket: new water pours in while old water slowly drains.

Step 2 — Normalize the update. Divide the current gradient by the square root of that running estimate. If a parameter has been seeing large gradients, its update gets scaled down. If its gradients have been small, the update gets amplified. The result: every parameter moves at a pace matched to its local geometry.

vt=ρ vt−1+(1−ρ) gt2v_t = \rho \, v_{t-1} + (1 - \rho) \, g_t^2
Step 1: update the running mean of squared gradients — v_t = running mean of squared gradients · ρ = decay rate (typically 0.9) · g_t² = current gradient squared · the blend gives 90% weight to history and 10% to the new observation
θt+1=θt−ηvt+ϵ gt\theta_{t+1} = \theta_t - \frac{\eta}{\sqrt{v_t + \epsilon}} \, g_t
Step 2: the adaptive update rule — η = global learning rate · √v_t = RMS of recent gradients (the normalizer) · ε ≈ 10⁻⁸ prevents division by zero · the gradient is divided by its own recent scale, so every parameter moves at roughly the same effective pace
Open in Lab
The bucket fills with new squared gradients and leaks old ones at rate ρ. Drag the slider to change ρ and watch how the bucket's memory span changes.
The demo wakes as you arrive…

The decay rate ρ: how much history to remember

The ρ\rho controls the moving average's memory span. With ρ=0.9\rho = 0.9 (Hinton's recommended default), the effective window is roughly 10 steps — gradients older than ≈10 iterations contribute negligibly. With ρ=0.99\rho = 0.99, the window stretches to about 100 steps, producing a smoother but slower-adapting estimate.

Think of ρ\rho as the "forgetting dial":

  • High ρ (near 1) → slow forgetting → stable estimates → slower adaptation to sudden landscape changes.
  • Low ρ (near 0) → fast forgetting → noisy estimates → rapid but jittery adaptation.

The sweet spot depends on the problem. For most deep learning tasks, ρ=0.9\rho = 0.9 with η=0.001\eta = 0.001 is a strong starting point.

Open in Lab
Adjust ρ and watch how the effective memory window expands or shrinks. The bar chart shows how much weight each past gradient receives.
The demo wakes as you arrive…

Comparing the optimizers: SGD, AdaGrad, RMSProp

The key differences come down to how each optimizer treats the past:

SGD uses a single global learning rate. It has no memory of past gradients' magnitude — every step is the same size (scaled by the current gradient).

AdaGrad remembers everything. It accumulates the entire history of squared gradients. Early on, this is helpful — rarely-updated parameters get a boost. But over time the accumulator only grows, and the effective learning rate monotonically decreases toward zero.

RMSProp remembers recently. It replaces AdaGrad's ever-growing sum with an exponential moving average. Old gradients fade out, so the optimizer can speed up again after passing through a region of large gradients. This makes it far better suited for non-convex and long training runs.

The algorithm in code

RMSProp from scratch — the full optimizer in 15 linespython

Simplified to show the idea — not the real implementation.

import numpy as np

def rmsprop(params, grads, cache, lr=0.001, rho=0.9, eps=1e-8):
    """One step of RMSProp for all parameters."""
    for p, g in zip(params, grads):
        # Step 1: update running mean of squared gradients
        cache[p] = rho * cache[p] + (1 - rho) * g ** 2

        # Step 2: normalize and update
        p -= lr * g / (np.sqrt(cache[p]) + eps)

# Usage:
# cache = {p: np.zeros_like(p) for p in params}  # init once
# for each mini-batch:
#     grads = compute_gradients(loss, params)
#     rmsprop(params, grads, cache)
# That's the entire algorithm. Adam adds momentum on top of this.

From RMSProp to Adam: adding momentum

RMSProp adapts the step size per parameter using the second moment (squared gradients). But it still follows the raw gradient direction, which can be noisy with mini-batches.

Adam (Adaptive Moment Estimation, 2014) combines two ideas:

  • From momentum: keep a moving average of the gradient itself (the first moment) to smooth the direction.
  • From RMSProp: keep a moving average of squared gradients (the second moment) to adapt the step size.

Adam also adds bias correction to account for the fact that both moving averages start from zero and are biased low in early steps. This combination made Adam the default optimizer for most deep learning tasks from 2015 onward.

In other words: RMSProp is Adam without momentum and without bias correction. Understanding RMSProp means you already understand half of Adam.

Open in Lab
See how Adam's first-moment smoothing stabilizes the update direction compared to RMSProp's raw gradient.
The demo wakes as you arrive…

Why it mattered

  1. 2011

    AdaGrad published

    Duchi, Hazan, and Singer introduce adaptive per-parameter learning rates via cumulative squared gradients. Excellent for sparse features but learning rate decays to zero.

  2. 2012

    RMSProp introduced in Coursera lecture

    Hinton proposes replacing AdaGrad's cumulative sum with an exponential moving average — fixing the decaying learning rate problem. Published only as lecture slides.

  3. 2012

    Adadelta published independently

    Zeiler independently proposes a similar fix to AdaGrad using exponential moving averages, plus a second moving average of parameter updates to eliminate the need for a global learning rate.

  4. 2014

    Adam combines RMSProp + momentum

    Kingma and Ba combine RMSProp's second-moment adaptation with momentum's first-moment smoothing, plus bias correction. Adam becomes the dominant optimizer in deep learning.

  5. 2015

    DQN uses RMSProp

    DeepMind's Atari-playing DQN agent uses RMSProp as its optimizer, demonstrating its effectiveness in reinforcement learning and bringing it to mainstream attention.

  6. 2017

    AdamW decouples weight decay

    Loshchilov and Hutter show that Adam's L2 regularization is not equivalent to weight decay, and propose AdamW. The RMSProp core remains central to modern optimizers.

CitationTieleman, Hinton. Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude. COURSERA: Neural Networks for Machine Learning, 2012.

Terms in this paper