Reinforcement Learning2015intermediate9 min read

Continuous Control with Deep Reinforcement Learning

التحكم المستمر بالتعلُّم المعزِّز العميق

Lillicrap, T. P. · Hunt, J. J. · Pritzel, A. · Heess, N. · Erez, T. · Tassa, Y. · Silver, D. · Wierstra, D. — ICLR

The problem

DQN proved that deep neural networks could master Atari games with discrete actions — choose button 1, 2, or 3. But most real-world control problems have continuous action spaces: how much torque to apply to a motor, what angle to steer, how hard to push. You can't enumerate every possible real number to find the maximum Q-value, so DQN's core trick — taking the argmax over actions — simply doesn't work. Discretizing the action space creates an exponential explosion: a robot arm with 7 joints, each split into 10 bins, has 10⁷ = 10 million discrete actions per step.

The contribution

DDPG: a model-free, off-, algorithm that extends DQN's ideas to continuous action spaces. The actor network directly outputs a deterministic continuous action; the critic network evaluates that action using the . Four stabilizing mechanisms imported from DQN and beyond — , target networks with soft updates, , and Ornstein–Uhlenbeck — make training stable. The same algorithm, architecture, and hyperparameters solve 20+ simulated physics tasks: cartpole, pendulum, locomotion, grasping, and car driving.

The impact

DDPG was the first algorithm to demonstrate that could handle high-dimensional continuous control. It became the baseline for continuous-action RL and directly inspired TD3 (which fixes its overestimation and instability) and SAC (which adds regularization for better exploration). Its actor-critic template with target networks and replay buffers remains the backbone of modern continuous-control RL, from robotic manipulation to autonomous driving.

Imagine a flight instructor and a student pilot. The student (the actor) grabs the control stick and flies the plane — choosing exact stick positions, not "left or right." The instructor (the critic) watches every move and gives a score: "that turn was smooth" or "you're losing altitude."

After each lesson, the student adjusts technique based on the instructor's feedback. But there's a twist: neither of them trusts their instincts fully at first, so they each keep a notebook copy of their best strategies. They only pencil in small updates to the notebook — never erasing everything at once — so neither the flying nor the judging lurches wildly from lesson to lesson.

DQN was the student who could only flip switches (discrete actions). DDPG is the student who learned to move the stick to any position on a continuous dial.

The problem: DQN cant turn a continuous dial

DQN conquered Atari by learning a Q-value for every (, action) pair, then picking the action with the highest Q. That argmax works when actions are buttons on a gamepad — a small finite set — but breaks when the action is a real number like a steering angle or a motor torque.

The naive fix is to chop the continuous range into bins: split steering from −1 to +1 into 100 slices. But if you have nn joints each with kk bins, the number of discrete actions is knk^n — an exponential explosion. A 7-joint robot at 10 bins per joint gives 10 million actions per step. The Q-network would need to evaluate them all to find the best one.

What we need is a way to directly output the best continuous action, without scanning every possibility.

Open in Lab
Left: discrete actions (DQN picks from a fixed set). Right: continuous actions (DDPG outputs exact values). Try adjusting the discretization bins to see the curse of dimensionality.
The demo wakes as you arrive…

The idea: let one network propose, another evaluate

DDPG's solution is the actor-critic architecture, adapted for deterministic continuous actions:

The Actor μ(s∣θμ)\mu(s \mid \theta^\mu) takes a state and directly outputs a continuous action — no probability distribution, just a single precise number (or of numbers for multi-dimensional actions). Think of it as a function that maps "what I see" to "what I do."

The Critic Q(s,a∣θQ)Q(s, a \mid \theta^Q) takes the state and the actor's proposed action, and outputs a single number: the estimated total future from taking that action in that state. It answers: "how good is this exact action?"

The key insight from the theorem: we can compute exactly how to nudge the actor's parameters to increase Q, by backpropagating through the critic into the actor. No sampling needed, no from stochastic policies — just a clean signal.

Open in Lab
Follow the data flow: state enters the actor, action flows to the critic alongside the state, and the gradient flows backward through both networks.
The demo wakes as you arrive…

The learning rules

The critic learns by minimizing the difference between its current Q-estimate and a one-step Bellman target — exactly like DQN, but the "next action" comes from the actor rather than an argmax scan:

L(θQ)=E(s,a,r,s′)∼R[(Q(s,a∣θQ)−y)2],y=r+γ Q′(s′,μ′(s′∣θμ′)∣θQ′)L(\theta^Q) = \mathbb{E}_{(s,a,r,s') \sim \mathcal{R}} \Big[ \big( Q(s,a \mid \theta^Q) - y \big)^2 \Big], \quad y = r + \gamma \, Q'\big(s', \mu'(s' \mid \theta^{\mu'}) \mid \theta^{Q'}\big)
Critic loss — one-step Bellman regression — y is the target: current reward plus the discounted value of the next state, where the next action is chosen by the target actor μ', and evaluated by the target critic Q'. The primes (') mark target networks — slowly-updated copies that stabilize learning.

The actor learns by following the gradient that increases Q — "change your output in the direction the critic says is better":

∇θμJ≈Es∼R[∇aQ(s,a∣θQ)∣a=μ(s)⋅∇θμμ(s∣θμ)]\nabla_{\theta^\mu} J \approx \mathbb{E}_{s \sim \mathcal{R}} \Big[ \nabla_a Q(s, a \mid \theta^Q) \Big|_{a=\mu(s)} \cdot \nabla_{\theta^\mu} \mu(s \mid \theta^\mu) \Big]
Deterministic policy gradient — the actor's update rule — Chain rule through the critic: first, how does Q change with the action (∇ₐQ)? Then, how does the actor's output change with its parameters (∇θμ)? Multiply to get the direction of improvement.

Four tricks that make it work

Raw actor-critic with neural networks is unstable — the critic chases a moving target while the actor chases a moving critic. DDPG borrows and extends four stabilization mechanisms:

1. Experience Replay — transitions (s,a,r,s′)(s, a, r, s') are stored in a buffer and sampled randomly for training. This breaks temporal correlations between consecutive samples (which would bias the gradient) and lets each experience be reused many times, dramatically improving data efficiency. Think of it as a textbook of past flights the student reviews rather than learning only from the current moment.

2. Target Networks with Soft Updates — instead of using the same networks to both compute the target and update the weights (which creates a feedback loop), DDPG keeps slowly-updated copies: θ′←τθ+(1−τ)θ′\theta' \leftarrow \tau\theta + (1-\tau)\theta'. With τ=0.001\tau = 0.001, the targets drift gently rather than jumping, preventing the oscillations and divergence that plagued early neural RL. This is different from DQN's periodic hard copy — the is smoother and was found to be more stable for continuous domains.

θ′←τ θ+(1−τ) θ′,τ≪1\theta' \leftarrow \tau \, \theta + (1 - \tau) \, \theta', \quad \tau \ll 1
Soft target update — the notebook rule — After each training step, the target parameters θ' take a tiny step toward the current parameters θ. With τ = 0.001, this means 99.9% old + 0.1% new each step — a very slow, very stable drift.

3. Batch Normalization — different physical tasks have wildly different state scales (position in meters, velocity in m/s, angles in radians). Batch normalization standardizes each across the , so the network sees inputs in a similar range regardless of the task. This lets one architecture and one set of hyperparameters work across many different domains without re-tuning.

4. Exploration Noise (Ornstein–Uhlenbeck) — a deterministic policy always outputs the same action for the same state, so it never explores on its own. DDPG adds temporally correlated noise from an Ornstein–Uhlenbeck process: at=μ(st∣θμ)+Nta_t = \mu(s_t \mid \theta^\mu) + \mathcal{N}_t. The OU process generates noise that drifts smoothly rather than jumping randomly — well-suited for physical systems with inertia, where jerky random actions would be wasted.

Open in Lab
Toggle each stabilization trick on/off to see how training stability changes. Notice how removing any single mechanism causes the learning curve to become noisy or diverge.
The demo wakes as you arrive…

The full algorithm at a glance

Open in Lab
Step through the DDPG training loop. Each step highlights the active component.
The demo wakes as you arrive…

The core idea in code

DDPG actor-critic training loop (simplified)python

Simplified to show the idea — not the real implementation.

import numpy as np

def soft_update(target_params, current_params, tau=0.001):
    """Slowly blend current network into target network."""
    for tp, cp in zip(target_params, current_params):
        tp[:] = tau * cp + (1 - tau) * tp

def ddpg_step(actor, critic, target_actor, target_critic,
              replay_buffer, batch_size=64, gamma=0.99, tau=0.001):
    """One training step of DDPG."""
    # 1. Sample a mini-batch from the replay buffer
    states, actions, rewards, next_states = replay_buffer.sample(batch_size)

    # 2. Compute target Q-values (using target networks)
    next_actions = target_actor(next_states)           # target actor picks next action
    target_Q = rewards + gamma * target_critic(next_states, next_actions)

    # 3. Update critic: minimize Bellman error
    critic_loss = mean_squared_error(critic(states, actions), target_Q)
    critic.update(critic_loss)

    # 4. Update actor: follow the gradient that increases Q
    #    Chain rule: d(Q)/d(theta_actor) = d(Q)/d(a) * d(a)/d(theta_actor)
    predicted_actions = actor(states)
    actor_loss = -critic(states, predicted_actions).mean()  # maximize Q
    actor.update(actor_loss)

    # 5. Soft-update both target networks
    soft_update(target_actor.params, actor.params, tau)
    soft_update(target_critic.params, critic.params, tau)

# During data collection:
# action = actor(state) + OU_noise()   ← exploration
# next_state, reward = env.step(action)
# replay_buffer.store(state, action, reward, next_state)

What DDPG can do

Using the same network architecture and hyperparameters, DDPG solved 20+ continuous control tasks in the MuJoCo simulator: balancing a cartpole by applying continuous force, swinging up a pendulum with a single motor, making a humanoid walk, dexterous in-hand manipulation, and driving a car — tasks that span from 1-dimensional to 26-dimensional action spaces.

Remarkably, DDPG's learned policies performed competitively with a planning algorithm that had full access to the physics model and its derivatives. The learned to match a privileged planner using only raw experience — no model of the world.

DDPG also demonstrated learning directly from raw pixel observations using convolutional layers, though low-dimensional state representations were more sample-efficient.

Open in Lab
Select a task to see how DDPG controls continuous actions in different domains.
The demo wakes as you arrive…

Limitations and what came next

DDPG opened the door but had clear weaknesses that its successors addressed:

Q-value overestimation — the critic tends to overestimate Q-values, and the actor exploits these errors, leading to brittle policies. TD3 fixes this with clipped double Q-learning: two critics, take the minimum.

sensitivity — DDPG required careful tuning of noise parameters, learning rates, and network sizes. Small changes could cause catastrophic failure.

Poor exploration — deterministic policy + OU noise is limited. SAC replaced this with a stochastic policy that maximizes both reward and entropy, automatically balancing exploration and .

Despite these issues, DDPG's actor-critic template with target networks and replay buffers remains the blueprint for TD3, SAC, and virtually every modern continuous-control RL algorithm.

CitationLillicrap, Hunt, Pritzel, Heess, Erez, Tassa, Silver, Wierstra. Continuous Control with Deep Reinforcement Learning. ICLR, 2016.

Terms in this paper