Reinforcement Learning2016intermediate12 min read

Asynchronous Methods for Deep Reinforcement Learning

الأساليب غير المتزامنة للتعلُّم المعزَّز العميق

Mnih, V. · Badia, A. P. · Mirza, M. · Graves, A. · Lillicrap, T. · Harley, T. · Silver, D. · Kavukcuoglu, K. — ICML

The problem

in 2015 relied on DQN's buffer — a giant memory bank that stores millions of past transitions and samples them randomly to break temporal . This worked, but had three costs: massive memory, restriction to algorithms only, and no way to leverage methods like that learn from their own fresh experience. was also slow: one , one , one .

The contribution

A lightweight parallel framework where multiple agents run independent copies of the environment on CPU threads and asynchronously push gradients to a shared global network. The diversity of parallel experiences decorrelates training data — replacing replay memory entirely. The best variant, Asynchronous Actor-Critic (A3C), combines n-step returns, a shared actor-critic network, and regularization. A3C beat DQN on Atari while training in half the time on a single multi-core CPU — no GPU required.

The impact

A3C proved that , not replay, is the essential ingredient for stable deep RL. Its ideas became the backbone of every subsequent scalable RL system: IMPALA scaled it to thousands of actors, refined its updates, and ICM added intrinsic curiosity on top of the same actor-critic skeleton. The synchronous variant A2C became a standard baseline. The paper democratized deep RL — suddenly you could train world-class agents on a laptop CPU.

Before A3C, training an RL agent was like one student cramming from a single textbook — they had to keep a diary of past exercises (the ) and restudy them in random order to avoid memorizing answers in sequence.

A3C replaces the diary with a study group: sixteen students each work on different chapters simultaneously, and every few minutes each one walks to a shared whiteboard and writes up what they learned. Because they are all in different chapters, the whiteboard gets diverse, decorrelated updates — no diary needed. And the whiteboard (the shared network) improves much faster than any single student could manage alone.

The problem: experience replay works, but at a cost

DQN solved a critical problem in deep RL: when a learns from consecutive game frames, the data is highly correlated — frame 100 and frame 101 look almost identical. This correlation destabilizes learning and causes the network to forget old skills as it overfits to recent experience.

DQN's solution was the experience replay buffer: store millions of past transitions (s,a,r,s′)(s, a, r, s') in memory and sample random mini-batches for training. This breaks temporal correlation brilliantly, but introduces three limitations:

  • Memory cost. Storing millions of transitions requires gigabytes of RAM.
  • Off-policy only. Replay samples are from an older version of the policy, so the must be off-policy. On-policy methods like policy gradients — which learn from their own fresh actions — cannot use stale replay data.
  • Single agent. One agent interacts with one environment, producing data sequentially. Scaling means bigger replay buffers and bigger GPUs, not more parallelism.
Open in Lab
Left: DQN stores and re-samples from one stream. Right: A3C runs many streams in parallel — no replay needed.
The demo wakes as you arrive…

The idea: parallelism replaces replay

The key insight is beautifully simple: if correlated data is the problem, don't store and shuffle — just collect data from many places at once.

Launch NN copies of the environment, each on its own CPU thread, each with its own agent holding a local copy of the network. Every agent explores independently: one might be in level 1, another in level 5, another dying on a cliff. At any given moment, the NN agents are in NN different states — their data is naturally decorrelated, just like random replay samples, but with zero memory cost.

Each agent plays for a few steps, computes gradients locally, then pushes those gradients to a single shared global network — asynchronously, without waiting for the others. It then pulls the latest global parameters and continues. The global network improves continuously from diverse, uncorrelated streams.

Open in Lab
Watch 4 workers explore independently, compute gradients, and push them to the shared network at different times. Click "Step" to advance.
The demo wakes as you arrive…

The architecture: actor and critic share a brain

A3C uses the actor-critic pattern — two roles inside a single neural network:

  • The actor () outputs a probability over actions: π(at∣st;θ)\pi(a_t | s_t; \theta). It answers: "Given what I see, what should I do?"
  • The critic () outputs a single number — the estimated state value: V(st;θv)V(s_t; \theta_v). It answers: "How good is this situation, regardless of what I do next?"

Crucially, the actor and critic share their lower layers — the convolutional or fully-connected feature extractor that processes raw pixels or state vectors. Only the final output heads branch apart: one head produces action probabilities, the other produces the value estimate. This sharing is efficient: both roles benefit from the same learned of the world.

Open in Lab
Click "Actor" or "Critic" to highlight each head's role in the shared network.
The demo wakes as you arrive…

The advantage: how much better was that action?

A raw says: "this action got RR — reinforce it." But RR alone is noisy. Was the reward high because the action was good, or because the state was already favorable?

The advantage function subtracts the critic's baseline to isolate the action's contribution: "How much better was this action compared to what I normally expect from this state?"

If A>0A > 0, the action was better than average — reinforce it. If A<0A < 0, it was worse — discourage it. This centering dramatically reduces the of the gradient estimate, making learning faster and more stable.

A(st,at)=∑i=0k−1γirt+i+γkV(st+k;θv)−V(st;θv)A(s_t, a_t) = \sum_{i=0}^{k-1} \gamma^i r_{t+i} + \gamma^k V(s_{t+k}; \theta_v) - V(s_t; \theta_v)
N-step advantage estimate — The first two terms are the n-step return — observed rewards plus the critic's estimate of the future. Subtracting V(sₜ) centers the signal: positive means "better than expected", negative means "worse than expected."

Think of it as a restaurant review system: the critic is the restaurant's average rating (3.5 stars), and the advantage is how much this specific meal deviated from that average. A 5-star meal at a 3.5-star restaurant gets A=+1.5A = +1.5 — you'd recommend it. A 2-star meal gets A=−1.5A = -1.5 — you'd warn against it. Without the baseline, you'd give raw star ratings that vary wildly depending on the restaurant's overall quality.

Open in Lab
Adjust the reward and the critic's value estimate to see how the advantage changes. Notice how centering reduces signal noise.
The demo wakes as you arrive…

Three losses, one update

A3C trains the shared network with three objectives simultaneously:

1. Policy (actor) — pushes the policy toward actions with positive advantage. This is the standard policy gradient, weighted by the advantage to reduce variance.

2. Value loss (critic) — trains the value head to accurately predict returns. A simple squared error between the predicted value and the observed .

3. — prevents the policy from collapsing to a single action too early. By adding the entropy of the action distribution to the objective, the algorithm encourages : a uniform policy has high entropy, a deterministic one has zero.

L=−log⁡π(at∣st;θ) A(st,at)+c1(Rt(n)−V(st;θv))2−c2H(π(⋅∣st;θ))L = -\log \pi(a_t|s_t;\theta)\, A(s_t,a_t) + c_1 \bigl(R_t^{(n)} - V(s_t;\theta_v)\bigr)^2 - c_2 H\bigl(\pi(\cdot|s_t;\theta)\bigr)
Combined A3C loss — Three terms: policy gradient × advantage (maximize), value error (minimize), entropy (maximize). The coefficients c₁ and c₂ balance the three objectives.
Open in Lab
Toggle each loss component on/off to see its effect on the total loss landscape.
The demo wakes as you arrive…

Entropy: the exploration thermostat

Without the entropy bonus, a policy gradient algorithm tends to become overconfident quickly: once it finds an action that gets some reward, it pushes all probability mass onto that action and stops trying alternatives. This is premature — the agent exploits too early and never discovers better strategies.

The entropy term H(π)=−∑aπ(a∣s)log⁡π(a∣s)H(\pi) = -\sum_a \pi(a|s) \log \pi(a|s) is maximized when all actions are equally likely and zero when one action has all the probability. Adding β⋅H(π)\beta \cdot H(\pi) to the objective is like a thermostat: it keeps the policy "warm" enough to keep exploring. The coefficient β\beta controls how much exploration the agent maintains — higher β\beta means more random exploration, lower means more .

Open in Lab
Drag the β slider to see how the entropy coefficient affects the policy distribution. Too low → premature collapse. Too high → random behavior.
The demo wakes as you arrive…

N-step returns: balancing bias and variance

How far ahead should the agent look before updating? This is the at the heart of learning:

  • 1-step return (rt+γV(st+1)r_t + \gamma V(s_{t+1})): updates quickly but relies heavily on the critic's estimate (which may be wrong early in training) — low variance, high bias.
  • Full return (∑γirt+i\sum \gamma^i r_{t+i}): uses only real rewards — no critic bias — but the sum of many stochastic rewards is noisy — low bias, high variance.
  • N-step return: a middle ground. Use kk real rewards, then bootstrap from the critic. A3C typically uses k=5k = 5 or k=20k = 20, getting the best of both worlds.
Rt(n)=∑i=0k−1γirt+i+γkV(st+k;θv)R_t^{(n)} = \sum_{i=0}^{k-1} \gamma^i r_{t+i} + \gamma^k V(s_{t+k}; \theta_v)
N-step return with bootstrap — k real reward steps, then bootstrap the remainder from the critic. As k → ∞ this becomes full Monte Carlo; as k → 1 it becomes pure TD.
Open in Lab
Drag n from 1 to 20 and watch how the return estimate changes. Longer n → more real rewards, less critic dependency, but higher variance.
The demo wakes as you arrive…

Putting it all together

Here is the complete A3C training loop, as each worker thread executes it:

  1. Sync — copy the global network's parameters to the local network.
  2. Collect — interact with the local environment for up to tmaxt_{max} steps (or until the episode ends), storing states, actions, and rewards.
  3. Compute returns — calculate the n-step return Rt(n)R_t^{(n)} for each step, working backwards from the last state.
  4. Compute advantages — for each step: At=Rt(n)−V(st)A_t = R_t^{(n)} - V(s_t).
  5. Compute gradients — differentiate the combined loss (policy + value + entropy) with respect to local parameters.
  6. Push — apply those gradients to the global network (asynchronously).
  7. Repeat from step 1.

The beauty is that all workers do this loop independently, at their own pace, with no synchronization barrier. The global network sees a continuous stream of diverse gradient updates.

A3C worker loop — pseudocodepython

Simplified to show the idea — not the real implementation.

import numpy as np

def a3c_worker(global_net, env, t_max=5, gamma=0.99, beta=0.01):
    """One A3C worker thread. Runs independently, pushes grads to global_net."""
    local_net = global_net.copy()           # step 1: sync

    while not done_training:
        local_net.load(global_net.params)    # re-sync each iteration
        states, actions, rewards = [], [], []

        s = env.current_state()
        for step in range(t_max):            # step 2: collect
            pi = local_net.policy(s)         # action probabilities
            a = np.random.choice(len(pi), p=pi)
            r, s_next, done = env.step(a)
            states.append(s); actions.append(a); rewards.append(r)
            s = s_next
            if done: break

        # step 3: compute n-step returns (backwards)
        R = 0 if done else local_net.value(s)
        returns = []
        for r in reversed(rewards):
            R = r + gamma * R
            returns.insert(0, R)

        # steps 4-5: advantages and gradients
        for s_t, a_t, R_t in zip(states, actions, returns):
            V = local_net.value(s_t)
            advantage = R_t - V                           # step 4
            policy_loss = -np.log(local_net.policy(s_t)[a_t]) * advantage
            value_loss  = (R_t - V) ** 2
            entropy     = -sum(p * np.log(p) for p in local_net.policy(s_t))
            loss = policy_loss + 0.5 * value_loss - beta * entropy

        global_net.apply_gradients(loss)     # step 6: async push

Results: faster, cheaper, better

A3C was evaluated on 57 Atari games, continuous control tasks (TORCS, MuJoCo), and 3D navigation (Labyrinth). The results were striking:

  • Half the training time of DQN on Atari — and on a CPU, not a GPU.
  • Near-linear speedup with more threads: 16 threads ≈ 14× faster than 1 thread.
  • On-policy learning matched or beat off-policy replay-based methods.
  • Robust to hyperparameters — a wide range of learning rates gave good results.
  • Generalized across domains — discrete Atari, continuous motor control, and 3D navigation, all with the same algorithm.
Open in Lab
Training speedup as the number of parallel workers increases. Notice the near-linear scaling.
The demo wakes as you arrive…

Why it changed everything

A3C's legacy extends far beyond the algorithm itself. It proved three principles that shaped all subsequent deep RL:

  • Parallelism is sufficient for stability. No replay buffer needed — just diverse, concurrent experience.
  • On-policy methods can compete. Before A3C, the consensus was that off-policy methods with replay were necessary. A3C showed that on-policy actor-critic, empowered by parallelism, could match or exceed them.
  • Deep RL doesn't need GPUs. By leveraging CPU threads, A3C democratized the field. Suddenly, a researcher with a laptop could run experiments that previously required a cluster.
  1. 2013

    DQN

    Deep Q-Network used experience replay and a target network to stabilize deep RL. Required a large replay buffer and was limited to off-policy, discrete actions.

  2. 2016

    A3C

    Parallelism replaced replay. On-policy actor-critic trained on multi-core CPUs, beating DQN in half the time.

  3. 2017

    ICM (Curiosity-Driven Exploration)

    Added an intrinsic curiosity module on top of A3C's actor-critic, rewarding the agent for encountering unpredictable states.

  4. 2018

    IMPALA

    Scaled A3C's paradigm to thousands of actors with a centralized learner and V-trace off-policy correction. DeepMind's workhorse for large-scale RL.

  5. 2017

    PPO

    Proximal Policy Optimization refined A3C's policy updates with a clipped surrogate objective, becoming the default algorithm for LLM alignment via RLHF.

Every time a modern LLM is aligned via RLHF, the PPO algorithm running underneath is a direct descendant of A3C's actor-critic framework. The parallel training paradigm A3C introduced is now so fundamental that it's invisible — like plumbing in a building, it's everywhere but nobody thinks about it anymore.

CitationMnih, Badia, Mirza, Graves, Lillicrap, Harley, Silver, Kavukcuoglu. Asynchronous Methods for Deep Reinforcement Learning. ICML, 2016.

Terms in this paper