Reinforcement Learning2020advanced11 min read

Conservative Q-Learning for Offline Reinforcement Learning

التعلُّم المُحافِظ لدالة Q للتعلُّم المعزَّز دون اتصال

Kumar, A. · Zhou, A. · Tucker, G. · Levine, S. — NeurIPS

The problem

promises to extract good policies from previously collected datasets without any further interaction. This is crucial for domains like healthcare, autonomous driving, and robotics where online exploration is expensive or dangerous. However, standard RL methods fail catastrophically in the offline setting: they overestimate the Q-values of state-action pairs that are absent from the , because the learned visits states the never explored. This distributional shift causes the to hallucinate high values for actions, leading to policies that look great on paper but fail in practice.

The contribution

Conservative (CQL): a simple regularizer added to any Q-learning or algorithm that learns a conservative Q-function — one that provably lower-bounds the true Q-value. CQL penalizes Q-values on out-of-distribution actions (pushing them down) while boosting Q-values on in-distribution actions (those seen in the dataset). The result is a Q-function that never overestimates the policy's true value, enabling safe policy improvement from static data. CQL achieves 2–5× higher returns than prior offline RL methods across D4RL benchmarks.

The impact

CQL became the dominant baseline for offline RL research, cited in virtually every subsequent paper in the field. Its conservative principle — deliberately underestimate what you do not know — influenced Decision Transformer, IQL, and the broader pessimism-in-the-face-of- uncertainty paradigm. CQL made offline RL practical for robotics and real-world control where online exploration is infeasible, and its idea generalized beyond RL to any setting where threatens learned models.

Imagine a doctor who has only ever read patient case files but has never treated a living patient. When she encounters a new case, she might be tempted to prescribe a bold, experimental treatment because her textbook knowledge suggests it could work brilliantly.

But a wise mentor tells her: "You have never seen this treatment in action. Assume it will work slightly worse than the evidence suggests. Stick to treatments that appear reliably in your case files."

That is exactly what CQL does for agents: it forces the to be pessimistic about actions it has never tried, so it gravitates toward actions that the data actually supports — and avoids catastrophic overconfidence.

The challenge: learning without a world to practice in

In standard reinforcement learning, an agent interacts with an environment: it takes actions, observes rewards, and iteratively improves. This loop — act, observe, learn, repeat — is the engine of RL. But in many real-world domains, this loop is impossible. You cannot let a self-driving car crash thousands of times to learn. You cannot experiment with drug dosages on real patients. You cannot let a robot break expensive hardware during exploration.

Offline RL breaks the loop. Instead of interacting with the environment, the agent receives a fixed dataset D\mathcal{D} of transitions (s,a,r,s′)(s, a, r, s') collected by some previous behavior policy πβ\pi_\beta. The goal is the same — learn the best policy — but the agent can never try new actions or visit new states. It must extract the best possible strategy from the data it has, like a chess student studying grandmaster games without ever touching a board.

Open in Lab
Online RL interacts with the environment in a loop. Offline RL gets a fixed dataset and must learn without any further interaction.
The demo wakes as you arrive…

The trap: Q-value overestimation under distribution shift

Why do standard Q-learning algorithms fail offline? The answer lies in a vicious cycle between distributional shift and .

A Q-function Q(s,a)Q(s, a) estimates the expected cumulative from taking action aa in state ss and then acting optimally. Q-learning updates this estimate using the , which involves a max⁡\max over actions at the next state:

Q(s,a)←r+γmax⁡a′Q(s′,a′)Q(s,a) \leftarrow r + \gamma \max_{a'} Q(s', a')
Standard Bellman backup — the source of overestimation — The max⁡\max operator selects the highest Q-value over all possible next actions a′a'. When QQ is approximate (as with neural networks), this consistently picks actions where the approximation error is positive, creating a systematic upward bias.

In online RL, this overestimation is annoying but manageable — the agent eventually tries the overestimated action, gets a low reward, and corrects its Q-value. The environment acts as a reality check.

In offline RL, there is no reality check. The agent never tries the action, so the overestimated Q-value is never corrected. Worse, the policy then selects these overestimated actions precisely because they appear to have the highest value. This creates a feedback loop:

  1. The Q-function overestimates actions not in the dataset (out-of-distribution actions).
  2. The policy selects those overestimated actions.
  3. The Bellman backup propagates the inflated values to other states.
  4. The entire Q-function becomes unreliable.

The agent ends up with a policy that confidently chooses actions the dataset has no evidence for — like a student who aces a practice test by memorizing wrong answers.

Open in Lab
Watch the overestimation spiral: without environment correction, Q-values for out-of-distribution actions inflate uncontrollably.
The demo wakes as you arrive…

The CQL idea: be pessimistic about the unknown

The core idea of CQL is elegantly simple: add a penalty that pushes Q-values down on actions the policy might take, and pushes them up on actions that actually appear in the dataset. The first force prevents overestimation of untested actions. The second force preserves accurate estimates for actions the data supports.

Think of it as a two-handed sculptor: one hand presses down on clay where there is no evidence (suppressing hallucinated values), while the other hand lifts clay where the data provides solid ground (preserving real values). The result is a Q-function landscape that is deliberately conservative — it underestimates the true value, but never dangerously overestimates it.

Formally, CQL adds a regularizer R(θ)\mathcal{R}(\theta) to the standard Bellman error objective. The regularizer has two parts. The first part minimizes Q-values under a chosen sampling distribution μ(a∣s)\mu(a|s) — this pushes down values on actions the policy considers. The second part maximizes Q-values under the dataset distribution π^β(a∣s)\hat{\pi}_\beta(a|s) — this lifts up values on actions we have real data for:

min⁡Q    α(Es∼D[log⁡∑aexp⁡(Q(s,a))]−Es,a∼D[Q(s,a)])⏟CQL regularizer+12Es,a,s′∼D[(Q(s,a)−B^πQ(s,a))2]⏟Standard Bellman error\min_Q \;\; \alpha \underbrace{\Big( \mathbb{E}_{s \sim \mathcal{D}} \big[ \log \sum_a \exp(Q(s,a)) \big] - \mathbb{E}_{s,a \sim \mathcal{D}}[Q(s,a)] \Big)}_{\text{CQL regularizer}} + \frac{1}{2} \underbrace{\mathbb{E}_{s,a,s' \sim \mathcal{D}} \Big[ \big(Q(s,a) - \hat{\mathcal{B}}^\pi Q(s,a) \big)^2 \Big]}_{\text{Standard Bellman error}}
CQL objective (CQL(H) variant) — conservative regularizer + Bellman error — The first term inside the regularizer is a soft-maximum (log⁡∑exp⁡\log\sum\exp) over all actions — it pushes down Q-values broadly. The second term lifts Q-values on dataset actions specifically. The hyperparameter α\alpha controls how conservative the Q-function is. B^π\hat{\mathcal{B}}^\pi is the empirical Bellman operator.
Open in Lab
Adjust α to see how the CQL regularizer balances pushing down OOD Q-values against preserving in-distribution values. Higher α = more conservative.
The demo wakes as you arrive…

The guarantee: a provable lower bound on policy value

The key theoretical result of CQL is that the learned Q-function Q^π\hat{Q}^\pi lower-bounds the true Q-function QπQ^\pi at every state-action pair that the policy visits. This means that when the agent evaluates a policy using Q^\hat{Q}, it will always underestimate (never overestimate) the policy's true expected return:

V^π(s)=Ea∼π(⋅∣s)[Q^π(s,a)]≤Vπ(s)∀s\hat{V}^\pi(s) = \mathbb{E}_{a \sim \pi(\cdot|s)}[\hat{Q}^\pi(s,a)] \leq V^\pi(s) \quad \forall s
CQL's lower bound guarantee — estimated value never exceeds true value — V^π(s)\hat{V}^\pi(s) is the estimated value of the policy at state ss, computed using the conservative Q-function. Vπ(s)V^\pi(s) is the true policy value. The inequality holds under the distribution of states the policy visits. This prevents the agent from selecting policies that look good only because of inflated Q-values.

Why does this matter practically? Because in offline RL you cannot evaluate a policy by running it in the real environment. You must judge policies using your Q-function alone. If your Q-function overestimates, you will choose a bad policy that appears good. CQL's means your Q-function is a trustworthy, if cautious, judge.

Think of it as a financial auditor who always reports conservative estimates: the company might be worth more than the audit says, but it will never be worth less. You can safely make investment decisions based on this audit — you might miss upside, but you will never be blindsided by losses.

Two flavors: CQL for discrete and continuous actions

CQL has two practical variants depending on whether the is discrete (like Atari games) or continuous (like robotic control).

CQL(H) — for continuous actions. This variant uses the current policy π\pi as the sampling distribution μ\mu, so the regularizer pushes down Q-values specifically on actions the learned policy favors. The "H" stands for the entropy bonus inherited from (SAC), making CQL(H) essentially SAC with a conservative Q-function. The policy and Q-function are updated alternately.

CQL(ρ) — for discrete actions (or when you want a fixed, policy-independent regularizer). Here μ\mu is a fixed prior distribution (e.g. uniform over actions). This simplifies implementation because the regularizer does not depend on the current policy. The log⁡∑exp⁡\log\sum\exp trick over the discrete action set directly applies.

Open in Lab
Compare CQL(H) for continuous control and CQL(ρ) for discrete actions. Toggle between variants to see how the regularizer operates differently.
The demo wakes as you arrive…

Implementation: adding CQL to existing algorithms

One of CQL's strongest selling points is its simplicity. To convert any Q-learning or actor-critic algorithm into its conservative variant, you only need to add a few lines to the Q-function loss computation. The existing Bellman error loss stays unchanged; you simply add the CQL regularizer on top.

The implementation requires: sampling actions from the current policy (or uniformly for the discrete case), computing QQ values on those actions, and adding the log⁡∑exp⁡\log\sum\exp penalty minus the mean dataset QQ to the loss. This is typically 5–10 lines of code on top of a standard SAC or implementation.

CQL regularizer — the core addition to any Q-learning losspython

Simplified to show the idea — not the real implementation.

# ── CQL regularizer (add to standard Bellman error loss) ──
# Sample random actions for the logsumexp penalty
random_actions = torch.FloatTensor(batch_size, num_random, action_dim).uniform_(-1, 1)
# Q-values on random actions (pushes Q down broadly)
q_random = critic(states, random_actions)  # [B, num_random]
# Q-values on dataset actions (lifts Q up on in-distribution)
q_data = critic(states, actions)            # [B, 1]
# CQL penalty: logsumexp(Q) - E_data[Q]
cql_penalty = torch.logsumexp(q_random, dim=1).mean() - q_data.mean()
# Total loss = standard Bellman error + alpha * CQL penalty
total_loss = bellman_loss + alpha * cql_penalty

Results: CQL dominates offline RL benchmarks

CQL was evaluated on the D4RL suite — the standard testbed for offline RL. D4RL includes locomotion tasks (HalfCheetah, Hopper, Walker2d) with datasets of varying quality: random, medium, medium-replay, and expert. It also includes Atari games for discrete control evaluation.

Across the board, CQL achieved 2–5× higher normalized returns than prior methods including BCQ, BEAR, and . The gains were especially pronounced on medium-quality and mixed datasets — exactly the settings where distribution shift is most severe and where prior methods struggled the most.

On Atari games, CQL outperformed the online DQN algorithm trained for 200M steps using only 1% of the data (2M transitions). This showed that conservative offline learning can match or exceed online learning in data efficiency when the dataset is informative.

Open in Lab
CQL vs prior offline RL methods on D4RL locomotion tasks. Toggle between dataset types to see where CQL's advantage is greatest.
The demo wakes as you arrive…

Connections: how CQL builds on Q-Learning and SAC

CQL does not replace Q-Learning or SAC — it augments them. The relationship is additive: take your favorite Q-learning variant, keep its Bellman error loss unchanged, and simply add the CQL regularizer.

From Q-Learning, CQL inherits the core mechanism: iteratively improving a Q-function using the Bellman equation with a . The max⁡\max over next actions that causes overestimation in offline settings is still there — CQL counteracts it rather than removing it.

From SAC, CQL(H) inherits the maximum-entropy framework: the policy maximizes expected return plus an entropy bonus, encouraging exploration in online settings. Offline, this entropy term becomes part of the CQL regularizer itself — the log⁡∑exp⁡\log\sum\exp term is essentially the soft-maximum from SAC, repurposed to penalize rather than encourage diverse action selection.

This modular design is why CQL was adopted so quickly: researchers did not need to learn a new algorithm. They added one penalty term to code they already had.

Open in Lab
CQL's architecture: the standard SAC/DQN pipeline with the CQL regularizer injected at the Q-function loss computation step.
The demo wakes as you arrive…

Why this matters: from theory to real robots

Before CQL, offline RL was largely an academic curiosity. Algorithms worked on simple benchmarks but collapsed on realistic datasets with mixed-quality data, partial state coverage, and diverse behavior policies. CQL changed this by providing a principled solution — conservative estimation — that scales to complex domains.

In robotics, CQL enabled learning manipulation and locomotion policies from heterogeneous datasets combining human demonstrations, scripted controllers, and random exploration. In healthcare, the conservative principle aligned naturally with medical caution: a treatment policy should never be more optimistic than the clinical evidence supports.

CQL also established the paradigm of pessimism in the face of uncertainty that now permeates offline RL. Every major subsequent method — IQL, TD3+BC, Decision Transformer — either directly builds on CQL's insights or was developed as an alternative to its approach. The paper's 4000+ citations reflect its role as the field's inflection point.

Timeline: the road to and from CQL

  1. 2005

    Batch RL foundations

    Ernst, Lange, and Riedmiller formalized batch (offline) RL as learning from a fixed dataset without interaction. Fitted Q-Iteration showed that offline Bellman updates work in principle but suffer from distribution mismatch.

  2. 2018

    Soft Actor-Critic (SAC)

    Haarnoja et al. introduced maximum-entropy RL, adding an entropy bonus to the standard RL objective. SAC became the strongest online continuous-control method and later the foundation for CQL(H).

  3. 2019

    BCQ & BEAR — constraint-based offline RL

    Fujimoto et al. (BCQ) and Kumar et al. (BEAR) constrained the policy to stay close to the behavior policy, preventing OOD action selection. Effective but overly restrictive — they could not improve much beyond the data-collecting policy.

  4. 2020

    CQL (this paper)

    Shifted the approach from constraining the policy to regularizing the Q-function. Instead of preventing OOD actions, CQL makes their Q-values low so the policy avoids them naturally. Achieved 2–5× better returns than BCQ/BEAR.

  5. 2021

    Decision Transformer & IQL

    Chen et al. reframed offline RL as sequence modeling. Kostrikov et al. proposed Implicit Q-Learning which avoids querying OOD actions entirely. Both built on CQL's insight that handling distribution shift is the central challenge.

  6. 2023

    Offline RL at scale — robotics & real-world deployment

    Teams at Google and Berkeley deployed CQL-based methods on real robots, learning manipulation from thousands of demonstration trajectories. The conservative principle proved essential for safe real-world deployment.

CitationKumar, Zhou, Tucker, Levine. Conservative Q-Learning for Offline Reinforcement Learning. NeurIPS, 2020.

Terms in this paper