Language Models2025advanced11 min read

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

DeepSeek-R1: تحفيز قدرات الاستدلال في النماذج اللغوية الكبيرة عبر التعلم المعزز

DeepSeek-AI · Guo, D. · Yang, D. · Zhang, H. · Song, J. · Wang, P. · Zhu, Q. · Xu, R. · Zhang, R. · Ma, S. · Bi, X. — arXiv / Nature

The problem

Teaching LLMs to reason typically requires massive human-annotated step-by-step solutions (supervised ). This is expensive, it does not scale, and it caps the at human-level reasoning — the model can never discover strategies the annotators did not think of. Moreover, neural reward models used in RLHF are susceptible to as scales up.

The contribution

DeepSeek-R1-Zero: a trained with pure — no supervised fine-tuning at all. Using (GRPO) and rule-based rewards (accuracy + format), the model spontaneously develops sophisticated reasoning behaviors: , verification, and dynamic strategy switching. DeepSeek-R1 builds on this with a multi-stage pipeline ( → RL → → SFT → RL) to fix readability and language mixing while preserving reasoning power. The reasoning patterns can be distilled into smaller models (1.5B–70B).

The impact

DeepSeek-R1 proved that reinforcement learning alone can unlock reasoning capabilities rivaling OpenAI's o1 — without relying on expensive human annotations. It achieved 79.8% on AIME 2024, 97.3% on MATH-500, and outperformed 96.3% of human competitors on Codeforces. The open-source release under MIT license democratized access to frontier reasoning models and sparked a wave of RL-for-reasoning research worldwide.

Traditional LLM training is like a driving school: a human instructor sits beside the student, demonstrating every turn and lane change. The student can never be better than the instructor.

DeepSeek-R1's approach is like a racing simulator: give the student a car, a track, and a lap timer. Say nothing about technique. After thousands of laps, the student discovers racing lines, braking points, and overtaking strategies the instructor never thought of — because the only feedback is the clock.

The "aha moment" is when the driver, mid-lap, slams the brakes, mutters "wait, that corner was wrong", reverses mental course, and finds a faster line — all on their own.

The problem: human demonstrations cap reasoning

Before DeepSeek-R1, the standard recipe for teaching LLMs to reason was a two-stage pipeline: first, supervised fine-tuning (SFT) on thousands of human-written step-by-step solutions, then (RLHF) to polish outputs. This approach has three fundamental limitations:

  • Scalability wall. High-quality reasoning traces are expensive to produce. Human experts spend minutes per problem, and the resulting data cannot cover the vast space of possible reasoning strategies.

  • Ceiling effect. The model learns to imitate human reasoning, so it can never surpass the quality of its training data. Novel strategies — like spending extra tokens to verify an answer, or spontaneously switching approaches — are never demonstrated and therefore never learned.

  • Reward hacking. Neural reward models used in RLHF can be exploited: the model finds shortcuts that score highly without genuinely reasoning. Retraining the reward model requires more compute and human effort, creating a vicious cycle.

Open in Lab
Compare the traditional SFT→RLHF pipeline with DeepSeek's pure RL approach. Click each stage to see the difference.
The demo wakes as you arrive…

The engine: Group Relative Policy Optimization (GRPO)

DeepSeek-R1-Zero starts from a pre-trained base model (DeepSeek-V3-Base, 671B parameters with 37B active via ) and applies reinforcement learning directly — no supervised fine-tuning. The RL algorithm is GRPO (Group Relative Policy Optimization), which was designed to simplify PPO by removing the need for a separate value model.

Here is the intuition: for each question, GRPO generates a group of candidate answers (typically 16). Some answers are correct, others wrong. Instead of needing a learned to estimate how good a partial answer is, GRPO simply compares each answer to its group-mates: if your reward is above the group average, you are reinforced; if below, you are penalized. The of each output is computed relative to its group — hence "Group Relative."

Think of it as a class ranking: instead of an external grader (the value model in PPO), the students grade themselves relative to each other. The best-performing solution in each batch gets the strongest positive signal.

JGRPO(θ)=E[1G∑i=1Gmin⁡(πθ(oi∣q)πθold(oi∣q)Ai,  clip(πθ(oi∣q)πθold(oi∣q),1−ε,1+ε)Ai)−β DKL(πθ∥πref)]J_{GRPO}(\theta) = \mathbb{E}\left[\frac{1}{G}\sum_{i=1}^{G} \min\left(\frac{\pi_\theta(o_i|q)}{\pi_{\theta_{old}}(o_i|q)} A_i,\; \text{clip}\left(\frac{\pi_\theta(o_i|q)}{\pi_{\theta_{old}}(o_i|q)}, 1-\varepsilon, 1+\varepsilon\right) A_i\right) - \beta\, D_{KL}(\pi_\theta \| \pi_{ref})\right]
GRPO objective — maximize advantage across the group while staying near the reference — G = group size (16 outputs per question) · Aᵢ = advantage of output i relative to group · clip keeps updates stable · D_KL prevents drifting too far from the reference policy
Ai=ri−mean({r1,r2,…,rG})std({r1,r2,…,rG})A_i = \frac{r_i - \text{mean}(\{r_1, r_2, \ldots, r_G\})}{\text{std}(\{r_1, r_2, \ldots, r_G\})}
Group-relative advantage — no value model needed — Each output's reward rᵢ is simply z-scored within its group. If you scored higher than the group mean, your advantage is positive; lower, it is negative. This replaces the expensive value model required by PPO.
Open in Lab
Step through GRPO's process. Click "Sample" to generate a group of answers, then watch advantages get computed.
The demo wakes as you arrive…

Reward design: simple rules, no neural judges

A critical design choice in DeepSeek-R1-Zero is the deliberate avoidance of neural reward models for reasoning tasks. Neural reward models are prone to reward hacking — the model learns to exploit blind spots in the judge rather than genuinely improve. Retraining the reward model consumes computational resources and adds pipeline complexity.

Instead, DeepSeek-R1-Zero uses rule-based rewards with two components:

  • Accuracy reward: did the model get the right answer? For math, the final answer is extracted and compared against ground truth. For code, a compiler runs the solution against test cases. Binary: 1 if correct, 0 if wrong.

  • Format reward: did the model wrap its reasoning in the expected <think>...</think> and <answer>...</answer> tags? This ensures the reasoning process is explicitly delineated, making it interpretable and analyzable.

The total reward is simply the sum of these two signals. No learned reward model, no human preference data, no reward-model retraining loop. The simplicity is the point: it eliminates the reward hacking failure mode that plagues neural approaches.

Rrule=Raccuracy+RformatR_{rule} = R_{accuracy} + R_{format}
Total rule-based reward — deliberately simple — Accuracy: binary (0 or 1) based on answer correctness · Format: binary based on proper use of think/answer tags · No neural reward model involved
Open in Lab
Toggle reward components on and off to see how they affect the training signal.
The demo wakes as you arrive…

Emergent behaviors: the model teaches itself to think

The most striking finding of DeepSeek-R1-Zero is that sophisticated reasoning strategies emerge spontaneously from pure RL — with no human demonstration. As training progresses, three phenomena unfold:

  • Self-extending thinking time. The model's average response length grows steadily from ~3,000 tokens to over 10,000 tokens. Nobody told it to think longer; it discovered on its own that spending more tokens on reasoning leads to higher rewards.

  • Reflective reasoning. Words like "wait", "mistake", "however", "verify", and "retry" increase 5–7× during training. The model learns to pause, question its own logic, and correct course — a form of metacognition that was never explicitly taught.

  • The "aha moment." Around training step 8,000, the model suddenly begins using phrases like "Wait, wait. That's an aha moment" — rethinking its approach mid-solution. The frequency of the word "wait" spikes dramatically, marking a phase transition in reasoning strategy. This moment is not programmed; it emerges from the reward signal alone.

AIME 2024 accuracy climbs from 15.6% to 77.9% during training. With majority voting over 16 samples, it reaches 86.7% — surpassing the average human competitor score.

Open in Lab
Watch how response length and reflective reasoning grow during RL training. Drag the step slider to explore.
The demo wakes as you arrive…

From R1-Zero to R1: a multi-stage refinement

DeepSeek-R1-Zero proves that pure RL works, but it has practical problems: the reasoning text is hard to read, and the model mixes Chinese and English unpredictably within a single chain-of-thought. It also cannot handle non-reasoning tasks like creative writing or open-domain QA well.

DeepSeek-R1 fixes these issues with a four-stage pipeline:

  • Stage 1 — Cold start SFT. A small dataset of thousands of readable, first-person reasoning traces is collected. Human annotators rewrite R1-Zero outputs into a conversational style ("I need to think about..." rather than "We observe that..."), then an LLM rewrites more data in this style. This is used for a brief SFT phase that gives the RL a better starting point.

  • Stage 2 — First RL with language consistency. Reasoning-focused RL training, just like R1-Zero, but with an added language consistency reward: the proportion of target-language words in the chain-of-thought. This nudges the model to stay in one language throughout its reasoning.

  • Stage 3 — Rejection sampling + SFT. The RL checkpoint generates solutions; only correct, well-formatted ones are kept. About 600K reasoning samples plus 200K non-reasoning samples (writing, QA, translation) form a combined SFT dataset.

  • Stage 4 — Second RL. A final round of RL using both rule-based rewards (reasoning) and model-based rewards (helpfulness/safety), creating a model that reasons well and follows instructions.

Rlanguage=N(Wordstarget)N(Words)R_{language} = \frac{N(\text{Words}_{target})}{N(\text{Words})}
Language consistency reward — fraction of target-language words — Counts the proportion of words in the target language (e.g. Chinese or English) within the chain-of-thought. Added to the total reward to discourage language mixing. Simple but effective — ablations show it maintains language consistency with only marginal reasoning performance loss.
Open in Lab
Click each stage to explore the multi-stage pipeline from V3-Base to the final DeepSeek-R1.
The demo wakes as you arrive…

Distillation: reasoning in smaller packages

A large model that reasons brilliantly is useful, but a small model that reasons well is transformative — it can run on a single , on-device, or at the edge. DeepSeek distilled R1's reasoning capabilities into six smaller models ranging from 1.5B to 70B parameters, built on Qwen and Llama base architectures.

The process is straightforward: fine-tune the smaller base model on the same ~800K high-quality reasoning trajectories generated by R1. No RL is needed for the student — the SFT data already encodes the reasoning patterns that RL discovered.

The key finding: distilled reasoning outperforms RL-from-scratch on small models. When researchers tried applying RL directly to Qwen-32B (the same approach that worked for the 671B model), the results were worse than simply distilling from R1. This suggests that discovering advanced reasoning strategies requires scale, but once discovered, those strategies can be efficiently transferred to smaller architectures.

Open in Lab
Compare distilled models against their base counterparts on AIME and MATH benchmarks.
The demo wakes as you arrive…

The same idea in code

GRPO advantage computation, simplifiedpython

Simplified to show the idea — not the real implementation.

import numpy as np

def compute_grpo_advantages(rewards: list[float]) -> list[float]:
    """Compute group-relative advantages for a batch of outputs.

    Each output's reward is z-scored within its group:
    advantage_i = (r_i - mean(group)) / std(group)

    If your reward is above the group average, you get reinforced.
    If below, you get penalized. No value model needed.
    """
    rewards = np.array(rewards)
    mean = rewards.mean()
    std = rewards.std() + 1e-8  # avoid division by zero
    return ((rewards - mean) / std).tolist()

def rule_based_reward(answer: str, ground_truth: str, has_tags: bool) -> float:
    """DeepSeek-R1-Zero's entire reward function.
    No neural reward model — just correctness + format.
    """
    accuracy = 1.0 if answer.strip() == ground_truth.strip() else 0.0
    format_ok = 1.0 if has_tags else 0.0
    return accuracy + format_ok

# Example: 4 sampled outputs for one math question
rewards = [
    rule_based_reward("42", "42", True),   # correct + formatted → 2.0
    rule_based_reward("42", "42", False),   # correct but no tags → 1.0
    rule_based_reward("17", "42", True),    # wrong + formatted   → 1.0
    rule_based_reward("99", "42", False),   # wrong + no tags     → 0.0
]
advantages = compute_grpo_advantages(rewards)
# [1.34, 0.0, 0.0, -1.34] — the correct+formatted answer
# gets the strongest positive signal; the wrong+unformatted
# gets the strongest negative signal.

Results and impact

DeepSeek-R1 achieves performance comparable to OpenAI's o1-1217 across reasoning benchmarks while being fully open-source (MIT license). On mathematics, it scores 79.8% on AIME 2024 (vs o1's 79.2%) and 97.3% on MATH-500 (vs o1's 96.4%). On coding, it reaches a Codeforces rating of 2029, placing it in the 96.3rd percentile of human competitors. On knowledge benchmarks like MMLU-Pro, it achieves 84.0%.

The distilled models are equally impressive: DeepSeek-R1-Distill-Qwen-32B surpasses OpenAI-o1-mini on most reasoning benchmarks, despite being dramatically smaller. The 7B distilled model outperforms many non-reasoning models ten times its size.

The training cost is remarkably efficient: approximately 147K H800 GPU hours total (~$294K USD), a fraction of what comparable closed-source reasoning models require.

  1. 2017

    Transformer (Attention Is All You Need)

    The architecture that made everything possible — self-attention replaces recurrence, enabling massive parallelism and the scaling era.

  2. 2017

    PPO (Proximal Policy Optimization)

    The RL algorithm that became standard for RLHF — stable policy updates via clipped surrogate objective. Requires a learned value model.

  3. 2022

    InstructGPT / RLHF

    SFT then RLHF became the standard post-training recipe. Powerful but dependent on human demonstrations and neural reward models.

  4. 2022

    Chain-of-Thought Prompting

    Wei et al. showed that prompting LLMs with "Let's think step by step" dramatically improves reasoning. But the model does not learn to reason — it is steered.

  5. 2024

    DeepSeek-V3 (671B MoE)

    The base model for R1. 671B total parameters with 37B active per token via Mixture of Experts. Multi-head Latent Attention for efficient inference.

  6. 2025

    DeepSeek-R1 (this paper)

    Pure RL unlocks emergent reasoning. GRPO + rule-based rewards produce self-reflection and verification without any human demonstrations. Open-source under MIT.

  7. 2025

    OpenAI o1

    OpenAI's reasoning model, the primary comparison point. Powerful but closed-source and expensive. DeepSeek-R1 matches or exceeds it on most reasoning benchmarks.

CitationDeepSeek-AI et al.. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv / Nature, 2025.

Terms in this paper