Language Models2025advanced11 min read
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
DeepSeek-R1: تحفيز قدرات الاستدلال في النماذج اللغوية الكبيرة عبر التعلم المعزز
DeepSeek-AI · Guo, D. · Yang, D. · Zhang, H. · Song, J. · Wang, P. · Zhu, Q. · Xu, R. · Zhang, R. · Ma, S. · Bi, X. — arXiv / Nature
The problem
Teaching LLMs to reason typically requires massive human-annotated step-by-step solutions (supervised ). This is expensive, it does not scale, and it caps the at human-level reasoning — the model can never discover strategies the annotators did not think of. Moreover, neural reward models used in RLHF are susceptible to as scales up.
The contribution
DeepSeek-R1-Zero: a trained with pure — no supervised fine-tuning at all. Using (GRPO) and rule-based rewards (accuracy + format), the model spontaneously develops sophisticated reasoning behaviors: , verification, and dynamic strategy switching. DeepSeek-R1 builds on this with a multi-stage pipeline ( → RL → → SFT → RL) to fix readability and language mixing while preserving reasoning power. The reasoning patterns can be distilled into smaller models (1.5B–70B).
The impact
DeepSeek-R1 proved that reinforcement learning alone can unlock reasoning capabilities rivaling OpenAI's o1 — without relying on expensive human annotations. It achieved 79.8% on AIME 2024, 97.3% on MATH-500, and outperformed 96.3% of human competitors on Codeforces. The open-source release under MIT license democratized access to frontier reasoning models and sparked a wave of RL-for-reasoning research worldwide.
Traditional LLM training is like a driving school: a human instructor sits beside the student, demonstrating every turn and lane change. The student can never be better than the instructor.
DeepSeek-R1's approach is like a racing simulator: give the student a car, a track, and a lap timer. Say nothing about technique. After thousands of laps, the student discovers racing lines, braking points, and overtaking strategies the instructor never thought of — because the only feedback is the clock.
The "aha moment" is when the driver, mid-lap, slams the brakes, mutters "wait, that corner was wrong", reverses mental course, and finds a faster line — all on their own.
The problem: human demonstrations cap reasoning
Before DeepSeek-R1, the standard recipe for teaching LLMs to reason was a two-stage pipeline: first, supervised fine-tuning (SFT) on thousands of human-written step-by-step solutions, then (RLHF) to polish outputs. This approach has three fundamental limitations:
-
Scalability wall. High-quality reasoning traces are expensive to produce. Human experts spend minutes per problem, and the resulting data cannot cover the vast space of possible reasoning strategies.
-
Ceiling effect. The model learns to imitate human reasoning, so it can never surpass the quality of its training data. Novel strategies — like spending extra tokens to verify an answer, or spontaneously switching approaches — are never demonstrated and therefore never learned.
-
Reward hacking. Neural reward models used in RLHF can be exploited: the model finds shortcuts that score highly without genuinely reasoning. Retraining the reward model requires more compute and human effort, creating a vicious cycle.
The engine: Group Relative Policy Optimization (GRPO)
DeepSeek-R1-Zero starts from a pre-trained base model (DeepSeek-V3-Base, 671B parameters with 37B active via ) and applies reinforcement learning directly — no supervised fine-tuning. The RL algorithm is GRPO (Group Relative Policy Optimization), which was designed to simplify PPO by removing the need for a separate value model.
Here is the intuition: for each question, GRPO generates a group of candidate answers (typically 16). Some answers are correct, others wrong. Instead of needing a learned to estimate how good a partial answer is, GRPO simply compares each answer to its group-mates: if your reward is above the group average, you are reinforced; if below, you are penalized. The of each output is computed relative to its group — hence "Group Relative."
Think of it as a class ranking: instead of an external grader (the value model in PPO), the students grade themselves relative to each other. The best-performing solution in each batch gets the strongest positive signal.
Reward design: simple rules, no neural judges
A critical design choice in DeepSeek-R1-Zero is the deliberate avoidance of neural reward models for reasoning tasks. Neural reward models are prone to reward hacking — the model learns to exploit blind spots in the judge rather than genuinely improve. Retraining the reward model consumes computational resources and adds pipeline complexity.
Instead, DeepSeek-R1-Zero uses rule-based rewards with two components:
-
Accuracy reward: did the model get the right answer? For math, the final answer is extracted and compared against ground truth. For code, a compiler runs the solution against test cases. Binary: 1 if correct, 0 if wrong.
-
Format reward: did the model wrap its reasoning in the expected
<think>...</think>and<answer>...</answer>tags? This ensures the reasoning process is explicitly delineated, making it interpretable and analyzable.
The total reward is simply the sum of these two signals. No learned reward model, no human preference data, no reward-model retraining loop. The simplicity is the point: it eliminates the reward hacking failure mode that plagues neural approaches.
Emergent behaviors: the model teaches itself to think
The most striking finding of DeepSeek-R1-Zero is that sophisticated reasoning strategies emerge spontaneously from pure RL — with no human demonstration. As training progresses, three phenomena unfold:
-
Self-extending thinking time. The model's average response length grows steadily from ~3,000 tokens to over 10,000 tokens. Nobody told it to think longer; it discovered on its own that spending more tokens on reasoning leads to higher rewards.
-
Reflective reasoning. Words like "wait", "mistake", "however", "verify", and "retry" increase 5–7× during training. The model learns to pause, question its own logic, and correct course — a form of metacognition that was never explicitly taught.
-
The "aha moment." Around training step 8,000, the model suddenly begins using phrases like "Wait, wait. That's an aha moment" — rethinking its approach mid-solution. The frequency of the word "wait" spikes dramatically, marking a phase transition in reasoning strategy. This moment is not programmed; it emerges from the reward signal alone.
AIME 2024 accuracy climbs from 15.6% to 77.9% during training. With majority voting over 16 samples, it reaches 86.7% — surpassing the average human competitor score.
From R1-Zero to R1: a multi-stage refinement
DeepSeek-R1-Zero proves that pure RL works, but it has practical problems: the reasoning text is hard to read, and the model mixes Chinese and English unpredictably within a single chain-of-thought. It also cannot handle non-reasoning tasks like creative writing or open-domain QA well.
DeepSeek-R1 fixes these issues with a four-stage pipeline:
-
Stage 1 — Cold start SFT. A small dataset of thousands of readable, first-person reasoning traces is collected. Human annotators rewrite R1-Zero outputs into a conversational style ("I need to think about..." rather than "We observe that..."), then an LLM rewrites more data in this style. This is used for a brief SFT phase that gives the RL a better starting point.
-
Stage 2 — First RL with language consistency. Reasoning-focused RL training, just like R1-Zero, but with an added language consistency reward: the proportion of target-language words in the chain-of-thought. This nudges the model to stay in one language throughout its reasoning.
-
Stage 3 — Rejection sampling + SFT. The RL checkpoint generates solutions; only correct, well-formatted ones are kept. About 600K reasoning samples plus 200K non-reasoning samples (writing, QA, translation) form a combined SFT dataset.
-
Stage 4 — Second RL. A final round of RL using both rule-based rewards (reasoning) and model-based rewards (helpfulness/safety), creating a model that reasons well and follows instructions.
Distillation: reasoning in smaller packages
A large model that reasons brilliantly is useful, but a small model that reasons well is transformative — it can run on a single , on-device, or at the edge. DeepSeek distilled R1's reasoning capabilities into six smaller models ranging from 1.5B to 70B parameters, built on Qwen and Llama base architectures.
The process is straightforward: fine-tune the smaller base model on the same ~800K high-quality reasoning trajectories generated by R1. No RL is needed for the student — the SFT data already encodes the reasoning patterns that RL discovered.
The key finding: distilled reasoning outperforms RL-from-scratch on small models. When researchers tried applying RL directly to Qwen-32B (the same approach that worked for the 671B model), the results were worse than simply distilling from R1. This suggests that discovering advanced reasoning strategies requires scale, but once discovered, those strategies can be efficiently transferred to smaller architectures.
The same idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
def compute_grpo_advantages(rewards: list[float]) -> list[float]:
"""Compute group-relative advantages for a batch of outputs.
Each output's reward is z-scored within its group:
advantage_i = (r_i - mean(group)) / std(group)
If your reward is above the group average, you get reinforced.
If below, you get penalized. No value model needed.
"""
rewards = np.array(rewards)
mean = rewards.mean()
std = rewards.std() + 1e-8 # avoid division by zero
return ((rewards - mean) / std).tolist()
def rule_based_reward(answer: str, ground_truth: str, has_tags: bool) -> float:
"""DeepSeek-R1-Zero's entire reward function.
No neural reward model — just correctness + format.
"""
accuracy = 1.0 if answer.strip() == ground_truth.strip() else 0.0
format_ok = 1.0 if has_tags else 0.0
return accuracy + format_ok
# Example: 4 sampled outputs for one math question
rewards = [
rule_based_reward("42", "42", True), # correct + formatted → 2.0
rule_based_reward("42", "42", False), # correct but no tags → 1.0
rule_based_reward("17", "42", True), # wrong + formatted → 1.0
rule_based_reward("99", "42", False), # wrong + no tags → 0.0
]
advantages = compute_grpo_advantages(rewards)
# [1.34, 0.0, 0.0, -1.34] — the correct+formatted answer
# gets the strongest positive signal; the wrong+unformatted
# gets the strongest negative signal.Results and impact
DeepSeek-R1 achieves performance comparable to OpenAI's o1-1217 across reasoning benchmarks while being fully open-source (MIT license). On mathematics, it scores 79.8% on AIME 2024 (vs o1's 79.2%) and 97.3% on MATH-500 (vs o1's 96.4%). On coding, it reaches a Codeforces rating of 2029, placing it in the 96.3rd percentile of human competitors. On knowledge benchmarks like MMLU-Pro, it achieves 84.0%.
The distilled models are equally impressive: DeepSeek-R1-Distill-Qwen-32B surpasses OpenAI-o1-mini on most reasoning benchmarks, despite being dramatically smaller. The 7B distilled model outperforms many non-reasoning models ten times its size.
The training cost is remarkably efficient: approximately 147K H800 GPU hours total (~$294K USD), a fraction of what comparable closed-source reasoning models require.
2017
Transformer (Attention Is All You Need)
The architecture that made everything possible — self-attention replaces recurrence, enabling massive parallelism and the scaling era.
2017
PPO (Proximal Policy Optimization)
The RL algorithm that became standard for RLHF — stable policy updates via clipped surrogate objective. Requires a learned value model.
2022
InstructGPT / RLHF
SFT then RLHF became the standard post-training recipe. Powerful but dependent on human demonstrations and neural reward models.
2022
Chain-of-Thought Prompting
Wei et al. showed that prompting LLMs with "Let's think step by step" dramatically improves reasoning. But the model does not learn to reason — it is steered.
2024
DeepSeek-V3 (671B MoE)
The base model for R1. 671B total parameters with 37B active per token via Mixture of Experts. Multi-head Latent Attention for efficient inference.
2025
DeepSeek-R1 (this paper)
Pure RL unlocks emergent reasoning. GRPO + rule-based rewards produce self-reflection and verification without any human demonstrations. Open-source under MIT.
2025
OpenAI o1
OpenAI's reasoning model, the primary comparison point. Powerful but closed-source and expensive. DeepSeek-R1 matches or exceeds it on most reasoning benchmarks.
CitationDeepSeek-AI et al.. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv / Nature, 2025.
Terms in this paper
- Group Relative Policy Optimizationالتحسين التجميعي النسبي للسياسة
- Rule-Based Rewardمكافأة مبنية على القواعد
- Cold Startالبداية الباردة
- Rejection Samplingالترشيح بالرفض
- Emergent Behaviorالسلوك الناشئ
- Self-Reflectionالمراجعة الذاتية
- Distillationالتقطير
- Reward Hackingاختراق المكافأة
- Overthinkingالإفراط في التفكير