Language Models2024intermediate10 min read
Learning to Reason with LLMs
كيف تتعلّم النماذج اللغوية الكبيرة الاستدلال
OpenAI — OpenAI Blog
The problem
Standard LLMs generate answers in a single — fast but shallow. On tasks requiring multi-step reasoning — competition mathematics, algorithmic coding, PhD-level science — even the best models (GPT-4o, Claude 3.5 Sonnet) plateau far below human experts. Scaling pre-training data and parameters alone no longer yields proportional gains on reasoning benchmarks. A fundamentally different scaling axis was needed.
The contribution
OpenAI o1: a trained with large-scale to produce an internal before answering. Through RL, the model learns to recognize and correct its own mistakes, break hard steps into simpler ones, and try different approaches when stuck. Performance scales smoothly with both (more RL) and (more thinking). On AIME 2024, o1 scored 83.3% (cons@64) vs GPT-4o's 13.4%. On GPQA Diamond, o1 surpassed human PhD-level experts for the first time. On Codeforces, it reached the 89th percentile.
The impact
o1 introduced a new scaling paradigm: instead of only making models bigger, make them think longer. This "test-time compute" axis opened a second dimension of improvement independent of pre-training scale. The approach spawned a family of reasoning models — DeepSeek-R1, QwQ, and others — and established chain-of-thought reinforcement learning as a core training technique. It also demonstrated that readable internal reasoning chains create new opportunities for AI safety and monitoring.
A standard LLM answering a hard question is like a chess grandmaster forced to move instantly — no thinking time allowed. They'll often choose a plausible move, but on complex positions they blunder.
o1 gives the grandmaster a clock: the model can sit, think, explore lines, back up from dead ends, and only then commit to an answer. The longer the clock runs, the better the move. That thinking happens in a — a scratchpad the model writes to itself before speaking.
The problem: fast answers hit a ceiling
Language models like GPT-4o generate text by token in a single left-to-right pass. This is Daniel Kahneman's System 1: fast, intuitive, automatic. For everyday questions it works beautifully, but for problems requiring deep reasoning — a multi-step proof, a tricky cipher, an algorithmic puzzle — the model must commit to each token the instant it generates it, with no opportunity to plan ahead, reconsider, or backtrack.
The result: on competition math (AIME 2024), GPT-4o solved only 12% of problems. On PhD-level science questions (GPQA Diamond), it scored around 50%. Scaling parameters and data further was hitting diminishing returns on these reasoning-heavy benchmarks.
The idea: teach the model to think before answering
The key insight behind o1 is shifting from System 1 to System 2 thinking. Instead of answering immediately, the model first writes an extended chain of thought — an internal reasoning trace that can span hundreds or thousands of tokens. Through reinforcement learning, the model discovers effective thinking strategies on its own:
- Recognizing and correcting mistakes — the model catches its own errors mid-reasoning and backtracks.
- Breaking hard steps into simpler ones — decomposing a complex problem into manageable sub-problems.
- Trying different approaches — when one strategy fails, the model pivots to another.
Crucially, nobody hand-wrote these strategies. The RL algorithm — trained on reasoning tasks where the answer can be verified — taught the model to discover them. This is highly data-efficient: the model learns how to think, not just what to know.
Chain of thought in action
To see what chain-of-thought reasoning looks like, consider a cipher puzzle. GPT-4o sees the encoded message and immediately tries to identify a pattern — but guesses wrong, gives up, and asks the user for help.
o1 approaches the same problem differently. It thinks for 5 seconds, systematically testing hypotheses: Are the ciphertext words twice as long as the plaintext? Do letter pairs average to the answer? It discovers the encoding rule — each plaintext letter is the average of two ciphertext letters — verifies it against the example, then applies it to decode: "THERE ARE THREE R'S IN STRAWBERRY".
This is a task GPT-4o could not solve at all. The chain of thought gave o1 the working memory to try, verify, and apply a decoding strategy — the same process a human puzzle-solver would follow.
How well does it work?
o1 was evaluated on a diverse set of reasoning-heavy benchmarks. The improvements over GPT-4o are dramatic:
- Competition Math (AIME 2024): GPT-4o solved 12% of problems. o1 solved 83.3% with — placing it among the top 500 US students and above the USAMO cutoff.
- PhD-Level Science (GPQA Diamond): o1 scored 78% on questions in chemistry, physics, and biology — surpassing human PhD experts for the first time on this benchmark.
- Competition Code (Codeforces): o1 reached the 89th percentile (Elo 1673), up from GPT-4o's 11th percentile. A further fine-tuned version for IOI reached the 93rd percentile.
- MMMU (vision + reasoning): o1 scored 78.2%, the first model competitive with human experts on this multimodal reasoning benchmark.
However, o1 is not universally better. On some natural language tasks — where fast, intuitive responses suffice — human evaluators preferred GPT-4o. The lesson: is powerful but not always necessary.
Safety: thinking out loud helps alignment
Chain of thought reasoning creates a new opportunity for safety. When a model writes out its reasoning step by step, we can read its thought process — like having a window into the model's mind. OpenAI integrated safety policies directly into the chain of thought, teaching o1 to reason about safety rules in context.
The results were significant: on evaluations, o1-preview scored 93.4% safe completions on challenging adversarial prompts, compared to 71.4% for GPT-4o. On the StrongREJECT benchmark, o1-preview achieved a goodness score of 0.84 versus 0.22 for GPT-4o. Importantly, this safety improvement did not come at the cost of — o1 was also more compliant on benign edge cases.
Hidden chains of thought
OpenAI chose not to show the raw chain of thought to users, instead displaying a model-generated summary. This decision balances three considerations:
- Safety monitoring. A hidden chain of thought lets researchers read the model's unfiltered reasoning and watch for signs of manipulation, deception, or misalignment — without the model learning to censor itself for the audience.
- Reasoning freedom. If the model knows its thoughts are shown to users, it may learn to perform for the audience rather than reason honestly. Keeping thoughts private preserves the integrity of the reasoning process.
- Competitive considerations. The chain of thought contains details about the model's learned strategies that represent significant research investment.
The tradeoff is transparency: users cannot verify the reasoning behind an answer. To partially compensate, o1 is trained to reproduce useful ideas from the hidden chain of thought in its final visible answer.
The reinforcement learning training loop
How does the model learn to reason? The core training loop has three parts. First, the model receives a problem — a math question, a coding challenge, a science problem — where the answer can be objectively verified. Second, the model generates a long chain of thought followed by a final answer. Third, the answer is checked: correct answers receive a positive reward signal, wrong answers receive a negative one.
Over millions of such episodes, the RL algorithm shapes the model's chain-of-thought behavior. Strategies that lead to correct answers — careful decomposition, error checking, trying alternatives — get reinforced. Strategies that lead to wrong answers — rushing, skipping verification, committing to a bad approach — get suppressed.
Importantly, OpenAI reports that this process is highly data-efficient. The model does not need trillions of tokens of reasoning examples. Instead, it learns how to reason from a comparatively small set of verifiable problems — and generalizes this skill to new domains.
The code perspective
Simplified to show the idea — not the real implementation.
def rl_reasoning_step(model, problem, verifier):
"""One step of RL training for chain-of-thought reasoning."""
# 1. Model generates chain of thought + final answer
chain_of_thought = model.generate_reasoning(problem)
final_answer = model.extract_answer(chain_of_thought)
# 2. Verifier checks if answer is correct
is_correct = verifier.check(problem, final_answer)
# 3. Reward signal: +1 for correct, -1 for wrong
reward = +1.0 if is_correct else -1.0
# 4. RL update: reinforce reasoning strategies that led to
# correct answers, suppress strategies that led to wrong ones
model.update_policy(chain_of_thought, reward)
return reward
# After millions of episodes, the model learns:
# - to break hard problems into simpler sub-steps
# - to check its own work and catch mistakes
# - to try a different approach when stuck
# These strategies emerge from RL, not from hand-written rules.What it means for the field
2022
Chain-of-Thought Prompting (Wei et al.)
Demonstrated that prompting LLMs to "think step by step" dramatically improves reasoning performance — but relied on prompt engineering, not learned behavior.
2024
OpenAI o1
First model trained via RL to produce extended chains of thought autonomously. Introduced the test-time compute scaling axis. Surpassed human experts on GPQA Diamond.
2025
DeepSeek-R1
Open-weights reasoning model that replicated o1-style chain-of-thought RL training. Demonstrated the approach is not architecture-specific but a general training paradigm.
2025
o1-pro, o3, o4-mini
Successive models that pushed reasoning further, with o3 achieving near-perfect scores on ARC-AGI and Frontier Math — benchmarks designed to resist scaling.
The path from chain-of-thought prompting to o1 mirrors a recurring pattern in AI: a technique that starts as a clever trick () gets baked into the model itself through training. What users once had to manually prompt, the model now does spontaneously — and better, because it was optimized end-to-end for it.
CitationOpenAI. Learning to Reason with LLMs. OpenAI Blog, 2024.
Terms in this paper
- Chain of Thoughtسلسلة التفكير
- Test-Time Computeحوسبة وقت الاستدلال
- Train-Time Computeحوسبة وقت التدريب
- Hidden Chain of Thoughtسلسلة التفكير المخفية
- Reward Hackingاختراق المكافأة
- Consensus Samplingأخذ العينات بالإجماع
- System 2 Thinkingتفكير النظام الثاني