Language Models2022intermediate11 min read
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
التحفيز بسلسلة التفكير يستخرج الاستدلال من النماذج اللغوية الكبيرة
Wei, J. · Wang, X. · Schuurmans, D. · Bosma, M. · Ichter, B. · Xia, F. · Chi, E. · Le, Q. · Zhou, D. — NeurIPS
The problem
Large language models could answer simple factual questions via , but consistently failed on tasks requiring multi-step reasoning — arithmetic word problems, commonsense chains, and symbolic manipulation. Scaling model size alone did not fix this: performance curves stayed flat. Meanwhile, training models with step-by-step rationales worked but required expensive labeled datasets for every new task.
The contribution
Chain-of-thought : augment few-shot exemplars with intermediate reasoning steps instead of direct answers. No training, no finetuning — just append "let me think step by step" style demonstrations to the prompt. On the GSM8K math benchmark, PaLM 540B with just 8 chain-of-thought exemplars achieved 56.9% accuracy, surpassing finetuned GPT-3 with a verifier (55%). The method works across arithmetic, commonsense, and tasks, and performance emerges as a function of .
The impact
Chain-of-thought prompting revealed that only provides a lower bound on what language models can do. It launched an entire research direction — , , ReAct, and eventually the reasoning-focused models (o1, Claude's extended thinking) that define the current frontier. The insight that "showing the work" unlocks reasoning became one of the most cited ideas in modern AI.
Imagine a math teacher who writes only the final answer on the board — "42" — and expects students to reverse-engineer the logic. That's standard prompting: you show the model input-output pairs and hope it figures out the reasoning.
Now imagine the teacher writes every step: "first we find the rate, then multiply by time, then subtract the discount…" Students don't just copy — they learn the pattern of thinking and apply it to new problems. That's chain-of-thought prompting: instead of training a model with expensive labeled data, you simply show it a few worked examples, and the model imitates the step-by-step reasoning.
The surprise: this only works once the model is large enough — like needing a student mature enough to follow along. Below ~100B parameters the model generates fluent but nonsensical steps.
The problem: models that can talk but can't think
By 2022, large language models had mastered an impressive range of tasks through few-shot prompting — show a model a few question-answer pairs, and it generalizes to new questions. But this magic had a hard ceiling: multi-step reasoning. A model that could summarize a paragraph or translate a sentence would fail on "Roger has 5 tennis balls. He buys 2 cans with 3 balls each. How many does he have?" — answering confidently and incorrectly.
Two approaches existed but each had a fatal flaw:
-
-augmented training — train models on step-by-step solutions. This works, but requires thousands of manually annotated rationales per task, making it expensive and non-transferable.
-
Standard few-shot prompting — show the model a few input-output pairs. This is cheap and general, but performance on reasoning tasks stayed flat even as models grew from 1B to 100B+ parameters. More capacity didn't mean more reasoning.
The question was: can we get the benefits of step-by-step reasoning without the cost of task-specific training?
The idea: show the work, don't just show the answer
The fix is almost embarrassingly simple. Instead of few-shot exemplars that pair a question with a bare answer, pair each question with a — a series of intermediate natural-language reasoning steps — followed by the answer. The model sees these demonstrations and mimics the pattern: it generates its own chain of thought for new questions, decomposing them into manageable steps before producing a final answer.
Concretely, each exemplar becomes a triple: ⟨input, chain of thought, output⟩. The chain of thought is not a formal proof or a computer program — it is ordinary language that mirrors how a human might talk through the problem. For example:
Q: The cafeteria had 23 apples. They used 20 for lunch and bought 6 more. How many do they have?
A: The cafeteria had 23 apples originally. They used 20 to make lunch. So they had 23 − 20 = 3. They bought 6 more, so they have 3 + 6 = 9. The answer is 9.
No model was finetuned. No gradient was updated. The same frozen model that fails under standard prompting succeeds under chain-of-thought prompting — the only change is the content of the prompt.
The surprise: reasoning emerges only at scale
The most striking finding in the paper is that chain-of-thought prompting is an of model scale. For models with fewer than ~100B parameters, adding chain-of-thought exemplars actually hurts performance — smaller models generate fluent but logically broken reasoning chains that lead to worse answers than simply guessing.
But once models cross the ~100B threshold, chain-of-thought prompting triggers a dramatic upward curve. On GSM8K, PaLM 540B jumps from 17.9% (standard) to 56.9% (chain of thought) — more than tripling accuracy. This pattern repeats across arithmetic, commonsense, and symbolic tasks: flat scaling curve under standard prompting, steep scaling curve under chain-of-thought prompting.
The implication is profound: standard prompting only provides a lower bound on what large language models can do. These models had latent reasoning abilities that standard evaluation completely missed — it took a different prompting strategy to unlock them.
Experiments: arithmetic, commonsense, and symbolic reasoning
The authors tested chain-of-thought prompting across three families of reasoning tasks, using five language models (GPT-3, LaMDA, PaLM, UL2, Codex) at various scales:
— Five math word problem benchmarks (GSM8K, SVAMP, ASDiv, AQuA, MAWPS). On GSM8K, PaLM 540B with chain of thought scored 56.9%, surpassing the prior state of the art set by finetuned GPT-3 with a verifier (55%). On easier one-step problems, gains were smaller — chain of thought helps most where the problems are hardest.
— Five benchmarks (CSQA, StrategyQA, Date Understanding, Sports Understanding, SayCan robot planning). PaLM 540B with chain of thought achieved new state of the art on StrategyQA (75.6% vs 69.4%) and beat an unaided sports enthusiast on Sports Understanding (95.4% vs 84%).
Symbolic reasoning — Two synthetic tasks: last-letter concatenation ("Amy Brown" → "yn") and coin-flip state tracking. Chain of thought enabled near-perfect in-domain performance and, crucially, — solving 4-step problems after seeing only 2-step exemplars.
What actually matters? The ablation study
To understand why chain-of-thought prompting works, the authors tested three ablations — each removing one hypothesized ingredient:
-
Equation only — the model outputs just the math equation, no natural language reasoning. Result: helps on easy problems but fails on GSM8K, because the semantics are too complex to translate directly into an equation without natural language thinking.
-
Variable compute only — the model outputs a sequence of dots ("...") equal in length to the chain of thought, adding extra computation without reasoning content. Result: no improvement over baseline. Extra tokens alone don't help — the content of intermediate steps matters.
-
Chain of thought after answer — the reasoning appears after the final answer instead of before it. Result: no improvement. The model needs to reason before answering. This rules out the hypothesis that chain of thought merely helps the model access relevant knowledge — the sequential reasoning process itself is essential.
Robustness: does it depend on perfect prompts?
A natural concern is sensitivity to the exact wording of the chain-of-thought exemplars. The authors tested this extensively:
-
Different annotators — three co-authors independently wrote chains of thought for the same exemplars. All outperformed standard prompting by a large margin despite notable variance between annotators.
-
Different exemplars — three sets of 8 exemplars randomly sampled from GSM8K's training set (with reasoning chains written by crowd workers, not ML researchers) all substantially outperformed the baseline.
-
Different exemplar orders — standard deviations across random orderings were small for most tasks.
-
Different models — gains held across GPT-3, LaMDA, and PaLM, though the magnitude varied.
The message: chain of thought does not depend on a particular linguistic style, specific exemplars, or a single model architecture. The step-by-step pattern is robust.
When it fails: anatomy of chain-of-thought errors
Chain of thought doesn't always produce correct reasoning. The authors manually examined 50 correct and 50 incorrect outputs from LaMDA 137B on GSM8K:
Of the 50 correct outputs, 49 had genuinely correct reasoning chains (only 1 arrived at the right answer by coincidence).
Of the 50 incorrect outputs, 46% were "almost correct" — the reasoning was sound but had a minor calculator error, symbol mapping mistake, or one missing step. The other 54% had major errors in semantic understanding or coherence.
When PaLM scaled from 62B to 540B, it fixed a large portion of both one-step-missing errors and semantic understanding errors — suggesting that scale helps models both stay on track and grasp meaning more deeply. But there is no guarantee of correct reasoning, and generated chains of thought should never be treated as factually reliable without verification.
The same idea in code
Simplified to show the idea — not the real implementation.
# Chain-of-thought prompting: the entire method in a single prompt
# STANDARD prompting: input → output pairs only
standard_prompt = """
Q: Roger has 5 tennis balls. He buys 2 cans of 3 balls each.
How many does he have now?
A: The answer is 11.
Q: The cafeteria had 23 apples. They used 20 for lunch
and bought 6 more. How many do they have?
A: The answer is 27. ← WRONG! Model jumps to answer.
"""
# CHAIN-OF-THOUGHT prompting: input → reasoning steps → output
cot_prompt = """
Q: Roger has 5 tennis balls. He buys 2 cans of 3 balls each.
How many does he have now?
A: Roger started with 5 balls. 2 cans of 3 is 6 balls.
5 + 6 = 11. The answer is 11.
Q: The cafeteria had 23 apples. They used 20 for lunch
and bought 6 more. How many do they have?
A: They started with 23. Used 20, so 23 - 20 = 3.
Bought 6 more, so 3 + 6 = 9. The answer is 9. ← CORRECT!
"""
# That's it. No training. No finetuning.
# The same frozen model, different prompt, dramatically different results.
# GPT, Claude, Gemini — they all run some form of this idea today.What it unlocked
2022
Chain-of-Thought Prompting
Wei et al. show that 8 worked examples with reasoning steps unlock multi-step reasoning in large language models — no training required. Standard prompting is just a lower bound.
2022
Self-Consistency (Wang et al.)
Sample multiple chains of thought and take a majority vote on the final answer. Improves over single-chain prompting without any additional training.
2022
Zero-Shot CoT — "Let's think step by step"
Kojima et al. discover that simply appending "Let's think step by step" to the prompt — with no exemplars at all — triggers chain-of-thought reasoning in large models.
2023
Tree of Thoughts (Yao et al.)
Generalize chain of thought from a single linear chain to a branching tree of reasoning paths, with lookahead and backtracking for deliberate problem solving.
2023
ReAct (Yao et al.)
Combine chain-of-thought reasoning with actions (web search, API calls). The model reasons about what to do next, acts, observes the result, then reasons again.
2024
Reasoning-native models (o1, extended thinking)
Chain-of-thought reasoning moves from prompt engineering into the model itself. OpenAI's o1 and Claude's extended thinking internalize multi-step reasoning as a core capability.
This paper's deepest contribution is not a technique — it is a lens. Before chain of thought, the assumption was that if a model couldn't solve a task, it lacked the capability. After chain of thought, the question became: "Have we asked the right way?" That shift in perspective — from model limitations to prompting limitations — reshaped how the entire field thinks about what language models can do.
CitationWei, Wang, Schuurmans, Bosma, Ichter, Xia, Chi, Le, Zhou. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. NeurIPS, 2022.
Terms in this paper
- Chain of Thoughtسلسلة التفكير
- Standard Promptingالتحفيز المعياري
- Few-Shot Promptingالتحفيز بأمثلة قليلة
- Emergent Abilityقدرة ناشئة
- Arithmetic Reasoningالاستدلال الحسابي
- Commonsense Reasoningاستدلال الحس السليم
- Symbolic Reasoningالاستدلال الرمزي
- Rationaleالمبرر
- Length Generalizationالتعميم الطولي
- Model Scaleحجم النموذج
- Self-Consistencyالاتّساق الذاتي
- Tree of Thoughtsشجرة الأفكار