Language Models2022intermediate11 min read

Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

التحفيز بسلسلة التفكير يستخرج الاستدلال من النماذج اللغوية الكبيرة

Wei, J. · Wang, X. · Schuurmans, D. · Bosma, M. · Ichter, B. · Xia, F. · Chi, E. · Le, Q. · Zhou, D. — NeurIPS

The problem

Large language models could answer simple factual questions via , but consistently failed on tasks requiring multi-step reasoning — arithmetic word problems, commonsense chains, and symbolic manipulation. Scaling model size alone did not fix this: performance curves stayed flat. Meanwhile, training models with step-by-step rationales worked but required expensive labeled datasets for every new task.

The contribution

Chain-of-thought : augment few-shot exemplars with intermediate reasoning steps instead of direct answers. No training, no finetuning — just append "let me think step by step" style demonstrations to the prompt. On the GSM8K math benchmark, PaLM 540B with just 8 chain-of-thought exemplars achieved 56.9% accuracy, surpassing finetuned GPT-3 with a verifier (55%). The method works across arithmetic, commonsense, and tasks, and performance emerges as a function of .

The impact

Chain-of-thought prompting revealed that only provides a lower bound on what language models can do. It launched an entire research direction — , , ReAct, and eventually the reasoning-focused models (o1, Claude's extended thinking) that define the current frontier. The insight that "showing the work" unlocks reasoning became one of the most cited ideas in modern AI.

Imagine a math teacher who writes only the final answer on the board — "42" — and expects students to reverse-engineer the logic. That's standard prompting: you show the model input-output pairs and hope it figures out the reasoning.

Now imagine the teacher writes every step: "first we find the rate, then multiply by time, then subtract the discount…" Students don't just copy — they learn the pattern of thinking and apply it to new problems. That's chain-of-thought prompting: instead of training a model with expensive labeled data, you simply show it a few worked examples, and the model imitates the step-by-step reasoning.

The surprise: this only works once the model is large enough — like needing a student mature enough to follow along. Below ~100B parameters the model generates fluent but nonsensical steps.

The problem: models that can talk but can't think

By 2022, large language models had mastered an impressive range of tasks through few-shot prompting — show a model a few question-answer pairs, and it generalizes to new questions. But this magic had a hard ceiling: multi-step reasoning. A model that could summarize a paragraph or translate a sentence would fail on "Roger has 5 tennis balls. He buys 2 cans with 3 balls each. How many does he have?" — answering confidently and incorrectly.

Two approaches existed but each had a fatal flaw:

  • -augmented training — train models on step-by-step solutions. This works, but requires thousands of manually annotated rationales per task, making it expensive and non-transferable.

  • Standard few-shot prompting — show the model a few input-output pairs. This is cheap and general, but performance on reasoning tasks stayed flat even as models grew from 1B to 100B+ parameters. More capacity didn't mean more reasoning.

The question was: can we get the benefits of step-by-step reasoning without the cost of task-specific training?

Open in Lab
Compare standard prompting (left) with chain-of-thought prompting (right). Notice how the model jumps to the wrong answer without intermediate steps.
The demo wakes as you arrive…

The idea: show the work, don't just show the answer

The fix is almost embarrassingly simple. Instead of few-shot exemplars that pair a question with a bare answer, pair each question with a — a series of intermediate natural-language reasoning steps — followed by the answer. The model sees these demonstrations and mimics the pattern: it generates its own chain of thought for new questions, decomposing them into manageable steps before producing a final answer.

Concretely, each exemplar becomes a triple: ⟨input, chain of thought, output⟩. The chain of thought is not a formal proof or a computer program — it is ordinary language that mirrors how a human might talk through the problem. For example:

Q: The cafeteria had 23 apples. They used 20 for lunch and bought 6 more. How many do they have?

A: The cafeteria had 23 apples originally. They used 20 to make lunch. So they had 23 − 20 = 3. They bought 6 more, so they have 3 + 6 = 9. The answer is 9.

No model was finetuned. No gradient was updated. The same frozen model that fails under standard prompting succeeds under chain-of-thought prompting — the only change is the content of the prompt.

Open in Lab
Step through a chain-of-thought reasoning process. Click each stage to see how the model decomposes the problem.
The demo wakes as you arrive…

The surprise: reasoning emerges only at scale

The most striking finding in the paper is that chain-of-thought prompting is an of model scale. For models with fewer than ~100B parameters, adding chain-of-thought exemplars actually hurts performance — smaller models generate fluent but logically broken reasoning chains that lead to worse answers than simply guessing.

But once models cross the ~100B threshold, chain-of-thought prompting triggers a dramatic upward curve. On GSM8K, PaLM 540B jumps from 17.9% (standard) to 56.9% (chain of thought) — more than tripling accuracy. This pattern repeats across arithmetic, commonsense, and symbolic tasks: flat scaling curve under standard prompting, steep scaling curve under chain-of-thought prompting.

The implication is profound: standard prompting only provides a lower bound on what large language models can do. These models had latent reasoning abilities that standard evaluation completely missed — it took a different prompting strategy to unlock them.

Open in Lab
Drag the model-size slider and watch the gap widen. Chain-of-thought prompting only begins helping beyond ~100B parameters.
The demo wakes as you arrive…

Experiments: arithmetic, commonsense, and symbolic reasoning

The authors tested chain-of-thought prompting across three families of reasoning tasks, using five language models (GPT-3, LaMDA, PaLM, UL2, Codex) at various scales:

— Five math word problem benchmarks (GSM8K, SVAMP, ASDiv, AQuA, MAWPS). On GSM8K, PaLM 540B with chain of thought scored 56.9%, surpassing the prior state of the art set by finetuned GPT-3 with a verifier (55%). On easier one-step problems, gains were smaller — chain of thought helps most where the problems are hardest.

— Five benchmarks (CSQA, StrategyQA, Date Understanding, Sports Understanding, SayCan robot planning). PaLM 540B with chain of thought achieved new state of the art on StrategyQA (75.6% vs 69.4%) and beat an unaided sports enthusiast on Sports Understanding (95.4% vs 84%).

Symbolic reasoning — Two synthetic tasks: last-letter concatenation ("Amy Brown" → "yn") and coin-flip state tracking. Chain of thought enabled near-perfect in-domain performance and, crucially, — solving 4-step problems after seeing only 2-step exemplars.

Open in Lab
Explore results across all three reasoning domains. Click a task category to see the performance breakdown.
The demo wakes as you arrive…

What actually matters? The ablation study

To understand why chain-of-thought prompting works, the authors tested three ablations — each removing one hypothesized ingredient:

  • Equation only — the model outputs just the math equation, no natural language reasoning. Result: helps on easy problems but fails on GSM8K, because the semantics are too complex to translate directly into an equation without natural language thinking.

  • Variable compute only — the model outputs a sequence of dots ("...") equal in length to the chain of thought, adding extra computation without reasoning content. Result: no improvement over baseline. Extra tokens alone don't help — the content of intermediate steps matters.

  • Chain of thought after answer — the reasoning appears after the final answer instead of before it. Result: no improvement. The model needs to reason before answering. This rules out the hypothesis that chain of thought merely helps the model access relevant knowledge — the sequential reasoning process itself is essential.

Open in Lab
Compare chain-of-thought prompting against the three ablations. Each bar shows the solve rate on GSM8K.
The demo wakes as you arrive…

Robustness: does it depend on perfect prompts?

A natural concern is sensitivity to the exact wording of the chain-of-thought exemplars. The authors tested this extensively:

  • Different annotators — three co-authors independently wrote chains of thought for the same exemplars. All outperformed standard prompting by a large margin despite notable variance between annotators.

  • Different exemplars — three sets of 8 exemplars randomly sampled from GSM8K's training set (with reasoning chains written by crowd workers, not ML researchers) all substantially outperformed the baseline.

  • Different exemplar orders — standard deviations across random orderings were small for most tasks.

  • Different models — gains held across GPT-3, LaMDA, and PaLM, though the magnitude varied.

The message: chain of thought does not depend on a particular linguistic style, specific exemplars, or a single model architecture. The step-by-step pattern is robust.

When it fails: anatomy of chain-of-thought errors

Chain of thought doesn't always produce correct reasoning. The authors manually examined 50 correct and 50 incorrect outputs from LaMDA 137B on GSM8K:

Of the 50 correct outputs, 49 had genuinely correct reasoning chains (only 1 arrived at the right answer by coincidence).

Of the 50 incorrect outputs, 46% were "almost correct" — the reasoning was sound but had a minor calculator error, symbol mapping mistake, or one missing step. The other 54% had major errors in semantic understanding or coherence.

When PaLM scaled from 62B to 540B, it fixed a large portion of both one-step-missing errors and semantic understanding errors — suggesting that scale helps models both stay on track and grasp meaning more deeply. But there is no guarantee of correct reasoning, and generated chains of thought should never be treated as factually reliable without verification.

The same idea in code

Chain-of-thought prompting — the complete patternpython

Simplified to show the idea — not the real implementation.

# Chain-of-thought prompting: the entire method in a single prompt

# STANDARD prompting: input → output pairs only
standard_prompt = """
Q: Roger has 5 tennis balls. He buys 2 cans of 3 balls each.
   How many does he have now?
A: The answer is 11.

Q: The cafeteria had 23 apples. They used 20 for lunch
   and bought 6 more. How many do they have?
A: The answer is 27.  ← WRONG! Model jumps to answer.
"""

# CHAIN-OF-THOUGHT prompting: input → reasoning steps → output
cot_prompt = """
Q: Roger has 5 tennis balls. He buys 2 cans of 3 balls each.
   How many does he have now?
A: Roger started with 5 balls. 2 cans of 3 is 6 balls.
   5 + 6 = 11. The answer is 11.

Q: The cafeteria had 23 apples. They used 20 for lunch
   and bought 6 more. How many do they have?
A: They started with 23. Used 20, so 23 - 20 = 3.
   Bought 6 more, so 3 + 6 = 9. The answer is 9.  ← CORRECT!
"""

# That's it. No training. No finetuning.
# The same frozen model, different prompt, dramatically different results.
# GPT, Claude, Gemini — they all run some form of this idea today.

What it unlocked

  1. 2022

    Chain-of-Thought Prompting

    Wei et al. show that 8 worked examples with reasoning steps unlock multi-step reasoning in large language models — no training required. Standard prompting is just a lower bound.

  2. 2022

    Self-Consistency (Wang et al.)

    Sample multiple chains of thought and take a majority vote on the final answer. Improves over single-chain prompting without any additional training.

  3. 2022

    Zero-Shot CoT — "Let's think step by step"

    Kojima et al. discover that simply appending "Let's think step by step" to the prompt — with no exemplars at all — triggers chain-of-thought reasoning in large models.

  4. 2023

    Tree of Thoughts (Yao et al.)

    Generalize chain of thought from a single linear chain to a branching tree of reasoning paths, with lookahead and backtracking for deliberate problem solving.

  5. 2023

    ReAct (Yao et al.)

    Combine chain-of-thought reasoning with actions (web search, API calls). The model reasons about what to do next, acts, observes the result, then reasons again.

  6. 2024

    Reasoning-native models (o1, extended thinking)

    Chain-of-thought reasoning moves from prompt engineering into the model itself. OpenAI's o1 and Claude's extended thinking internalize multi-step reasoning as a core capability.

This paper's deepest contribution is not a technique — it is a lens. Before chain of thought, the assumption was that if a model couldn't solve a task, it lacked the capability. After chain of thought, the question became: "Have we asked the right way?" That shift in perspective — from model limitations to prompting limitations — reshaped how the entire field thinks about what language models can do.

CitationWei, Wang, Schuurmans, Bosma, Ichter, Xia, Chi, Le, Zhou. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. NeurIPS, 2022.

Terms in this paper