Language Models2022intermediate10 min read

Self-Consistency Improves Chain of Thought Reasoning in Language Models

الاتّساق الذاتي يُحسّن الاستدلال بسلسلة التفكير في النماذج اللغوية

Wang, X. · Wei, J. · Schuurmans, D. · Le, Q. · Chi, E. H. · Narang, S. · Chowdhery, A. · Zhou, D. — ICLR

The problem

showed that large language models can solve complex tasks by generating intermediate steps. But the standard approach uses — picking the single most likely at each step — which produces only one . If that path contains an error, the final answer is wrong and there is no mechanism to recover. Greedy decoding also suffers from repetitiveness and local optima, ignoring the fact that most reasoning problems can be solved through multiple valid approaches.

The contribution

A new decoding strategy called that replaces greedy decoding in chain-of-thought prompting. The method samples a diverse set of reasoning paths from the 's , then selects the most consistent answer via majority voting. Self-consistency is entirely unsupervised — it requires no additional , , verifiers, or human annotation. When combined with PaLM-540B or GPT-3, it achieves state-of-the-art results on GSM8K (+17.9%), SVAMP (+11.0%), AQuA (+12.2%), StrategyQA (+6.4%), and ARC-challenge (+3.9%).

The impact

Self-consistency established -then-aggregating as a standard technique for improving language model reasoning. It showed that diversity in reasoning paths is a feature, not noise. The paper directly inspired Tree of Thoughts (which structures exploration further) and influenced the paradigm seen in OpenAI o1 and similar systems. Today, majority voting over sampled reasoning paths is a default tool in every LLM reasoning pipeline.

Imagine you're lost in a foreign city and ask five strangers for directions to the train station. Each person describes a different route — through the park, past the cathedral, along the river. The routes differ, but if four out of five all end at the same building, you can be quite confident that building is the station.

That's self-consistency. Instead of trusting one route (greedy decoding), you collect many routes (sampled reasoning paths) and trust the destination most of them agree on ().

The problem: greedy decoding bets everything on one path

Chain-of-thought (CoT) prompting was a breakthrough: by prompting a language model with step-by-step examples, the model generates its own intermediate reasoning steps before reaching an answer. For a question like "Janet's ducks lay 16 eggs per day. She eats 3 for breakfast and bakes muffins with 4. She sells the rest at $2 each. How much does she earn daily?", the model doesn't just blurt out a number — it reasons through the problem.

But standard CoT uses greedy decoding: at every step, the model picks the single most probable next token. This produces one reasoning chain. If that chain contains a mistake — say it computes 16 − 4 − 3 = 26 instead of 16 − 3 − 4 = 9 — the final answer is wrong, and there is no second chance.

The core insight of this paper is that reasoning problems usually have multiple valid solution paths. A student can solve 16 − 3 − 4 step by step, or compute 3 + 4 = 7 first and then 16 − 7 = 9, or reason about ratios. These paths look different but converge on the same correct answer. Greedy decoding ignores this convergence by committing to exactly one path.

Open in Lab
Compare greedy decoding (one path) with self-consistency (multiple paths → majority vote).
The demo wakes as you arrive…

The method: sample, then vote

Self-consistency works in three simple steps. Step 1 — : give the language model the same chain-of-thought prompt used in standard CoT. Nothing changes here. Step 2 — Sample: instead of greedy decoding, sample multiple independent completions from the model's decoder using sampling, , or . Each sample produces a different reasoning path that may arrive at a different final answer. Step 3 — Aggregate: parse the final answer from each sampled path and take the majority vote — the answer that appears most often wins.

The method is deliberately simple. It needs no additional models, no verifiers, no re-rankers, no fine-tuning, and no human annotations. It works on top of any pre-trained language model as a pure -time strategy.

Open in Lab
The three-step self-consistency pipeline: prompt → sample diverse paths → majority vote.
The demo wakes as you arrive…

Formal description: marginalizing out reasoning paths

Formally, self-consistency introduces a into the generation process. Given a prompt and question, the model generates a pair (rᵢ, aᵢ) where rᵢ is the reasoning path (a sequence of tokens explaining the solution steps) and aᵢ is the final answer parsed from that path. After sampling m such pairs from the decoder, self-consistency marginalizes out the reasoning paths by selecting the most frequent answer.

a∗=arg⁡max⁡a∑i=1m1(ai=a)a^{*} = \arg\max_{a} \sum_{i=1}^{m} \mathbf{1}(a_i = a)
Self-consistency majority vote — The final answer a* is the one that appears most frequently among the m sampled answers. Each reasoning path rᵢ is "marginalized out" — we do not care which path produced the answer, only how many paths agree on it.

The authors also investigated weighting each answer by its generation probability. They found that a normalized weighted sum performs comparably to simple majority vote, because the model assigns similar probabilities to different valid reasoning paths. In other words, the model considers the sampled paths "equally likely," making the unweighted majority vote both simpler and equally effective.

Results: consistent gains across tasks and scales

The authors evaluated self-consistency across four language models (UL2-20B, LaMDA-137B, GPT-3-175B, PaLM-540B) on benchmarks (GSM8K, SVAMP, AQuA, MultiArith, AddSub, ASDiv) and benchmarks (CommonsenseQA, StrategyQA, ARC).

The results are striking. On GSM8K, PaLM-540B improved from 56.5% (greedy CoT) to 74.4% with self-consistency — a gain of nearly 18 percentage points. On AQuA, GPT-3 code-davinci-002 jumped from 39.8% to 52.0%. On StrategyQA, PaLM-540B improved from 75.3% to 81.6%. These gains are larger than those achieved by training additional verifier models, despite self-consistency requiring zero extra training.

Two patterns emerge from the results. First, gains are larger on harder tasks where greedy decoding is lower — there is more room for diverse paths to recover from errors. Second, gains increase with model scale: larger models generate higher- quality reasoning paths, so the diversity is more informative.

Open in Lab
Accuracy gains from self-consistency over greedy CoT across benchmarks (PaLM-540B).
The demo wakes as you arrive…

More paths, better answers

The authors studied how accuracy scales with the number of sampled reasoning paths. The pattern is consistent: accuracy improves rapidly from 1 to 10 paths, with diminishing but still positive returns up to 40 paths. Even just 5 sampled paths delivers a substantial improvement over greedy decoding — making self-consistency practical even when compute is limited.

This curve reveals something important. The gains don't come from a single "lucky" sample. They come from the statistical structure of the problem: as you add more paths, the probability that the majority answer is correct increases monotonically. In practice, 10–20 paths captures most of the benefit.

Open in Lab
Drag the slider to see how accuracy improves as more reasoning paths are sampled.
The demo wakes as you arrive…

The paper carefully compares self-consistency against several alternatives. Sample-and-rank samples multiple sequences and picks the one with the highest log-probability. It helps slightly, but far less than self-consistency, because ranking selects a single path — it doesn't leverage agreement across paths.

is even worse. Beams are diverse by construction, but beam search tends to produce very similar outputs. The diversity that makes self-consistency powerful is exactly what beam search suppresses.

Prompt-order ensembles randomly permute the exemplars in the prompt and take majority vote over greedy decodings. This captures sensitivity to prompt ordering but doesn't diversify the reasoning process itself. Self-consistency with 40 sampled paths outperforms 40 prompt permutations by a wide margin (27.7% vs 19.2% on GSM8K with LaMDA-137B).

Model ensembles (majority vote over outputs from multiple different language models) actually perform worse than self-consistency on a single model, because weaker models in the ensemble drag down the average.

Open in Lab
Self-consistency compared to sample-and-rank, beam search, and prompt ensembles.
The demo wakes as you arrive…

Robustness: it just works

Self-consistency is remarkably robust. It works across different sampling strategies — temperature sampling at T=0.3, 0.5, 0.7, top-k with k=20 or k=40, and nucleus sampling with p=0.9 or p=0.95 — all delivering strong gains. No careful tuning of sampling parameters is needed.

Even when the chain-of-thought prompts contain mistakes (wrong intermediate numbers but correct final answers), self-consistency still improves accuracy. On GSM8K with LaMDA-137B, imperfect prompts with greedy decoding scored 14.9%, but adding self-consistency brought it up to 23.4%.

The authors also showed that self-consistency works with zero-shot chain-of-thought (where the prompt simply says "Let's think step by step" without any examples). On GSM8K with PaLM-540B, zero-shot CoT scored 43.0% — with self-consistency, it jumped to 69.2%, a gain of over 26 points.

Perhaps most usefully, the consistency level (what fraction of sampled paths agree on the final answer) is strongly correlated with accuracy. This means self-consistency can double as an uncertainty estimator: when consistency is low, the model is likely unsure. It gives the model a way to "know when it doesn't know."

Consistency as a confidence signal

Beyond improving accuracy, self-consistency reveals how confident the model is. The authors plotted the relationship between consistency (percentage of sampled paths agreeing on the majority answer) and accuracy (whether that answer is correct). The correlation is remarkably strong: questions where 90% or more of the paths agree are almost always answered correctly, while low-consistency questions are much more likely to be wrong.

This is valuable in practice. A system using self-consistency can flag low-confidence answers for human review, route them to a more capable model, or simply indicate uncertainty to the user. The consistency score is a free byproduct of the method — no additional is needed.

Open in Lab
Higher consistency among sampled paths correlates strongly with answer correctness.
The demo wakes as you arrive…

Legacy: from self-consistency to test-time compute

  1. 2022

    Chain-of-Thought Prompting (CoT)

    Wei et al. showed that prompting language models with step-by-step examples enables multi-step reasoning. The default decoding strategy was greedy.

  2. 2022

    Self-Consistency (this paper)

    Replaced greedy decoding with sample-and-vote. Showed that diversity in reasoning paths is a powerful signal, not noise — establishing the "test-time compute" idea.

  3. 2023

    Tree of Thoughts (ToT)

    Yao et al. extended the idea by structuring exploration as a tree with backtracking, evaluating partial thoughts, and searching for optimal reasoning paths explicitly.

  4. 2024

    OpenAI o1 and test-time compute scaling

    OpenAI o1 operationalized the idea at scale — spending more compute at inference time to reason through problems via hidden chains of thought, echoing the fundamental insight that more reasoning paths yield better answers.

Self-consistency demonstrated a simple but profound principle: you can trade inference-time compute for accuracy. This "test-time compute" paradigm — spending more computation at prediction time rather than at training time — has become one of the most active research directions in modern AI. Every system that samples multiple reasoning chains and aggregates their answers, from coding assistants to mathematical solvers, owes an intellectual debt to this paper.

Self-Consistency in ~15 lines of Pythonpython

Simplified to show the idea — not the real implementation.

from collections import Counter
def self_consistency(model, prompt, question, n_paths=40, temperature=0.7):
    """Sample n_paths reasoning chains, return the majority-vote answer."""
    answers = []
    for _ in range(n_paths):
        # Step 1 & 2: Sample a reasoning path with temperature > 0
        output = model.generate(prompt + question, temperature=temperature)
        # Parse the final answer from the reasoning path
        answer = parse_answer(output)  # e.g., extract text after "The answer is"
        answers.append(answer)
    # Step 3: Majority vote
    vote_counts = Counter(answers)
    best_answer = vote_counts.most_common(1)[0][0]
    confidence = vote_counts[best_answer] / n_paths  # consistency score
    return best_answer, confidence

CitationWang, Wei, Schuurmans, Le, Chi, Narang, Chowdhery, Zhou. Self-Consistency Improves Chain of Thought Reasoning in Language Models. ICLR, 2023.

Terms in this paper