Language Models2022intermediate10 min read
Self-Consistency Improves Chain of Thought Reasoning in Language Models
الاتّساق الذاتي يُحسّن الاستدلال بسلسلة التفكير في النماذج اللغوية
Wang, X. · Wei, J. · Schuurmans, D. · Le, Q. · Chi, E. H. · Narang, S. · Chowdhery, A. · Zhou, D. — ICLR
The problem
showed that large language models can solve complex tasks by generating intermediate steps. But the standard approach uses — picking the single most likely at each step — which produces only one . If that path contains an error, the final answer is wrong and there is no mechanism to recover. Greedy decoding also suffers from repetitiveness and local optima, ignoring the fact that most reasoning problems can be solved through multiple valid approaches.
The contribution
A new decoding strategy called that replaces greedy decoding in chain-of-thought prompting. The method samples a diverse set of reasoning paths from the 's , then selects the most consistent answer via majority voting. Self-consistency is entirely unsupervised — it requires no additional , , verifiers, or human annotation. When combined with PaLM-540B or GPT-3, it achieves state-of-the-art results on GSM8K (+17.9%), SVAMP (+11.0%), AQuA (+12.2%), StrategyQA (+6.4%), and ARC-challenge (+3.9%).
The impact
Self-consistency established -then-aggregating as a standard technique for improving language model reasoning. It showed that diversity in reasoning paths is a feature, not noise. The paper directly inspired Tree of Thoughts (which structures exploration further) and influenced the paradigm seen in OpenAI o1 and similar systems. Today, majority voting over sampled reasoning paths is a default tool in every LLM reasoning pipeline.
Imagine you're lost in a foreign city and ask five strangers for directions to the train station. Each person describes a different route — through the park, past the cathedral, along the river. The routes differ, but if four out of five all end at the same building, you can be quite confident that building is the station.
That's self-consistency. Instead of trusting one route (greedy decoding), you collect many routes (sampled reasoning paths) and trust the destination most of them agree on ().
The problem: greedy decoding bets everything on one path
Chain-of-thought (CoT) prompting was a breakthrough: by prompting a language model with step-by-step examples, the model generates its own intermediate reasoning steps before reaching an answer. For a question like "Janet's ducks lay 16 eggs per day. She eats 3 for breakfast and bakes muffins with 4. She sells the rest at $2 each. How much does she earn daily?", the model doesn't just blurt out a number — it reasons through the problem.
But standard CoT uses greedy decoding: at every step, the model picks the single most probable next token. This produces one reasoning chain. If that chain contains a mistake — say it computes 16 − 4 − 3 = 26 instead of 16 − 3 − 4 = 9 — the final answer is wrong, and there is no second chance.
The core insight of this paper is that reasoning problems usually have multiple valid solution paths. A student can solve 16 − 3 − 4 step by step, or compute 3 + 4 = 7 first and then 16 − 7 = 9, or reason about ratios. These paths look different but converge on the same correct answer. Greedy decoding ignores this convergence by committing to exactly one path.
The method: sample, then vote
Self-consistency works in three simple steps. Step 1 — : give the language model the same chain-of-thought prompt used in standard CoT. Nothing changes here. Step 2 — Sample: instead of greedy decoding, sample multiple independent completions from the model's decoder using sampling, , or . Each sample produces a different reasoning path that may arrive at a different final answer. Step 3 — Aggregate: parse the final answer from each sampled path and take the majority vote — the answer that appears most often wins.
The method is deliberately simple. It needs no additional models, no verifiers, no re-rankers, no fine-tuning, and no human annotations. It works on top of any pre-trained language model as a pure -time strategy.
Formal description: marginalizing out reasoning paths
Formally, self-consistency introduces a into the generation process. Given a prompt and question, the model generates a pair (rᵢ, aᵢ) where rᵢ is the reasoning path (a sequence of tokens explaining the solution steps) and aᵢ is the final answer parsed from that path. After sampling m such pairs from the decoder, self-consistency marginalizes out the reasoning paths by selecting the most frequent answer.
The authors also investigated weighting each answer by its generation probability. They found that a normalized weighted sum performs comparably to simple majority vote, because the model assigns similar probabilities to different valid reasoning paths. In other words, the model considers the sampled paths "equally likely," making the unweighted majority vote both simpler and equally effective.
Results: consistent gains across tasks and scales
The authors evaluated self-consistency across four language models (UL2-20B, LaMDA-137B, GPT-3-175B, PaLM-540B) on benchmarks (GSM8K, SVAMP, AQuA, MultiArith, AddSub, ASDiv) and benchmarks (CommonsenseQA, StrategyQA, ARC).
The results are striking. On GSM8K, PaLM-540B improved from 56.5% (greedy CoT) to 74.4% with self-consistency — a gain of nearly 18 percentage points. On AQuA, GPT-3 code-davinci-002 jumped from 39.8% to 52.0%. On StrategyQA, PaLM-540B improved from 75.3% to 81.6%. These gains are larger than those achieved by training additional verifier models, despite self-consistency requiring zero extra training.
Two patterns emerge from the results. First, gains are larger on harder tasks where greedy decoding is lower — there is more room for diverse paths to recover from errors. Second, gains increase with model scale: larger models generate higher- quality reasoning paths, so the diversity is more informative.
More paths, better answers
The authors studied how accuracy scales with the number of sampled reasoning paths. The pattern is consistent: accuracy improves rapidly from 1 to 10 paths, with diminishing but still positive returns up to 40 paths. Even just 5 sampled paths delivers a substantial improvement over greedy decoding — making self-consistency practical even when compute is limited.
This curve reveals something important. The gains don't come from a single "lucky" sample. They come from the statistical structure of the problem: as you add more paths, the probability that the majority answer is correct increases monotonically. In practice, 10–20 paths captures most of the benefit.
Comparisons: why not just rank or beam search?
The paper carefully compares self-consistency against several alternatives. Sample-and-rank samples multiple sequences and picks the one with the highest log-probability. It helps slightly, but far less than self-consistency, because ranking selects a single path — it doesn't leverage agreement across paths.
is even worse. Beams are diverse by construction, but beam search tends to produce very similar outputs. The diversity that makes self-consistency powerful is exactly what beam search suppresses.
Prompt-order ensembles randomly permute the exemplars in the prompt and take majority vote over greedy decodings. This captures sensitivity to prompt ordering but doesn't diversify the reasoning process itself. Self-consistency with 40 sampled paths outperforms 40 prompt permutations by a wide margin (27.7% vs 19.2% on GSM8K with LaMDA-137B).
Model ensembles (majority vote over outputs from multiple different language models) actually perform worse than self-consistency on a single model, because weaker models in the ensemble drag down the average.
Robustness: it just works
Self-consistency is remarkably robust. It works across different sampling strategies — temperature sampling at T=0.3, 0.5, 0.7, top-k with k=20 or k=40, and nucleus sampling with p=0.9 or p=0.95 — all delivering strong gains. No careful tuning of sampling parameters is needed.
Even when the chain-of-thought prompts contain mistakes (wrong intermediate numbers but correct final answers), self-consistency still improves accuracy. On GSM8K with LaMDA-137B, imperfect prompts with greedy decoding scored 14.9%, but adding self-consistency brought it up to 23.4%.
The authors also showed that self-consistency works with zero-shot chain-of-thought (where the prompt simply says "Let's think step by step" without any examples). On GSM8K with PaLM-540B, zero-shot CoT scored 43.0% — with self-consistency, it jumped to 69.2%, a gain of over 26 points.
Perhaps most usefully, the consistency level (what fraction of sampled paths agree on the final answer) is strongly correlated with accuracy. This means self-consistency can double as an uncertainty estimator: when consistency is low, the model is likely unsure. It gives the model a way to "know when it doesn't know."
Consistency as a confidence signal
Beyond improving accuracy, self-consistency reveals how confident the model is. The authors plotted the relationship between consistency (percentage of sampled paths agreeing on the majority answer) and accuracy (whether that answer is correct). The correlation is remarkably strong: questions where 90% or more of the paths agree are almost always answered correctly, while low-consistency questions are much more likely to be wrong.
This is valuable in practice. A system using self-consistency can flag low-confidence answers for human review, route them to a more capable model, or simply indicate uncertainty to the user. The consistency score is a free byproduct of the method — no additional is needed.
Legacy: from self-consistency to test-time compute
2022
Chain-of-Thought Prompting (CoT)
Wei et al. showed that prompting language models with step-by-step examples enables multi-step reasoning. The default decoding strategy was greedy.
2022
Self-Consistency (this paper)
Replaced greedy decoding with sample-and-vote. Showed that diversity in reasoning paths is a powerful signal, not noise — establishing the "test-time compute" idea.
2023
Tree of Thoughts (ToT)
Yao et al. extended the idea by structuring exploration as a tree with backtracking, evaluating partial thoughts, and searching for optimal reasoning paths explicitly.
2024
OpenAI o1 and test-time compute scaling
OpenAI o1 operationalized the idea at scale — spending more compute at inference time to reason through problems via hidden chains of thought, echoing the fundamental insight that more reasoning paths yield better answers.
Self-consistency demonstrated a simple but profound principle: you can trade inference-time compute for accuracy. This "test-time compute" paradigm — spending more computation at prediction time rather than at training time — has become one of the most active research directions in modern AI. Every system that samples multiple reasoning chains and aggregates their answers, from coding assistants to mathematical solvers, owes an intellectual debt to this paper.
Simplified to show the idea — not the real implementation.
from collections import Counter
def self_consistency(model, prompt, question, n_paths=40, temperature=0.7):
"""Sample n_paths reasoning chains, return the majority-vote answer."""
answers = []
for _ in range(n_paths):
# Step 1 & 2: Sample a reasoning path with temperature > 0
output = model.generate(prompt + question, temperature=temperature)
# Parse the final answer from the reasoning path
answer = parse_answer(output) # e.g., extract text after "The answer is"
answers.append(answer)
# Step 3: Majority vote
vote_counts = Counter(answers)
best_answer = vote_counts.most_common(1)[0][0]
confidence = vote_counts[best_answer] / n_paths # consistency score
return best_answer, confidenceCitationWang, Wei, Schuurmans, Le, Chi, Narang, Chowdhery, Zhou. Self-Consistency Improves Chain of Thought Reasoning in Language Models. ICLR, 2023.
Terms in this paper
- Self-Consistencyالاتّساق الذاتي
- Chain of Thoughtسلسلة التفكير
- Greedy Decodingفك الترميز الجشع
- Samplingاختيار العينات الاحتمالية
- Majority Voteالتصويت بالأغلبية
- Reasoningالاستدلال
- Temperatureالحرارة
- Arithmetic Reasoningالاستدلال الحسابي
- Commonsense Reasoningاستدلال الحس السليم
- Nucleus Samplingمعاينة النواة الاحتمالية
- Beam Searchبحث الحزمة
- Ensembleالنماذج التجميعية الهجينة
- Few-Shot Promptingالتحفيز بأمثلة قليلة
- Language Modelالنموذج اللغوي