Language Models2022intermediate10 min read
STaR: Bootstrapping Reasoning With Reasoning
STaR: تمهيد الاستدلال بالاستدلال
Zelikman, E. · Wu, Y. · Mu, J. · Goodman, N.D. — NeurIPS
The problem
By 2022, chain-of-thought prompting had shown that asking language models to "think step by step" dramatically improves their . But there was a catch: either you needed thousands of hand-written rationales for (expensive and task-specific), or you relied on few-shot prompting alone (which significantly underperforms fine-tuned models). There was no scalable way to teach a model to reason without massive human annotation effort.
The contribution
STaR (Self-Taught Reasoner): an iterative method where a language model generates chain-of-thought rationales for a dataset, keeps only those leading to correct answers, and fine-tunes on them. For problems it fails, it uses "rationalization" — generating a given the correct answer as a hint, then training as if it reasoned independently. Repeating this loop, a 6B-parameter GPT-J reaches 72.5% on CommonsenseQA, matching a 30× larger GPT-3 fine-tuned directly, and dramatically improves on arithmetic and GSM8K.
The impact
STaR introduced the paradigm of self-improving reasoning through iterative bootstrapping — the idea that a model can teach itself to reason better by learning from its own successful chains of thought. This paradigm directly influenced OpenAI's o1 and DeepSeek-R1, which scale test-time reasoning via reinforcement learning on self-generated rationales. STaR showed that even modest models can punch far above their weight when trained on their own reasoning traces, opening the door to scalable reasoning without human annotation.
Imagine a junior chef who can only cook three dishes. Her mentor gives her a recipe book with thousands of dishes but no instructions — just photos of the final plates.
The chef tries each dish using her existing skills. When a dish turns out right, she writes down exactly what she did — her own recipe. When a dish fails, the mentor shows her what the final plate should look like, and she reverse-engineers a recipe from the result. She adds both kinds of recipes to her personal cookbook and practices again.
After a few rounds, she's cooking dishes she never could before — taught entirely by her own successes and reverse-engineered solutions. That is STaR.
The problem: reasoning is powerful but expensive to teach
Chain-of-thought prompting showed that when language models write out their reasoning step by step before answering, they perform dramatically better on tasks like math, , and code evaluation. But there was a fundamental dilemma:
Path A — Fine-tune on rationales: Build a large dataset of problems paired with step-by-step human reasoning. This works well but is extremely expensive: you need thousands of manually written rationales, and they don't transfer across tasks.
Path B — Few-shot prompting: Show the model a handful of reasoning examples in its prompt. Much cheaper, but the accuracy gap compared to fine-tuned models is large — you're leaving significant performance on the table.
Neither path was scalable. The field needed a way to get fine-tuning quality from few-shot cost — to somehow amplify a few reasoning examples into thousands, without hiring an army of annotators.
The idea: let the model teach itself to reason
STaR's key insight is that a model's own correct reasoning is valid training data. If the model generates a chain-of-thought rationale and arrives at the right answer, that rationale is worth learning from — even though no human wrote it.
The algorithm is a simple loop with four steps:
-
Generate: Prompt the model (with a few rationale examples) to solve every problem in the dataset, producing a step-by-step rationale followed by an answer.
-
Filter: Keep only the rationales that led to the correct answer. Discard the rest.
-
Fine-tune: Train the base model on the filtered correct rationales.
-
Repeat: Use the improved model to generate new, better rationales. Go back to step 1.
Each iteration, the model solves more problems, generates more training data, and becomes better at reasoning. It's a virtuous cycle — improvements in reasoning produce better training data, which produce further improvements.
Rationalization: learning from failures too
The basic loop has a fundamental limitation: the model only trains on problems it already solves correctly. Once it stops solving new problems, improvement stalls. It receives no learning signal from its failures.
STaR solves this with rationalization. For every problem the model gets wrong, it gets a second chance: the correct answer is provided as a hint, and the model is asked to generate a rationale that explains why that answer is correct. This is like reverse engineering — given the destination, find the path.
The crucial trick: when this rationalized rationale is added to the training set, the hint is removed. The model is trained as if it had arrived at the rationale independently. This forces the model to internalize reasoning patterns for problems that were previously beyond its reach.
The formal view: STaR as approximate policy gradient
Before seeing the math, here is the intuition: think of the model as an agent that chooses a reasoning path (the rationale) before giving a final answer. If the answer is correct, the reasoning path gets rewarded. If it's wrong, the path is discarded. This is exactly the structure of a in reinforcement learning — the rationale is the "action," the correct/incorrect answer is the "reward," and STaR is optimizing the policy to produce more rewarding reasoning paths.
The full algorithm
Simplified to show the idea — not the real implementation.
def star(model_base, dataset, few_shot_prompts, num_iterations):
"""STaR: Self-Taught Reasoner bootstrapping loop."""
model = copy(model_base)
for iteration in range(num_iterations):
correct_rationales = []
rationalized = []
for question, answer in dataset:
# Step 1: Generate rationale with few-shot prompting
prompt = few_shot_prompts + [question]
rationale, predicted = model.generate(prompt)
if predicted == answer:
# Correct → keep this self-generated rationale
correct_rationales.append((question, rationale, answer))
else:
# Wrong → rationalize: give the answer as a hint
hint_prompt = few_shot_prompts + [question + f" (Answer: {answer})"]
rationale_hint, pred_hint = model.generate(hint_prompt)
if pred_hint == answer:
# Store WITHOUT the hint — as if the model solved it alone
rationalized.append((question, rationale_hint, answer))
# Step 3: Fine-tune base model on all correct rationales
train_data = correct_rationales + rationalized
model = finetune(model_base, train_data) # always from base!
return modelResults: small model, big performance
STaR was evaluated on three domains using GPT-J (6B parameters) — a model far smaller than the state-of-the-art at the time:
CommonsenseQA: STaR reached 72.5% accuracy, a huge leap from the few-shot baseline (36.6%) and the directly fine-tuned GPT-J (60.0%). This matched GPT-3's fine-tuned result of 73.0% — despite GPT-3 being 30× larger with 175B parameters. STaR trained on only 86.7% of the data (78.2% from rationale generation, 8.5% from rationalization).
Arithmetic: After 16 iterations, STaR achieved 89.5% accuracy on multi-digit addition, up from near-zero few-shot performance. With rationalization, the model learned multiple digit lengths simultaneously rather than sequentially, and even generalized to unseen 9- and 10-digit problems.
GSM8K (grade school math): STaR improved from 3.1% (few-shot) to 10.7% with rationalization, using only 28.7% of the dataset. The model often found simpler solutions than the human-written ground truth, sometimes solving 7-step problems in a single step.
Rationale quality: humans prefer STaR
Beyond accuracy, the authors asked a natural question: are STaR's rationales actually good? To find out, they ran a human evaluation on Prolific. Twenty crowdworkers were shown questions with three rationales in random order — one from few-shot prompting, one from STaR, and one from human annotators — and asked to rank them.
The result was striking: participants preferred STaR-generated rationales over few-shot rationales 30% more often (p=0.039), and preferred STaR over human-written rationales 74% more often (p‹0.001). The authors note this likely reflects the difficulty of eliciting high-quality rationales from crowdworkers rather than superhuman reasoning, but it demonstrates that the bootstrapping process genuinely improves rationale quality, not just accuracy.
Why it works: the virtuous cycle
STaR works because of a synergy between generation and training. Better rationales produce more correct answers, which produces more training data, which produces a better model. Each component reinforces the others:
Rationale generation acts as a data amplifier: a few hand-written examples get multiplied into thousands of self-generated reasoning traces. The model's pre-existing language understanding does the heavy lifting — it already knows facts and patterns, it just needs to learn to chain them together.
Filtering by correctness acts as a quality gate: only reasoning that actually works survives into the training set. This is a simple but powerful form of reward — the answer serves as an automatic verifier.
Rationalization acts as a curriculum expander: it brings hard problems into the training set that the model wouldn't otherwise encounter, preventing the loop from stagnating on easy problems it already solves.
What STaR opened
2022
STaR — Self-Taught Reasoner
Iterative bootstrapping of reasoning from a few examples. A 6B model matches a 175B model on CommonsenseQA. First method to let an LLM improve itself through its own reasoning.
2023
ReST — Reinforced Self-Training
Extended STaR's approach with an explicit RL framework, using reward-filtered generation for math and code tasks with stronger temperature sampling.
2024
Quiet-STaR — Thinking Before Speaking
By the same authors: the model generates internal rationales at *every* token, not just when prompted. Generalizes STaR from specific tasks to general language modeling.
2024
OpenAI o1
Scales test-time reasoning with hidden chains of thought trained via RL. The lineage from STaR's self-improving rationale loop to o1's extended thinking is direct.
2025
DeepSeek-R1
Demonstrates that reasoning can emerge from pure RL on reasoning tasks without supervised fine-tuning. The self-improvement paradigm that STaR pioneered taken to its logical conclusion.
STaR's deepest contribution is not a benchmark result — it's a paradigm. Before STaR, improving reasoning required more human annotations. After STaR, the question became: how can we build systems where reasoning improves reasoning? This self-referential loop — a model using its own outputs to become better at the thing that generates those outputs — is the seed that grew into the modern reasoning-focused AI systems like o1 and R1.
CitationZelikman, Wu, Mu, Goodman. STaR: Bootstrapping Reasoning With Reasoning. NeurIPS, 2022.
Terms in this paper
- Rationaleالمبرر
- Bootstrappingالتمهيد الذاتي
- Fine-Tuningالضبط الدقيق
- Few-Shot Learningالتعلّم بأمثلة قليلة
- Self-Improvementالتحسين الذاتي
- Reasoningالاستدلال
- Policy Gradientتدرج السياسة التشغيلية
- In-Context Learningالتعلم في السياق
- Commonsense Reasoningاستدلال الحس السليم