Language Models2023intermediate12 min read
Textbooks Are All You Need
الكتب المدرسية هي كل ما تحتاجه
Gunasekar, S. · Zhang, Y. · Aneja, J. · Mendes, C. C. T. · Del Giorno, A. · Gopi, S. · Javaheripi, M. · Kauffmann, P. · de Rosa, G. · Saarikivi, O. · Salim, A. · Shah, S. · Behl, H. S. · Wang, X. · Bubeck, S. · Eldan, R. · Kalai, A. T. · Lee, Y. T. · Li, Y. — arXiv
The problem
By 2023, the dominant recipe for building better code-generation models was straightforward: use more data and bigger models. predicted that performance improves as you increase compute, data, and parameters. But this scaling approach has diminishing returns and enormous environmental and financial costs. Meanwhile, the training data itself — scraped from GitHub, StackOverflow, and other web sources — is noisy, repetitive, poorly documented, and often non-instructive. A human student forced to learn from such data would be lost in a sea of boilerplate, incomplete snippets, and undocumented functions.
The contribution
phi-1, a 1.3B-parameter trained on only 7 billion tokens of carefully curated "textbook quality" data — a combination of filtered web code (6B tokens) and synthetically generated textbooks and exercises using GPT-3.5 (1B tokens). Despite being 10× smaller in model size and 100× smaller in dataset size than competitors, phi-1 achieves 50.6% pass@1 on and 55.5% on MBPP, outperforming nearly all open-source code models. The paper also reveals surprising emergent capabilities after on a small exercises dataset, including the ability to use APIs and libraries never seen during fine-tuning.
The impact
This paper shifted the AI community's focus from "scale everything" to "curate your data." It demonstrated that can break existing scaling laws, achieving frontier performance at a fraction of the cost. phi-1 directly inspired the phi model family (phi-1.5, phi-2, phi-3) and energized research on generation, data curation, and efficient training. It proved that small, well-trained models can compete with giants — democratizing access to high-performance AI.
Imagine two medical students preparing for their board exams. Student A is handed a complete university library — millions of pages, most of them irrelevant, many contradictory, some downright wrong. Student B receives a single textbook: concise, expertly curated, with clear explanations and practice exercises at the end of each chapter.
Who passes first? Student B — by a wide margin. Not because they read more, but because every page they read taught them something.
This paper applies the same principle to training language models for code. Instead of feeding a model terabytes of messy internet code, they curated a small "textbook" of high-quality data and trained a tiny model that outperforms giants.
The problem: web code is a terrible textbook
The standard recipe for training code models by 2023 was to scrape millions of files from open-source repositories like The Stack, combine them with Q&A from StackOverflow, and train a massive Transformer on hundreds of billions of tokens. Models like StarCoder (15.5B parameters, 1 trillion tokens) and CodeGen (16.1B parameters, 577B tokens) followed this playbook.
But Gunasekar et al. noticed something important by actually looking at this data. Most code snippets from the web suffer from four critical flaws: they are not self-contained (depending on external modules the model never sees), they involve trivial boilerplate (setting constants, configuring GUIs), they bury algorithmic logic inside poorly documented functions, and they are skewed toward narrow topics. A human learner would find this data frustrating and inefficient.
The authors conjectured that language models suffer from the same problems. If the data does not clearly map natural language to code in a clean, instructive way, the model wastes capacity learning noise rather than reasoning patterns.
The solution: three curated datasets
The phi-1 training pipeline uses three carefully crafted datasets, totalling less than 7 billion tokens — a fraction of what competitors use. Think of the pipeline as a three-stage education system: first a filtered library of the best available code, then a synthetic textbook written by an AI tutor, and finally a set of exercises for practice.
1. Filtered Code (≈6B tokens). The authors started with the Python subset of The Stack and StackOverflow — over 35 million files. They used GPT-4 to annotate a small subset (≈100K samples) for "educational value," then trained a random forest classifier on the embeddings from a pretrained CodeGen model to filter the rest. This classifier acts as a librarian, keeping only the books worth reading.
2. Synthetic Textbooks (fewer than 1B tokens). They prompted GPT-3.5 to generate Python textbooks — natural language explanations interspersed with code snippets, covering topics from basic algorithms to data structures. To ensure diversity, they constrained topics and target audiences in the prompts, much like asking an author to write chapters for different readers.
3. CodeExercises (≈180M tokens). A small synthetic dataset of Python exercises where each exercise is a function docstring that needs to be completed. Diversity was achieved by constraining function names. This dataset is used only for fine-tuning.
Quality filtering: using GPT-4 as a librarian
The filtering process is clever in its simplicity. Instead of manually reviewing millions of files, the authors used GPT-4 to judge the "educational value" of a small sample — about 100K snippets from The Stack and StackOverflow. The prompt asked GPT-4 to evaluate whether each code snippet would be useful for "a student whose goal is to learn basic coding concepts."
They then trained a random forest classifier using the output embeddings of a pretrained CodeGen 350M model as features. This classifier acts as a fast proxy for GPT-4's judgment, allowing them to filter millions of files efficiently. The result: from 35 million files, they kept only those with high educational value — roughly 6 billion tokens worth.
Even without the synthetic data, this filtering alone boosted HumanEval performance from 12.19% (unfiltered Stack) to 17.68%. Combined with synthetic textbooks, it reached 20.12% for a 350M model — before any fine-tuning.
Synthetic data: writing the AI textbook
The synthetic textbook is perhaps the most creative part of the pipeline. The authors used GPT-3.5 to generate Python textbook chapters — passages of natural language explanation interspersed with working code. The key challenge was diversity: if you simply ask GPT-3.5 to "write a Python textbook," you get the same sorting algorithms and basic patterns repeated endlessly.
To solve this, inspired by the TinyStories work, the authors injected randomness into the prompts. They constrained the topic (e.g., "write about graph algorithms for data scientists") and the target audience (e.g., "for beginners with statistics background") so that each generated chapter approaches coding from a different angle.
The CodeExercises dataset followed a similar strategy but focused on function completion — each exercise is a docstring describing what the function should do, and GPT-3.5 generates both the docstring and the solution. Diversity was achieved by constraining function names, forcing the model to invent exercises for unusual function signatures.
Simplified to show the idea — not the real implementation.
# From the synthetic textbook dataset:
# "A matrix is singular if its determinant is zero."
import numpy as np
def is_singular(A):
"""Check if a matrix is singular (determinant = 0)."""
det = np.linalg.det(A)
return det == 0
A = np.array([[1, 2], [2, 4]])
print(is_singular(A)) # True — rows are linearly dependent
Simplified to show the idea — not the real implementation.
def valid_guessing_letters(word: str, guesses: List[str]) -> List[str]:
"""
Returns valid guessing letters: letters not yet
guessed that are present in the word.
Parameters:
word (str): The word to guess.
guesses (List[str]): Already guessed letters.
Returns:
List[str]: Valid guessing letters.
"""
valid_letters = []
for letter in word:
if letter not in guesses and letter not in valid_letters:
valid_letters.append(letter)
return valid_letters
Model architecture: conventional Transformer, unconventional data
The phi-1 architecture is deliberately conventional — the entire point of the paper is that the data is what matters, not architectural novelty. phi-1 is a decoder-only Transformer with 24 layers, hidden dimension 2048, MLP inner dimension 8192, and 32 attention heads. It uses rotary position embeddings (RoPE), FlashAttention for efficient attention computation, and parallel MHA + MLP configurations following CodeGen and PaLM.
The smaller variant, phi-1-small, has 350M parameters with 20 layers, hidden dimension 1024, and 16 attention heads. Both models use the CodeGen tokenizer.
Training used 8 NVIDIA A100 GPUs with fp16 precision, optimizer, linear warmup followed by linear decay, and dropout of 0.1 on attention and residuals. Pre-training phi-1-base took under 4 days. Fine-tuning to produce phi-1 required only 7 additional hours.
Breaking scaling laws: less is more
The results are striking. phi-1 with just 1.3B parameters and 7B training tokens achieves 50.6% on HumanEval — outperforming StarCoder (15.5B parameters, 1T tokens, 33.6%), CodeGen-16.1B (29.3%), and even GPT-3.5 (47%). Only GPT-4 (67%) and WizardCoder (57.3% on HumanEval but worse on MBPP) surpass it.
On MBPP (Mostly Basic Python Programs), phi-1 scores 55.5%, beating every model in the comparison except GPT-4. This is especially remarkable because phi-1 was trained with roughly 100× fewer tokens and is 10× smaller than StarCoder.
Even phi-1-small at 350M parameters achieves 45% on HumanEval — competitive with models 50× its size. The key takeaway: data quality can shift scaling laws so dramatically that a small, well-fed model outperforms a large, poorly-fed one.
The surprise: emergent capabilities from fine-tuning
The most fascinating finding goes beyond scores. After fine-tuning on the small CodeExercises dataset — which contains only basic Python functions using standard libraries — phi-1 develops unexpected abilities that are not present in the fine-tuning data.
Specifically, the fine-tuned phi-1 can correctly use external libraries like PyGame (game development), Tkinter (GUI), PyTorch (deep learning), and Matplotlib (plotting) — none of which appear in CodeExercises. The base model phi-1-base, which saw these libraries during pre-training but was not fine-tuned, cannot use them coherently.
This suggests that fine-tuning on well-structured exercises does more than teach new tasks — it helps the model reorganize and consolidate knowledge it already acquired during pre-training but could not access reliably. Think of it like studying practice problems before an exam: you are not learning new material, you are making connections between concepts you already know.
Limitations: what phi-1 cannot do
phi-1 is not without weaknesses. Being specialized in Python, it cannot generate code in other languages. Its small parameter count limits its ability to handle complex tasks like building full Flask applications. It is also sensitive to prompt variations — performance drops significantly with longer prompts or grammatical errors in the instructions.
The structured nature of the training data means phi-1 is less robust to stylistic variations. It struggles with spatial reasoning and counting tasks, and its chat capabilities are limited compared to larger models. These limitations are not fundamental — but they highlight that data quality can complement, not entirely replace, model scale.
Ablation: each dataset component matters
The paper provides a clear ablation study through Figure 2.1. For 350M parameter models, training on unfiltered Stack data gives only 11% on HumanEval. Switching to the filtered CodeTextbook dataset (without synthetic exercises) jumps to 20%. Fine-tuning on CodeExercises pushes it to 45%.
For 1.3B parameter models, the pattern is even more striking. phi-1-base trained on CodeTextbook achieves 29% — matching Replit-Finetuned (2.7B parameters trained on 100× more tokens). Adding CodeExercises fine-tuning reaches 50.6%.
Each component builds on the last: filtering raises the floor, synthetic textbooks teach reasoning, and exercises trigger the consolidation of knowledge into executable skills.
Legacy: data quality becomes a research frontier
2021
Codex (OpenAI)
Established the code generation benchmark with HumanEval. Trained on 100B tokens of code — the "more data" approach.
2023
TinyStories (Eldan & Li)
Showed that a 10M-parameter model can generate coherent English when trained on high-quality synthetic stories. Inspired the phi-1 approach.
2023
phi-1 — this paper
Demonstrated that textbook-quality data can break scaling laws for code generation, achieving state-of-the-art with 1.3B parameters and 7B tokens.
2023
phi-1.5
Extended the textbook approach to common-sense reasoning in natural language, creating a 1.3B model competitive with 5× larger models.
2024
phi-2, phi-3, and the small model revolution
The phi family continued to grow, proving that data-centric training generalizes across domains. The broader community embraced synthetic data and data curation as first-class research priorities.
phi-1 showed the AI community that the path to better models is not always through bigger GPUs and more data. Sometimes it runs through better libraries, better textbooks, and better exercises. The paper's core message — "quality over quantity" — has become a guiding principle for efficient AI development, and its impact extends far beyond .
CitationGunasekar, Zhang, Aneja, Mendes, Del Giorno, Gopi, Javaheripi, Kauffmann, de Rosa, Saarikivi, Salim, Shah, Behl, Wang, Bubeck, Eldan, Kalai, Lee, Li. Textbooks Are All You Need. arXiv, 2023.
Terms in this paper
- Synthetic Dataالبيانات الاصطناعية
- Data Qualityجودة البيانات
- Scaling Lawsقوانين التوسعة
- Fine-Tuningالضبط الدقيق
- Code Generationتوليد الشيفرات
- HumanEvalHumanEval
- Transformerالمحوِّل
- Decoder-Only Modelنموذج فكّ الترميز فقط
- Knowledge Distillationتقطير المعرفة
- Emergent Abilitiesالقدرات المعرفية الناشئة فجأة
- Data Augmentationتعزيز البيانات
- Benchmarkالمعيار المرجعي
- Overfittingفرط التخصيص
- Generalizationالتعميم
- Tokenizationتجزئة النصوص