Language Models2023intermediate12 min read

Textbooks Are All You Need

الكتب المدرسية هي كل ما تحتاجه

Gunasekar, S. · Zhang, Y. · Aneja, J. · Mendes, C. C. T. · Del Giorno, A. · Gopi, S. · Javaheripi, M. · Kauffmann, P. · de Rosa, G. · Saarikivi, O. · Salim, A. · Shah, S. · Behl, H. S. · Wang, X. · Bubeck, S. · Eldan, R. · Kalai, A. T. · Lee, Y. T. · Li, Y. — arXiv

The problem

By 2023, the dominant recipe for building better code-generation models was straightforward: use more data and bigger models. predicted that performance improves as you increase compute, data, and parameters. But this scaling approach has diminishing returns and enormous environmental and financial costs. Meanwhile, the training data itself — scraped from GitHub, StackOverflow, and other web sources — is noisy, repetitive, poorly documented, and often non-instructive. A human student forced to learn from such data would be lost in a sea of boilerplate, incomplete snippets, and undocumented functions.

The contribution

phi-1, a 1.3B-parameter trained on only 7 billion tokens of carefully curated "textbook quality" data — a combination of filtered web code (6B tokens) and synthetically generated textbooks and exercises using GPT-3.5 (1B tokens). Despite being 10× smaller in model size and 100× smaller in dataset size than competitors, phi-1 achieves 50.6% pass@1 on and 55.5% on MBPP, outperforming nearly all open-source code models. The paper also reveals surprising emergent capabilities after on a small exercises dataset, including the ability to use APIs and libraries never seen during fine-tuning.

The impact

This paper shifted the AI community's focus from "scale everything" to "curate your data." It demonstrated that can break existing scaling laws, achieving frontier performance at a fraction of the cost. phi-1 directly inspired the phi model family (phi-1.5, phi-2, phi-3) and energized research on generation, data curation, and efficient training. It proved that small, well-trained models can compete with giants — democratizing access to high-performance AI.

Imagine two medical students preparing for their board exams. Student A is handed a complete university library — millions of pages, most of them irrelevant, many contradictory, some downright wrong. Student B receives a single textbook: concise, expertly curated, with clear explanations and practice exercises at the end of each chapter.

Who passes first? Student B — by a wide margin. Not because they read more, but because every page they read taught them something.

This paper applies the same principle to training language models for code. Instead of feeding a model terabytes of messy internet code, they curated a small "textbook" of high-quality data and trained a tiny model that outperforms giants.

The problem: web code is a terrible textbook

The standard recipe for training code models by 2023 was to scrape millions of files from open-source repositories like The Stack, combine them with Q&A from StackOverflow, and train a massive Transformer on hundreds of billions of tokens. Models like StarCoder (15.5B parameters, 1 trillion tokens) and CodeGen (16.1B parameters, 577B tokens) followed this playbook.

But Gunasekar et al. noticed something important by actually looking at this data. Most code snippets from the web suffer from four critical flaws: they are not self-contained (depending on external modules the model never sees), they involve trivial boilerplate (setting constants, configuring GUIs), they bury algorithmic logic inside poorly documented functions, and they are skewed toward narrow topics. A human learner would find this data frustrating and inefficient.

The authors conjectured that language models suffer from the same problems. If the data does not clearly map natural language to code in a clean, instructive way, the model wastes capacity learning noise rather than reasoning patterns.

Open in Lab
Compare high-quality "textbook" code (left) with typical web-scraped code (right). Toggle examples to see the difference in clarity and instructive value.
The demo wakes as you arrive…

The solution: three curated datasets

The phi-1 training pipeline uses three carefully crafted datasets, totalling less than 7 billion tokens — a fraction of what competitors use. Think of the pipeline as a three-stage education system: first a filtered library of the best available code, then a synthetic textbook written by an AI tutor, and finally a set of exercises for practice.

1. Filtered Code (≈6B tokens). The authors started with the Python subset of The Stack and StackOverflow — over 35 million files. They used GPT-4 to annotate a small subset (≈100K samples) for "educational value," then trained a random forest classifier on the embeddings from a pretrained CodeGen model to filter the rest. This classifier acts as a librarian, keeping only the books worth reading.

2. Synthetic Textbooks (fewer than 1B tokens). They prompted GPT-3.5 to generate Python textbooks — natural language explanations interspersed with code snippets, covering topics from basic algorithms to data structures. To ensure diversity, they constrained topics and target audiences in the prompts, much like asking an author to write chapters for different readers.

3. CodeExercises (≈180M tokens). A small synthetic dataset of Python exercises where each exercise is a function docstring that needs to be completed. Diversity was achieved by constraining function names. This dataset is used only for fine-tuning.

Open in Lab
Explore phi-1's three-stage data pipeline. Click each stage to see the dataset size, composition, and role in training.
The demo wakes as you arrive…

Quality filtering: using GPT-4 as a librarian

The filtering process is clever in its simplicity. Instead of manually reviewing millions of files, the authors used GPT-4 to judge the "educational value" of a small sample — about 100K snippets from The Stack and StackOverflow. The prompt asked GPT-4 to evaluate whether each code snippet would be useful for "a student whose goal is to learn basic coding concepts."

They then trained a random forest classifier using the output embeddings of a pretrained CodeGen 350M model as features. This classifier acts as a fast proxy for GPT-4's judgment, allowing them to filter millions of files efficiently. The result: from 35 million files, they kept only those with high educational value — roughly 6 billion tokens worth.

Even without the synthetic data, this filtering alone boosted HumanEval performance from 12.19% (unfiltered Stack) to 17.68%. Combined with synthetic textbooks, it reached 20.12% for a 350M model — before any fine-tuning.

Open in Lab
See what the quality filter keeps vs discards. High-value code is self-contained and instructive; low-value code is boilerplate or context-dependent.
The demo wakes as you arrive…

Synthetic data: writing the AI textbook

The synthetic textbook is perhaps the most creative part of the pipeline. The authors used GPT-3.5 to generate Python textbook chapters — passages of natural language explanation interspersed with working code. The key challenge was diversity: if you simply ask GPT-3.5 to "write a Python textbook," you get the same sorting algorithms and basic patterns repeated endlessly.

To solve this, inspired by the TinyStories work, the authors injected randomness into the prompts. They constrained the topic (e.g., "write about graph algorithms for data scientists") and the target audience (e.g., "for beginners with statistics background") so that each generated chapter approaches coding from a different angle.

The CodeExercises dataset followed a similar strategy but focused on function completion — each exercise is a docstring describing what the function should do, and GPT-3.5 generates both the docstring and the solution. Diversity was achieved by constraining function names, forcing the model to invent exercises for unusual function signatures.

Example: synthetic textbook excerpt — teaching matrix concepts with codepython

Simplified to show the idea — not the real implementation.

# From the synthetic textbook dataset:
# "A matrix is singular if its determinant is zero."
import numpy as np

def is_singular(A):
    """Check if a matrix is singular (determinant = 0)."""
    det = np.linalg.det(A)
    return det == 0

A = np.array([[1, 2], [2, 4]])
print(is_singular(A))  # True — rows are linearly dependent
Example: synthetic exercise — function completion taskpython

Simplified to show the idea — not the real implementation.

def valid_guessing_letters(word: str, guesses: List[str]) -> List[str]:
    """
    Returns valid guessing letters: letters not yet
    guessed that are present in the word.

    Parameters:
        word (str): The word to guess.
        guesses (List[str]): Already guessed letters.
    Returns:
        List[str]: Valid guessing letters.
    """
    valid_letters = []
    for letter in word:
        if letter not in guesses and letter not in valid_letters:
            valid_letters.append(letter)
    return valid_letters

Model architecture: conventional Transformer, unconventional data

The phi-1 architecture is deliberately conventional — the entire point of the paper is that the data is what matters, not architectural novelty. phi-1 is a decoder-only Transformer with 24 layers, hidden dimension 2048, MLP inner dimension 8192, and 32 attention heads. It uses rotary position embeddings (RoPE), FlashAttention for efficient attention computation, and parallel MHA + MLP configurations following CodeGen and PaLM.

The smaller variant, phi-1-small, has 350M parameters with 20 layers, hidden dimension 1024, and 16 attention heads. Both models use the CodeGen tokenizer.

Training used 8 NVIDIA A100 GPUs with fp16 precision, optimizer, linear warmup followed by linear decay, and dropout of 0.1 on attention and residuals. Pre-training phi-1-base took under 4 days. Fine-tuning to produce phi-1 required only 7 additional hours.

Open in Lab
Explore phi-1's architecture and training pipeline. Click stages to see hyperparameters and training details.
The demo wakes as you arrive…

Breaking scaling laws: less is more

The results are striking. phi-1 with just 1.3B parameters and 7B training tokens achieves 50.6% on HumanEval — outperforming StarCoder (15.5B parameters, 1T tokens, 33.6%), CodeGen-16.1B (29.3%), and even GPT-3.5 (47%). Only GPT-4 (67%) and WizardCoder (57.3% on HumanEval but worse on MBPP) surpass it.

On MBPP (Mostly Basic Python Programs), phi-1 scores 55.5%, beating every model in the comparison except GPT-4. This is especially remarkable because phi-1 was trained with roughly 100× fewer tokens and is 10× smaller than StarCoder.

Even phi-1-small at 350M parameters achieves 45% on HumanEval — competitive with models 50× its size. The key takeaway: data quality can shift scaling laws so dramatically that a small, well-fed model outperforms a large, poorly-fed one.

Open in Lab
Compare phi-1 against other code models by model size, dataset size, and HumanEval performance. Drag models to see how phi-1 breaks the expected scaling trend.
The demo wakes as you arrive…

The surprise: emergent capabilities from fine-tuning

The most fascinating finding goes beyond scores. After fine-tuning on the small CodeExercises dataset — which contains only basic Python functions using standard libraries — phi-1 develops unexpected abilities that are not present in the fine-tuning data.

Specifically, the fine-tuned phi-1 can correctly use external libraries like PyGame (game development), Tkinter (GUI), PyTorch (deep learning), and Matplotlib (plotting) — none of which appear in CodeExercises. The base model phi-1-base, which saw these libraries during pre-training but was not fine-tuned, cannot use them coherently.

This suggests that fine-tuning on well-structured exercises does more than teach new tasks — it helps the model reorganize and consolidate knowledge it already acquired during pre-training but could not access reliably. Think of it like studying practice problems before an exam: you are not learning new material, you are making connections between concepts you already know.

Open in Lab
Compare outputs of phi-1 (fine-tuned), phi-1-base (pre-trained only), and phi-1-small on tasks using external libraries. See how fine-tuning unlocks capabilities not present in the training data.
The demo wakes as you arrive…

Limitations: what phi-1 cannot do

phi-1 is not without weaknesses. Being specialized in Python, it cannot generate code in other languages. Its small parameter count limits its ability to handle complex tasks like building full Flask applications. It is also sensitive to prompt variations — performance drops significantly with longer prompts or grammatical errors in the instructions.

The structured nature of the training data means phi-1 is less robust to stylistic variations. It struggles with spatial reasoning and counting tasks, and its chat capabilities are limited compared to larger models. These limitations are not fundamental — but they highlight that data quality can complement, not entirely replace, model scale.

Ablation: each dataset component matters

The paper provides a clear ablation study through Figure 2.1. For 350M parameter models, training on unfiltered Stack data gives only 11% on HumanEval. Switching to the filtered CodeTextbook dataset (without synthetic exercises) jumps to 20%. Fine-tuning on CodeExercises pushes it to 45%.

For 1.3B parameter models, the pattern is even more striking. phi-1-base trained on CodeTextbook achieves 29% — matching Replit-Finetuned (2.7B parameters trained on 100× more tokens). Adding CodeExercises fine-tuning reaches 50.6%.

Each component builds on the last: filtering raises the floor, synthetic textbooks teach reasoning, and exercises trigger the consolidation of knowledge into executable skills.

Open in Lab
Each bar group shows how adding data quality components improves HumanEval performance. Click to compare 350M and 1.3B models.
The demo wakes as you arrive…

Legacy: data quality becomes a research frontier

  1. 2021

    Codex (OpenAI)

    Established the code generation benchmark with HumanEval. Trained on 100B tokens of code — the "more data" approach.

  2. 2023

    TinyStories (Eldan & Li)

    Showed that a 10M-parameter model can generate coherent English when trained on high-quality synthetic stories. Inspired the phi-1 approach.

  3. 2023

    phi-1 — this paper

    Demonstrated that textbook-quality data can break scaling laws for code generation, achieving state-of-the-art with 1.3B parameters and 7B tokens.

  4. 2023

    phi-1.5

    Extended the textbook approach to common-sense reasoning in natural language, creating a 1.3B model competitive with 5× larger models.

  5. 2024

    phi-2, phi-3, and the small model revolution

    The phi family continued to grow, proving that data-centric training generalizes across domains. The broader community embraced synthetic data and data curation as first-class research priorities.

phi-1 showed the AI community that the path to better models is not always through bigger GPUs and more data. Sometimes it runs through better libraries, better textbooks, and better exercises. The paper's core message — "quality over quantity" — has become a guiding principle for efficient AI development, and its impact extends far beyond .

CitationGunasekar, Zhang, Aneja, Mendes, Del Giorno, Gopi, Javaheripi, Kauffmann, de Rosa, Saarikivi, Salim, Shah, Behl, Wang, Bubeck, Eldan, Kalai, Lee, Li. Textbooks Are All You Need. arXiv, 2023.

Terms in this paper