Language Models2023beginner11 min read

Alpaca: A Strong, Replicable Instruction-Following Model

ألباكا: نموذج قوي وقابل للتكرار يتّبع التعليمات

Taori, R. · Gulrajani, I. · Zhang, T. · Dubois, Y. · Li, X. · Guestrin, C. · Liang, P. · Hashimoto, T. B. — Stanford CRFM Blog

The problem

By early 2023, instruction-following models like ChatGPT and GPT-3.5 had shown remarkable abilities, but they were closed-source, expensive to use via API, and impossible to study or modify. Meanwhile, Meta had released LLaMA — a family of strong open-weight base models — but these models could not follow instructions out of the box. The research community lacked a simple, cheap, and reproducible recipe for turning an open base model into an instruction-following assistant.

The contribution

Alpaca showed that a 7B-parameter open model (LLaMA 7B) can match the instruction-following ability of GPT-3.5 (text-davinci-003) using a three-step recipe costing under $600 total: (1) start with 175 human-written seed instructions, (2) use text-davinci-003 to generate 52K diverse instruction-output pairs via for less than $500 in API costs, (3) fine-tune LLaMA 7B on this data using standard tools for about $100 in compute. The entire pipeline — data, code, and recipe — was released openly.

The impact

Alpaca catalyzed the open-source instruction-tuning movement. Within weeks, dozens of projects replicated and extended the recipe — Vicuna, Koala, Dolly — proving that high-quality was not exclusive to trillion-dollar labs. It demonstrated that data quality and curation matter more than sheer model size, and that from a strong is a viable shortcut to . The project also sparked important debates about the legal and ethical implications of distilling proprietary model outputs.

Think of LLaMA as a brilliant new hire who read the entire company library during onboarding — they have vast knowledge, but when you ask them to draft an email or summarize a report, they stare blankly. What they need is not more reading, but a few days shadowing a senior colleague: watching 52,000 examples of how to turn a request into a polished response. After this short apprenticeship, the new hire can handle almost any task the senior colleague could — and the entire training program cost less than a plane ticket.

The problem: open models that cannot follow instructions

In March 2023 the landscape of large language models was sharply divided. On one side sat proprietary models like ChatGPT and text-davinci-003 — superb at following instructions but locked behind APIs, expensive to use at scale, and impossible for researchers to inspect or improve. On the other side sat Meta's newly released LLaMA family — strong base models available for research, but trained only as next- predictors. Ask LLaMA to "write a poem about the ocean" and it might continue with "and then write a poem about the sky" instead of actually writing a poem.

The gap was not in knowledge — LLaMA 7B already knew how poems work. The gap was in behavior: the model hadn't learned when and how to deploy its knowledge in response to user requests. Bridging this gap usually required () — a complex, expensive process involving human annotators ranking thousands of outputs.

Stanford asked a simpler question: could a straightforward on high-quality instruction data close most of this gap, cheaply and reproducibly?

Open in Lab
The Alpaca recipe: 175 seed instructions → text-davinci-003 generates 52K examples → fine-tune LLaMA 7B. Total cost: under $600.
The demo wakes as you arrive…

Step 1: generating 52K instructions from 175 seeds

The key insight of Self-Instruct is that a strong language model can generate its own training data. Alpaca builds on this idea with a simplified and cheaper pipeline.

The process begins with 175 human-written seed instructions — diverse examples covering tasks like brainstorming, classification, rewriting, coding, and open-ended generation. Each seed contains an instruction, an optional input, and an expected output.

These seeds are then used as in-context examples in prompts sent to text-davinci-003. The asks the model to generate 20 new instructions at once (instead of one at a time in the original Self-Instruct method). This aggressive batching dramatically reduces API cost. Generated pairs are filtered for quality and deduplicated using ROUGE-L similarity, yielding 52,002 unique instruction-output pairs at a total API cost of under $500.

Open in Lab
See how 175 seed instructions snowball into 52K examples through iterative generation.
The demo wakes as you arrive…

Step 2: fine-tuning LLaMA on the generated data

With 52K instruction-output pairs in hand, the training itself is straightforward supervised — no , no reward models, no complex alignment pipelines. The team used Hugging Face's standard training framework with Fully Sharded Data Parallel (FSDP) and .

The hyperparameters are remarkably simple: a of 128, a of 2e-5, 3 training epochs, a maximum sequence length of 512 tokens, and no . On a cluster of 4 A100 80GB GPUs, training took about 3 hours — roughly $100 in cloud compute.

The input format follows a simple template: each example is formatted as an instruction (optionally with an input) followed by the expected response. This flat format, without the multi-turn conversation structure used by ChatGPT, is sufficient for single-turn instruction following.

Alpaca training data format — the instruction templatepython

Simplified to show the idea — not the real implementation.

# Each example follows this template:
# If there's an input field, use "with input" format:
PROMPT_WITH_INPUT = (
    "Below is an instruction that describes a task, "
    "paired with an input that provides further context. "
    "Write a response that appropriately completes the request.\n\n"
    "### Instruction:\n{instruction}\n\n"
    "### Input:\n{input}\n\n"
    "### Response:\n{output}"
)

# If there's no input field, use "without input" format:
PROMPT_WITHOUT_INPUT = (
    "Below is an instruction that describes a task. "
    "Write a response that appropriately completes the request.\n\n"
    "### Instruction:\n{instruction}\n\n"
    "### Response:\n{output}"
)

# Example from the Alpaca 52K dataset:
example = {
    "instruction": "Give three tips for staying healthy.",
    "input": "",
    "output": "1. Eat a balanced diet...\n"
              "2. Exercise regularly...\n"
              "3. Get enough sleep..."
}
# That's it. No multi-turn. No system prompt. Just instruction → response.

The data: diversity is the secret ingredient

The Alpaca team found that their generated data was significantly more diverse than the original Self-Instruct dataset. They visualized this by plotting the root verbs and direct objects of all 52K instructions — the inner ring shows what the model is asked to do (write, explain, create, generate, describe, ...) and the outer ring shows what it acts on (a story, a list, a sentence, a code snippet, ...).

This diversity matters because models learn patterns from data: if most examples are "write a poem," the model overspecializes. Alpaca's diverse spread ensures broad coverage across task types, domains, and output formats, which is why a small 7B model can handle such a wide range of user requests after training.

Open in Lab
Explore the diversity of Alpaca's 52K instructions — inner ring shows verbs, outer ring shows objects.
The demo wakes as you arrive…

Evaluation: Alpaca vs. text-davinci-003

The Alpaca team evaluated their model using the Self-Instruct evaluation set — 252 diverse instructions covering a wide range of tasks. In a blind , five evaluators compared Alpaca 7B's outputs with those of text-davinci-003.

The result was striking: Alpaca 7B won or tied on 50% of the comparisons, despite being a 7B model fine-tuned on machine-generated data, competing against a model orders of magnitude more expensive to build. On many tasks — brainstorming, rewriting, basic question answering — evaluators genuinely could not distinguish the two models.

However, Alpaca had clear weaknesses. It struggled with complex reasoning, multi-step instructions, and factual accuracy. It also lacked safety training: the model could generate harmful or biased content because it had not undergone the reinforcement learning from human feedback stage that makes commercial models refuse dangerous requests.

Open in Lab
Compare Alpaca 7B against text-davinci-003 across different task categories.
The demo wakes as you arrive…

The bigger picture: knowledge distillation as a shortcut

What Alpaca really demonstrated is a form of knowledge distillation: using a strong teacher model (text-davinci-003) to transfer its instruction-following ability to a smaller (LLaMA 7B). The teacher doesn't share its weights — it shares its behavior through the 52K examples it generates.

This is a powerful pattern. The student model doesn't need to discover how to follow instructions from scratch; it learns by imitating the teacher's outputs. The result is a compressed version of the teacher's behavioral knowledge, packaged inside a much smaller and more accessible model.

However, this shortcut has a ceiling. The student can only approximate the teacher — it cannot exceed the teacher's ability on tasks where the training data is generated by that teacher. Later work like LIMA showed that with just 1,000 carefully curated human-written examples, you can get comparable results, suggesting that curation quality may matter even more than quantity.

What Alpaca unlocked

  1. 2022

    Self-Instruct

    Wang et al. proposed using a language model to generate its own instruction-following training data from a small set of human-written seeds.

  2. 2023

    LLaMA — Open Foundation Models

    Meta released LLaMA (7B–65B), strong base models available for research. These had no instruction-following ability but excellent base capabilities.

  3. 2023

    Alpaca — Cheap Instruction Tuning

    Stanford's recipe: Self-Instruct + LLaMA + \$600 = an instruction-following model competitive with GPT-3.5. Released data, code, and training recipe openly.

  4. 2023

    Vicuna, Koala, Dolly — The Open Wave

    Multiple teams replicated and improved the Alpaca recipe using different data sources (ShareGPT conversations, human-written data), proving the approach generalizes.

  5. 2023

    LIMA — Less Is More for Alignment

    Meta showed that just 1,000 carefully curated examples can match 52K generated ones, emphasizing curation quality over data quantity.

The idea in code

Generating instruction data with the Alpaca pipelinepython

Simplified to show the idea — not the real implementation.

import openai
import json
import random

# Step 1: Load the 175 human-written seed instructions
with open("seed_tasks.jsonl") as f:
    seed_tasks = [json.loads(line) for line in f]

# Step 2: Generate new instructions using text-davinci-003
def generate_instructions(seed_examples, num_to_generate=20):
    """Use in-context learning to generate new instruction-output pairs."""
    # Pick 3 random seeds as in-context examples
    examples = random.sample(seed_examples, 3)

    prompt = "You are asked to generate diverse task instructions.\n\n"
    for i, ex in enumerate(examples):
        prompt += f"###\n"
        prompt += f"{i+1}. Instruction: {ex['instruction']}\n"
        prompt += f"{i+1}. Output: {ex['output']}\n"
    prompt += f"###\nGenerate {num_to_generate} new instructions:\n"

    response = openai.Completion.create(
        model="text-davinci-003",
        prompt=prompt,
        max_tokens=2048,
        temperature=1.0,      # high temperature for diversity
        top_p=1.0,
    )
    return parse_instructions(response.choices[0].text)

# Step 3: Filter and deduplicate
# Uses ROUGE-L similarity to remove near-duplicates
# Total cost: ~\$500 for 52K examples

# Step 4: Fine-tune LLaMA 7B
# Standard supervised fine-tuning with Hugging Face
# LR=2e-5, batch=128, epochs=3, max_len=512
# Cost: ~\$100 on 4× A100 GPUs (3 hours)

Limitations and lessons

Alpaca was a proof of concept, not a production system. Its limitations reveal important lessons for the field:

: Like its teacher, Alpaca generates confident-sounding but incorrect answers. The model inherited text-davinci-003's knowledge but also its tendency to fabricate facts.

Safety: Without RLHF or content filtering, Alpaca would comply with harmful requests. This showed that instruction-following ability and safety alignment are orthogonal problems requiring different training stages.

Legal gray area: The 52K training examples were generated by a proprietary model (text-davinci-003) whose terms of service may prohibit using outputs to train competing models. Alpaca was released for research only, but it opened a broader debate about the intellectual property status of AI-generated training data.

Evaluation challenges: Human evaluation on 252 examples gives a rough signal but not a statistically rigorous comparison. Later benchmarks like AlpacaEval, MT-Bench, and Arena provided more systematic evaluation frameworks.

CitationTaori, Gulrajani, Zhang, Dubois, Li, Guestrin, Liang, Hashimoto. Alpaca: A Strong, Replicable Instruction-Following Model. Stanford CRFM Blog, 2023.

Terms in this paper