NLP2022intermediate14 min read

Self-Instruct: Aligning Language Models with Self-Generated Instructions

Self-Instruct: مواءمة النماذج اللغوية بتعليمات يولّدها النموذج بنفسه

Wang, Y. · Kordi, Y. · Mishra, S. · Liu, A. · Smith, N. A. · Khashabi, D. · Hajishirzi, H. — ACL

The problem

By late 2022, instruction-tuned language models like InstructGPT had shown remarkable ability to follow diverse instructions. But building these models required massive amounts of human-written instruction data — expensive to collect, limited in creativity, and often biased toward well-known NLP tasks like and summarization. This created a bottleneck: the model's generality was capped by the diversity of human annotations, and scaling up required even more costly labeling.

The contribution

Self-Instruct: a framework that lets a vanilla pretrained generate its own instruction-tuning data. Starting from only 175 hand-written seed tasks, the pipeline iteratively generates new instructions, classifies them, creates input-output instances, and filters low-quality examples — all using the same LM. Applied to GPT-3, it produced 52K diverse instructions and 82K instances. GPT-3 on this self-generated data yielded a 33% absolute improvement on Super-NaturalInstructions, nearly matching InstructGPT-001 which relied on private human-annotated data.

The impact

Self-Instruct demonstrated that language models contain enough knowledge to bootstrap their own instruction data — a finding that transformed the field. It directly inspired Stanford Alpaca, which applied the same idea to LLaMA and popularized open-source instruction-following models. The concept of synthetic instruction became a cornerstone technique, leading to LLaVA (multimodal instructions), Toolformer (tool-use instructions), and the broader movement toward RLAIF (AI-generated feedback replacing human feedback). The paper made accessible to anyone with API access to a capable language model.

Imagine a new teacher at a school who has read every textbook in the library but has never taught a class. The principal hands them a binder with 175 example lesson plans and says: "Now create thousands more." The teacher studies the examples, invents new lessons in subjects the examples never covered — creative writing, astronomy, cooking — writes sample exercises for each, throws away the ones that don't make sense, and then practices teaching from this self-made curriculum until they become a genuinely effective instructor.

No one told the teacher what to teach. The teacher figured it out from the pattern of what good lessons look like. Self-Instruct does this with a language model: it learns to follow instructions by generating its own data from a handful of seeds.

The bottleneck: human-written instructions do not scale

By 2022, the recipe for building instruction-following models was well established: take a pretrained language model, collect thousands of human-written (instruction, input, output) examples, and fine-tune the model on them. Projects like FLAN, T0, and Super-NaturalInstructions demonstrated that this instruction tuning unlocks impressive — the model can handle tasks it has never seen during training, as long as someone describes the task in natural language.

But the approach had a ceiling. Human annotators tend to generate familiar NLP tasks — , translation, — because those are what they know. Creative or unusual tasks — writing code, composing poetry, role-playing as a historical figure — are underrepresented. The annotators' imagination becomes the limit of what the model can do. Scaling up means hiring more annotators, which is expensive and still doesn't guarantee diversity.

InstructGPT (OpenAI) solved this partly by collecting real user queries, but that data was private, expensive to annotate with human feedback, and not reproducible by the research community. The field needed a way to generate diverse instruction data cheaply and openly.

The Self-Instruct pipeline: four steps to synthetic data

The core insight of Self-Instruct is deceptively simple: a pretrained language model already "knows" a vast number of tasks from its pretraining data — it just hasn't been organized to follow instructions about them. If you it carefully with a few example tasks, it can generate entirely new tasks, along with their inputs and outputs, that it has never been explicitly trained on.

The pipeline has four steps that run in an iterative loop. Think of it as a snowball rolling downhill — it starts with a small seed and grows larger with each pass.

Open in Lab
The four-step Self-Instruct pipeline. Click each step to see how it works — from seed tasks through instruction generation, classification, instance creation, and filtering.
The demo wakes as you arrive…

Step 1 — Instruction Generation. The process starts with a task pool seeded by 175 human-written tasks (each consisting of one instruction and one example). At each iteration, 8 tasks are randomly sampled from the pool — 6 from the original human-written set and 2 from previously generated tasks — and used as in-context examples to prompt the language model to generate a new instruction. This mix of human and machine examples promotes diversity while maintaining quality.

Step 2 — Classification Task Identification. Not all tasks are alike. Classification tasks (like sentiment analysis) have a small, fixed set of output labels, while generation tasks (like essay writing) produce open-ended text. The pipeline classifies each new instruction into one of these two categories using a prompt with 12 classification and 19 non-classification examples. This matters because the two types need different strategies for generating training instances.

Step 3 — Instance Generation. For each instruction, the model generates input-output pairs — the actual training examples. Non-classification tasks use an input-first approach: generate the input, then the output. But classification tasks tend to be biased when generated this way (e.g., a grammar checker mostly generates correct sentences). So classification tasks use an output-first approach: first generate the class labels, then generate an input example for each label. This balances the label distribution.

Step 4 — Filtering. Quality control happens on two fronts. First, diversity: a new instruction is only added to the pool if its -L similarity with every existing instruction is below 0.7. Second, validity: instances are discarded if the output repeats the input, if the instruction is too long or too short, or if it references capabilities the model doesn't have (like processing images). Instructions containing keywords like "image" or "picture" are also removed.

The bootstrapping loop: from 175 seeds to 52K tasks

The key to Self-Instruct is that the pipeline is iterative. Generated tasks that pass filtering are added back to the task pool, so the next round of generation draws from a richer, more diverse set of examples. This creates a virtuous cycle: more diversity in the pool leads to more creative generated tasks, which in turn further diversifies the pool.

Starting from 175 seed tasks, the pipeline generated 52,445 instructions with 82,439 input-output instances. Of these, 11,584 were classified as classification tasks and 40,861 as non-classification tasks. The average instruction was about 16 words long, inputs averaged 13 words, and outputs averaged 19 words.

The total cost of generating this entire using the GPT-3 API was approximately $600 — a fraction of what human at this scale would cost. Fine-tuning GPT-3 on the resulting data cost an additional $338.

Open in Lab
Watch the task pool grow from 175 seeds to thousands. Each iteration samples from the pool, generates new tasks, filters them, and feeds the good ones back in.
The demo wakes as you arrive…

Quality and diversity: how good is self-generated data?

A natural worry is that machine-generated data might be repetitive or low-quality. The authors investigated both concerns carefully.

Diversity: The generated instructions cover a wide range of task types. When analyzed by verb-noun structure, the top 20 root verbs and their noun objects account for only 14% of all instructions — meaning the remaining 86% use less common formulations. Tasks span from creative writing to , math problems to email composition, classification to open-ended brainstorming. The generated instructions also diverge significantly from the 175 seeds: most have low ROUGE-L overlap with their nearest seed instruction, confirming that the model is genuinely creating novel tasks rather than paraphrasing the seeds.

Quality: A manual review of 200 random instructions found that 92% described valid tasks, 79% had appropriate inputs, and 58% had fully correct outputs. Overall, 54% of sampled instances were valid across all three fields. While this means roughly half the data has some issue, the authors found that even partially correct examples — those with the right format but imperfect content — still provide useful training signal for teaching models to follow instructions.

Open in Lab
Explore the diversity of generated instructions. The inner ring shows the most common root verbs; the outer ring shows their noun objects. These top patterns cover only 14% of all generated instructions.
The demo wakes as you arrive…

Input-first vs output-first: two strategies for generating examples

One of the paper's practical innovations is recognizing that different task types need different generation strategies. Consider a grammar error detection task: if you ask the model to generate an input sentence first, it tends to produce grammatically correct sentences — because that's what language models naturally do. The resulting dataset would be heavily biased toward the "no error" class.

The solution is the output-first approach for classification tasks. Instead of asking "generate an input, then classify it," the pipeline says "here are the possible labels (e.g., grammatical / ungrammatical). Now generate an input for each label." This forces balanced coverage of all classes.

For non-classification tasks — essay writing, code generation, — the standard input-first approach works well. The model generates a plausible input (like a topic for an essay), then produces the corresponding output (the essay itself). This mirrors the natural flow of how such tasks are used in practice.

Open in Lab
Compare the two generation strategies. See how input-first creates biased classification data, while output-first ensures balanced labels.
The demo wakes as you arrive…

Formal framework: instruction data as task definitions

Formally, the instruction data consists of a set of instructions {It}\{I_t\}, where each instruction ItI_t defines a task tt in natural language. Each task has one or more input-output instances {(Xt,i,Yt,i)}i=1nt\{(X_{t,i}, Y_{t,i})\}_{i=1}^{n_t}. The model MM is trained to satisfy:

M(It,Xt,i)=Yt,i,∀ i∈{1,…,nt}M(I_t, X_{t,i}) = Y_{t,i}, \quad \forall\, i \in \{1, \dots, n_t\}
Instruction-following objective — the model produces the correct output given instruction and input — Given a task instruction ItI_t and an input Xt,iX_{t,i}, the model should produce output Yt,iY_{t,i}. Note that some tasks have no input (XX is empty) — for instance, "write an essay about climate change" is a complete instruction with no additional input needed.

The ROUGE-L filtering threshold ensures diversity by rejecting any new instruction InewI_{\text{new}} whose similarity to existing instructions exceeds a threshold:

ROUGE-L(Inew, Iexisting)<0.7∀ Iexisting∈P\text{ROUGE-L}(I_{\text{new}},\, I_{\text{existing}}) < 0.7 \quad \forall\, I_{\text{existing}} \in \mathcal{P}
Diversity filter — new instructions must differ from all existing ones — P\mathcal{P} is the task pool. Only instructions sufficiently different from everything already in the pool are admitted. ROUGE-L measures the longest common subsequence between two texts, so a threshold of 0.7 means at most 70% structural overlap is tolerated.

Results: closing the gap with private data

The authors evaluated Self-Instruct in two settings: automated benchmarks and expert .

Super-NaturalInstructions . Fine-tuning GPT-3 on the self-generated data (called GPT3SELF-INST_{\text{SELF-INST}}) improved ROUGE-L from 6.8 to 39.9 — a 33.1 point absolute improvement. This nearly matched InstructGPT-001 (40.8), which was trained on proprietary human-annotated data. GPT3SELF-INST_{\text{SELF-INST}} also outperformed GPT-3 fine-tuned on the T0 training data (37.9), despite T0's data requiring tremendous human effort to create.

Human evaluation on 252 novel tasks. The authors curated 252 user-oriented instructions spanning domains like email writing, programming, entertainment, and productivity. Expert evaluators rated responses on a four-level scale (A: satisfying, B: minor errors, C: significant errors, D: irrelevant). GPT3SELF-INST_{\text{SELF-INST}} outperformed all models trained on publicly available instruction data. The gap to InstructGPT-001 was only 5% when counting both A and B ratings as acceptable.

Scaling and quality. Performance improved consistently with more generated data, though gains plateaued around 16K instructions. Replacing the model's own outputs with outputs from InstructGPT-003 (a form of distillation) further boosted performance by 10%, suggesting room for improvement through better output quality.

Open in Lab
Compare model performance on Super-NaturalInstructions. Self-Instruct bridges most of the gap between vanilla GPT-3 and InstructGPT trained on private data.
The demo wakes as you arrive…

Prompting strategy: how the model generates new tasks

The prompting template for instruction generation is surprisingly simple. The model is shown a list of existing tasks and asked to continue the pattern. This works because GPT-3, having been trained on massive amounts of text including task descriptions and tutorials, can generalize the pattern of "here are some tasks, generate more like them."

Instruction generation prompt template (simplified)python

Simplified to show the idea — not the real implementation.

# The model sees 8 existing tasks and generates new ones
prompt = """Come up with a series of tasks:
Task 1: {sampled_instruction_1}
Task 2: {sampled_instruction_2}
Task 3: {sampled_instruction_3}
Task 4: {sampled_instruction_4}
Task 5: {sampled_instruction_5}
Task 6: {sampled_instruction_6}
Task 7: {sampled_instruction_7}
Task 8: {sampled_instruction_8}
Task 9:"""
# 6 tasks from human seeds + 2 from prior generated tasks
# Model generates new instructions until it stops

Scaling analysis: more data, better quality

The authors explored two axes of improvement: generating more data and improving data quality.

On the quantity axis, performance improves steadily from 175 instructions (just the seeds) to around 16K instructions, after which gains plateau. This aligns with findings from other instruction-tuning work — a few thousand diverse instructions are enough to teach the format, and adding more of the same quality doesn't help much.

On the quality axis, the authors tried a distillation experiment: they kept the same instructions and inputs but replaced the GPT-3-generated outputs with outputs from InstructGPT-003 (a much stronger model). This improved performance by approximately 10 percentage points, suggesting that the pipeline's value lies primarily in generating diverse instructions and inputs, while the outputs can be improved by a better teacher model.

Open in Lab
Performance improves with more instruction data but plateaus around 16K. Improving output quality via distillation provides an additional 10% gain.
The demo wakes as you arrive…

Why this matters: democratizing instruction tuning

Before Self-Instruct, building an instruction-following model required either access to expensive proprietary data (like OpenAI's user queries) or massive efforts (like the 1,600+ tasks in Super-NaturalInstructions). Self-Instruct showed a third path: let the model generate its own training data. This was transformative for three reasons.

First, it made instruction tuning accessible. Anyone with API access to a capable language model could generate their own instruction dataset for approximately $600. No crowdsourcing platform, no expert annotators, no proprietary data needed.

Second, it shifted the focus from quantity of human effort to quality of seed design. The 175 seed tasks matter more than any individual generated example — they set the distribution of what gets generated. This is a much more tractable problem than writing 50,000 instructions by hand.

Third, it revealed that pretrained language models have far more task knowledge than their zero-shot performance suggests. The knowledge exists in the weights; it just needs to be elicited through the right prompting strategy and then organized through fine-tuning.

Legacy: the synthetic data revolution

  1. 2022

    Self-Instruct (this paper)

    Demonstrated that a vanilla GPT-3 can bootstrap its own instruction data from 175 seeds, generating 52K instructions at ~\$600 and nearly matching InstructGPT-001.

  2. 2023

    Stanford Alpaca

    Applied Self-Instruct to Meta's LLaMA model, generating 52K instructions using GPT-3.5-Turbo for only \$500. Created the first widely-used open-source instruction-following model.

  3. 2023

    LLaVA — multimodal Self-Instruct

    Extended the Self-Instruct idea to vision-language tasks, generating instruction data for image understanding. Showed the bootstrapping concept generalizes beyond text.

  4. 2023

    Toolformer — self-generated tool-use data

    Applied bootstrapping to teach models to use external tools (calculators, search engines, APIs) by generating tool-use annotations from the model itself.

  5. 2023

    RLAIF: AI feedback replaces human feedback

    Constitutional AI and related work showed that AI-generated preference labels can replace human annotators in RLHF, extending the self-generation principle from instructions to feedback signals.

Self-Instruct planted a seed — quite literally — that grew into one of the most consequential ideas in modern AI: that language models can teach themselves. What started as a $600 experiment with 175 hand-written tasks has become the foundation for how the open-source community builds instruction-following models. The paper's greatest contribution is not the specific pipeline, but the demonstration that the data bottleneck can be broken by the model itself.

CitationWang, Kordi, Mishra, Liu, Smith, Khashabi, Hajishirzi. Self-Instruct: Aligning Language Models with Self-Generated Instructions. ACL, 2023.

Terms in this paper