NLP2022intermediate14 min read
Self-Instruct: Aligning Language Models with Self-Generated Instructions
Self-Instruct: مواءمة النماذج اللغوية بتعليمات يولّدها النموذج بنفسه
Wang, Y. · Kordi, Y. · Mishra, S. · Liu, A. · Smith, N. A. · Khashabi, D. · Hajishirzi, H. — ACL
The problem
By late 2022, instruction-tuned language models like InstructGPT had shown remarkable ability to follow diverse instructions. But building these models required massive amounts of human-written instruction data — expensive to collect, limited in creativity, and often biased toward well-known NLP tasks like and summarization. This created a bottleneck: the model's generality was capped by the diversity of human annotations, and scaling up required even more costly labeling.
The contribution
Self-Instruct: a framework that lets a vanilla pretrained generate its own instruction-tuning data. Starting from only 175 hand-written seed tasks, the pipeline iteratively generates new instructions, classifies them, creates input-output instances, and filters low-quality examples — all using the same LM. Applied to GPT-3, it produced 52K diverse instructions and 82K instances. GPT-3 on this self-generated data yielded a 33% absolute improvement on Super-NaturalInstructions, nearly matching InstructGPT-001 which relied on private human-annotated data.
The impact
Self-Instruct demonstrated that language models contain enough knowledge to bootstrap their own instruction data — a finding that transformed the field. It directly inspired Stanford Alpaca, which applied the same idea to LLaMA and popularized open-source instruction-following models. The concept of synthetic instruction became a cornerstone technique, leading to LLaVA (multimodal instructions), Toolformer (tool-use instructions), and the broader movement toward RLAIF (AI-generated feedback replacing human feedback). The paper made accessible to anyone with API access to a capable language model.
Imagine a new teacher at a school who has read every textbook in the library but has never taught a class. The principal hands them a binder with 175 example lesson plans and says: "Now create thousands more." The teacher studies the examples, invents new lessons in subjects the examples never covered — creative writing, astronomy, cooking — writes sample exercises for each, throws away the ones that don't make sense, and then practices teaching from this self-made curriculum until they become a genuinely effective instructor.
No one told the teacher what to teach. The teacher figured it out from the pattern of what good lessons look like. Self-Instruct does this with a language model: it learns to follow instructions by generating its own data from a handful of seeds.
The bottleneck: human-written instructions do not scale
By 2022, the recipe for building instruction-following models was well established: take a pretrained language model, collect thousands of human-written (instruction, input, output) examples, and fine-tune the model on them. Projects like FLAN, T0, and Super-NaturalInstructions demonstrated that this instruction tuning unlocks impressive — the model can handle tasks it has never seen during training, as long as someone describes the task in natural language.
But the approach had a ceiling. Human annotators tend to generate familiar NLP tasks — , translation, — because those are what they know. Creative or unusual tasks — writing code, composing poetry, role-playing as a historical figure — are underrepresented. The annotators' imagination becomes the limit of what the model can do. Scaling up means hiring more annotators, which is expensive and still doesn't guarantee diversity.
InstructGPT (OpenAI) solved this partly by collecting real user queries, but that data was private, expensive to annotate with human feedback, and not reproducible by the research community. The field needed a way to generate diverse instruction data cheaply and openly.
The Self-Instruct pipeline: four steps to synthetic data
The core insight of Self-Instruct is deceptively simple: a pretrained language model already "knows" a vast number of tasks from its pretraining data — it just hasn't been organized to follow instructions about them. If you it carefully with a few example tasks, it can generate entirely new tasks, along with their inputs and outputs, that it has never been explicitly trained on.
The pipeline has four steps that run in an iterative loop. Think of it as a snowball rolling downhill — it starts with a small seed and grows larger with each pass.
Step 1 — Instruction Generation. The process starts with a task pool seeded by 175 human-written tasks (each consisting of one instruction and one example). At each iteration, 8 tasks are randomly sampled from the pool — 6 from the original human-written set and 2 from previously generated tasks — and used as in-context examples to prompt the language model to generate a new instruction. This mix of human and machine examples promotes diversity while maintaining quality.
Step 2 — Classification Task Identification. Not all tasks are alike. Classification tasks (like sentiment analysis) have a small, fixed set of output labels, while generation tasks (like essay writing) produce open-ended text. The pipeline classifies each new instruction into one of these two categories using a prompt with 12 classification and 19 non-classification examples. This matters because the two types need different strategies for generating training instances.
Step 3 — Instance Generation. For each instruction, the model generates input-output pairs — the actual training examples. Non-classification tasks use an input-first approach: generate the input, then the output. But classification tasks tend to be biased when generated this way (e.g., a grammar checker mostly generates correct sentences). So classification tasks use an output-first approach: first generate the class labels, then generate an input example for each label. This balances the label distribution.
Step 4 — Filtering. Quality control happens on two fronts. First, diversity: a new instruction is only added to the pool if its -L similarity with every existing instruction is below 0.7. Second, validity: instances are discarded if the output repeats the input, if the instruction is too long or too short, or if it references capabilities the model doesn't have (like processing images). Instructions containing keywords like "image" or "picture" are also removed.
The bootstrapping loop: from 175 seeds to 52K tasks
The key to Self-Instruct is that the pipeline is iterative. Generated tasks that pass filtering are added back to the task pool, so the next round of generation draws from a richer, more diverse set of examples. This creates a virtuous cycle: more diversity in the pool leads to more creative generated tasks, which in turn further diversifies the pool.
Starting from 175 seed tasks, the pipeline generated 52,445 instructions with 82,439 input-output instances. Of these, 11,584 were classified as classification tasks and 40,861 as non-classification tasks. The average instruction was about 16 words long, inputs averaged 13 words, and outputs averaged 19 words.
The total cost of generating this entire using the GPT-3 API was approximately $600 — a fraction of what human at this scale would cost. Fine-tuning GPT-3 on the resulting data cost an additional $338.
Quality and diversity: how good is self-generated data?
A natural worry is that machine-generated data might be repetitive or low-quality. The authors investigated both concerns carefully.
Diversity: The generated instructions cover a wide range of task types. When analyzed by verb-noun structure, the top 20 root verbs and their noun objects account for only 14% of all instructions — meaning the remaining 86% use less common formulations. Tasks span from creative writing to , math problems to email composition, classification to open-ended brainstorming. The generated instructions also diverge significantly from the 175 seeds: most have low ROUGE-L overlap with their nearest seed instruction, confirming that the model is genuinely creating novel tasks rather than paraphrasing the seeds.
Quality: A manual review of 200 random instructions found that 92% described valid tasks, 79% had appropriate inputs, and 58% had fully correct outputs. Overall, 54% of sampled instances were valid across all three fields. While this means roughly half the data has some issue, the authors found that even partially correct examples — those with the right format but imperfect content — still provide useful training signal for teaching models to follow instructions.
Input-first vs output-first: two strategies for generating examples
One of the paper's practical innovations is recognizing that different task types need different generation strategies. Consider a grammar error detection task: if you ask the model to generate an input sentence first, it tends to produce grammatically correct sentences — because that's what language models naturally do. The resulting dataset would be heavily biased toward the "no error" class.
The solution is the output-first approach for classification tasks. Instead of asking "generate an input, then classify it," the pipeline says "here are the possible labels (e.g., grammatical / ungrammatical). Now generate an input for each label." This forces balanced coverage of all classes.
For non-classification tasks — essay writing, code generation, — the standard input-first approach works well. The model generates a plausible input (like a topic for an essay), then produces the corresponding output (the essay itself). This mirrors the natural flow of how such tasks are used in practice.
Formal framework: instruction data as task definitions
Formally, the instruction data consists of a set of instructions , where each instruction defines a task in natural language. Each task has one or more input-output instances . The model is trained to satisfy:
The ROUGE-L filtering threshold ensures diversity by rejecting any new instruction whose similarity to existing instructions exceeds a threshold:
Results: closing the gap with private data
The authors evaluated Self-Instruct in two settings: automated benchmarks and expert .
Super-NaturalInstructions . Fine-tuning GPT-3 on the self-generated data (called GPT3) improved ROUGE-L from 6.8 to 39.9 — a 33.1 point absolute improvement. This nearly matched InstructGPT-001 (40.8), which was trained on proprietary human-annotated data. GPT3 also outperformed GPT-3 fine-tuned on the T0 training data (37.9), despite T0's data requiring tremendous human effort to create.
Human evaluation on 252 novel tasks. The authors curated 252 user-oriented instructions spanning domains like email writing, programming, entertainment, and productivity. Expert evaluators rated responses on a four-level scale (A: satisfying, B: minor errors, C: significant errors, D: irrelevant). GPT3 outperformed all models trained on publicly available instruction data. The gap to InstructGPT-001 was only 5% when counting both A and B ratings as acceptable.
Scaling and quality. Performance improved consistently with more generated data, though gains plateaued around 16K instructions. Replacing the model's own outputs with outputs from InstructGPT-003 (a form of distillation) further boosted performance by 10%, suggesting room for improvement through better output quality.
Prompting strategy: how the model generates new tasks
The prompting template for instruction generation is surprisingly simple. The model is shown a list of existing tasks and asked to continue the pattern. This works because GPT-3, having been trained on massive amounts of text including task descriptions and tutorials, can generalize the pattern of "here are some tasks, generate more like them."
Simplified to show the idea — not the real implementation.
# The model sees 8 existing tasks and generates new ones
prompt = """Come up with a series of tasks:
Task 1: {sampled_instruction_1}
Task 2: {sampled_instruction_2}
Task 3: {sampled_instruction_3}
Task 4: {sampled_instruction_4}
Task 5: {sampled_instruction_5}
Task 6: {sampled_instruction_6}
Task 7: {sampled_instruction_7}
Task 8: {sampled_instruction_8}
Task 9:"""
# 6 tasks from human seeds + 2 from prior generated tasks
# Model generates new instructions until it stopsScaling analysis: more data, better quality
The authors explored two axes of improvement: generating more data and improving data quality.
On the quantity axis, performance improves steadily from 175 instructions (just the seeds) to around 16K instructions, after which gains plateau. This aligns with findings from other instruction-tuning work — a few thousand diverse instructions are enough to teach the format, and adding more of the same quality doesn't help much.
On the quality axis, the authors tried a distillation experiment: they kept the same instructions and inputs but replaced the GPT-3-generated outputs with outputs from InstructGPT-003 (a much stronger model). This improved performance by approximately 10 percentage points, suggesting that the pipeline's value lies primarily in generating diverse instructions and inputs, while the outputs can be improved by a better teacher model.
Why this matters: democratizing instruction tuning
Before Self-Instruct, building an instruction-following model required either access to expensive proprietary data (like OpenAI's user queries) or massive efforts (like the 1,600+ tasks in Super-NaturalInstructions). Self-Instruct showed a third path: let the model generate its own training data. This was transformative for three reasons.
First, it made instruction tuning accessible. Anyone with API access to a capable language model could generate their own instruction dataset for approximately $600. No crowdsourcing platform, no expert annotators, no proprietary data needed.
Second, it shifted the focus from quantity of human effort to quality of seed design. The 175 seed tasks matter more than any individual generated example — they set the distribution of what gets generated. This is a much more tractable problem than writing 50,000 instructions by hand.
Third, it revealed that pretrained language models have far more task knowledge than their zero-shot performance suggests. The knowledge exists in the weights; it just needs to be elicited through the right prompting strategy and then organized through fine-tuning.
Legacy: the synthetic data revolution
2022
Self-Instruct (this paper)
Demonstrated that a vanilla GPT-3 can bootstrap its own instruction data from 175 seeds, generating 52K instructions at ~\$600 and nearly matching InstructGPT-001.
2023
Stanford Alpaca
Applied Self-Instruct to Meta's LLaMA model, generating 52K instructions using GPT-3.5-Turbo for only \$500. Created the first widely-used open-source instruction-following model.
2023
LLaVA — multimodal Self-Instruct
Extended the Self-Instruct idea to vision-language tasks, generating instruction data for image understanding. Showed the bootstrapping concept generalizes beyond text.
2023
Toolformer — self-generated tool-use data
Applied bootstrapping to teach models to use external tools (calculators, search engines, APIs) by generating tool-use annotations from the model itself.
2023
RLAIF: AI feedback replaces human feedback
Constitutional AI and related work showed that AI-generated preference labels can replace human annotators in RLHF, extending the self-generation principle from instructions to feedback signals.
Self-Instruct planted a seed — quite literally — that grew into one of the most consequential ideas in modern AI: that language models can teach themselves. What started as a $600 experiment with 175 hand-written tasks has become the foundation for how the open-source community builds instruction-following models. The paper's greatest contribution is not the specific pipeline, but the demonstration that the data bottleneck can be broken by the model itself.
CitationWang, Kordi, Mishra, Liu, Smith, Khashabi, Hajishirzi. Self-Instruct: Aligning Language Models with Self-Generated Instructions. ACL, 2023.
Terms in this paper
- Instruction Tuningالضبط التعليمي
- Bootstrappingالتمهيد الذاتي
- Synthetic Dataالبيانات الاصطناعية
- Fine-Tuningالضبط الدقيق
- Few-Shotالنمط القليل العيّنات
- In-Context Learningالتعلم في السياق
- Classificationالتصنيف
- ROUGEمعيار روج الإحصائي
- Language Modelالنموذج اللغوي
- Generationالتوليد الآلي
- Promptالنص التوجيهي (الـمُحفّز)
- Downstream Taskالمهمة اللاحقة
- Data Augmentationتعزيز البيانات
- Knowledge Distillationتقطير المعرفة
- Zero-Shot Transferنقل بدون تدريب