Language Models2022intermediate12 min read
Finetuned Language Models Are Zero-Shot Learners
النماذج اللغوية المضبوطة بالتعليمات تتعلَّم بدون أمثلة
Wei, J. · Bosma, M. · Zhao, V. Y. · Guu, K. · Yu, A. W. · Lester, B. · Du, N. · Dai, A. M. · Le, Q. V. — ICLR
The problem
By 2021, large language models like GPT-3 excelled at — give them a handful of examples and they perform well. But their zero-shot performance lagged badly: without examples to anchor the format, the model struggled to understand what was being asked. The model needed the crutch of examples because its taught it to complete text, not follow instructions. Meanwhile, each new task still demanded either engineering or a separate finetuned model.
The contribution
: finetune a 137B- pretrained model (LaMDA-PT) on over 60 NLP datasets phrased as natural language instructions across 12 task clusters. The resulting model, FLAN, achieves strong zero-shot performance on unseen task types — outperforming zero-shot GPT-3 (175B) on 20 of 25 evaluated tasks. Ablation studies show that performance improves with more task clusters, that benefits emerge only at sufficient model scale (≥68B), and that natural language instructions during finetuning are critical.
The impact
FLAN established instruction tuning as a foundational technique for aligning large language models with human intent. It directly inspired InstructGPT, ChatGPT, and every instruction-following model that followed. The insight that multi-task finetuning with natural language instructions unlocks zero-shot is now the standard recipe behind Claude, GPT-4, Gemini, and modern AI assistants.
A GPT-3 without examples is like a brilliant new employee who has read every book in the company library, but has never been given an actual work assignment. Hand them a customer complaint and they might continue writing the complaint instead of answering it — they know about tasks but have never done one.
Instruction tuning is orientation week: you walk them through dozens of real tasks — file this report, answer that ticket, translate this email — each time giving a clear instruction. After orientation, hand them a completely new type of task they've never seen, and they know what to do. They've learned to follow instructions, not just mimic text.
The problem: knowing text is not the same as following instructions
By 2021, the NLP world had three paradigms, and each had a gap:
-
Pretrain → finetune (like BERT and T5): you need labeled data for every new task, and each task gets its own separate model. Works well, but doesn't scale to the endless variety of tasks people actually want to do.
-
Prompting (like GPT-3): one giant model handles many tasks via clever prompts. Few-shot prompting (giving 10–100 examples in context) works impressively, but zero-shot prompting — just asking the model in plain language — often fails. The model was pretrained to complete documents, not answer questions.
-
Missing: a single model that can handle a broad range of tasks just from reading instructions, without needing examples or task-specific training.
The idea: teach a model to follow instructions by practicing on many tasks
The idea behind instruction tuning is elegantly simple. Take a large pretrained language model and finetune it on a diverse collection of NLP tasks — but instead of feeding raw input-output pairs, describe each task using natural language instructions.
For example, instead of training on:
Input: "The movie was wonderful" → Output: "Positive"
You train on:
Input: "Is the following movie review positive or negative? 'The movie was wonderful' OPTIONS: - positive - negative" → Output: "Positive"
The authors collected 62 publicly available NLP datasets, grouped them into 12 task clusters (NLI, sentiment, translation, QA, , etc.), and wrote 10 different instruction templates per dataset. The model they finetuned was LaMDA-PT, a 137B-parameter -only .
Rigorous evaluation: cluster-level hold-out
How do you prove a model truly generalizes to unseen tasks? The authors used a strict cluster-level hold-out strategy. They didn't just hold out individual datasets — they held out entire task types.
To evaluate FLAN on (NLI), they removed all NLI datasets from instruction tuning and trained on everything else. This means the model never saw any form of NLI during training — not a single dataset about determining whether a hypothesis follows from a premise.
This is critical: if you only hold out one NLI dataset but train on others, the model has already learned "what NLI is." By holding out the entire cluster, you test whether instruction tuning teaches the model to follow instructions rather than just memorize task formats.
Results: FLAN beats zero-shot GPT-3 — and sometimes even few-shot GPT-3
FLAN's zero-shot performance told a striking story across task categories:
-
Natural language : FLAN dramatically outperformed GPT-3 on all five NLI datasets (ANLI R1–R3, CB, RTE). Why? NLI examples are awkwardly phrased as text continuations — "Does 'X' mean that 'Y'?" is much more natural as an instruction than as something you'd find in a web document.
-
Reading comprehension: FLAN outperformed GPT-3 on BoolQ, MultiRC, and OpenbookQA. The instruction format ("Read the passage and answer the question") maps naturally to how humans would describe these tasks.
-
Closed-book QA: FLAN outperformed GPT-3 on all four datasets (ARC-easy, ARC-challenge, Natural Questions, TriviaQA).
-
Translation: FLAN outperformed zero-shot GPT-3 on all six language pairs, though it fell short of few-shot GPT-3 in most translation directions — likely because the model's and pretraining data were predominantly English.
However, instruction tuning was less effective for tasks that already look like language modeling — such as commonsense reasoning and formatted as sentence completion, where instructions add little signal beyond what the format already conveys.
What makes instruction tuning work? Three ablations
The paper's ablation studies revealed three crucial factors:
1. More task clusters → better generalization. Adding more types of tasks to instruction tuning steadily improved performance on held-out clusters. With only 1 cluster (summarization), performance was modest. Adding translation, reading comprehension, sentiment, and more — each new cluster pushed the average higher. Performance showed no sign of saturating even at 7 clusters, suggesting that even more diversity could help.
2. Scale is a prerequisite. Instruction tuning only helps large models. At 422M and 2B parameters, instruction tuning hurt zero-shot performance on held-out tasks — the model used all its capacity to memorize the training tasks and had nothing left for generalization. At 8B, results were mixed. Only at 68B and 137B did instruction tuning consistently improve generalization. The insight: small models learn tasks; large models learn how to follow instructions.
3. Natural language instructions are essential. Three variants were compared: (a) no template at all (just raw inputs and outputs), (b) a dataset name tag (e.g., "[Translation: WMT'14]"), and (c) natural instructions (e.g., "Translate this sentence to French"). FLAN with natural instructions substantially outperformed both ablations, confirming that the way tasks are described matters as much as the tasks themselves.
A practical trick: the OPTIONS suffix
For classification tasks, FLAN introduced a simple but effective trick: appending an OPTIONS suffix that lists the valid output classes. For example:
"Premise: At my age you will probably have learnt one lesson. Hypothesis: It's not certain how many lessons you'll learn by your thirties. Does the premise entail the hypothesis? OPTIONS: - yes - it is not possible to tell - no"
Why does this matter? Without OPTIONS, a decoder-only model might split its probability mass across many ways of saying "yes" ("Yes", "True", "Correct", "It is true", etc.), making rank classification unreliable. The OPTIONS suffix focuses the model's output distribution on the exact valid choices, acting like a menu that narrows the search space for the model.
Under the hood: training details
FLAN was built on LaMDA-PT, a 137B-parameter decoder-only Transformer pretrained on 2.49 trillion tokens from web documents, dialog data, and Wikipedia. Around 10% of its pretraining data was non-English.
The instruction tuning procedure was surprisingly lightweight compared to pretraining: 30,000 steps with a batch size of 8,192 tokens, using the Adafactor with a of 3e-5. To balance different dataset sizes, each dataset was capped at 30,000 examples with a mixing rate maximum of 3,000. The entire instruction tuning took roughly 60 hours on a TPUv3 with 128 cores — less than 2% of the pretraining compute.
This efficiency is a key insight: instruction tuning is not a second round of massive pretraining. It's a lightweight step that reorients the model's existing knowledge toward following instructions.
Few-shot + instruction tuning: complementary, not competing
The paper also showed that instruction tuning and few-shot prompting are complementary. When few-shot exemplars are added at inference time, FLAN's performance improves further across all task clusters, especially for tasks with complex output formats like structured text generation and translation. The standard deviation across templates also drops with few-shot exemplars, meaning the model becomes less sensitive to how exactly the instruction is phrased.
Additionally, FLAN proved more responsive to — a technique that learns continuous "soft" prompt embeddings via . With only 32 training examples, prompt tuning on FLAN achieved over 10% improvement compared to prompt tuning on the base LaMDA-PT. This suggests that instruction tuning produces a model that is fundamentally more "ready" to be steered toward new tasks by any prompting method.
Limitations and honest assessment
The authors were transparent about FLAN's limitations:
-
Not universal. Instruction tuning did not help for tasks that already look like natural language modeling (sentence completion, coreference as gap-filling). When the task format itself already tells the model what to do, instructions add little.
-
Scale dependency. The technique only works at large model sizes (≥68B). For smaller models, instruction tuning hurts zero-shot generalization — the model memorizes training tasks instead of learning to generalize.
-
Subjectivity in clustering. Grouping datasets into task clusters involves judgment calls. The boundary between "reading comprehension" and "closed-book QA" is not always clear, and different groupings might yield different conclusions.
-
Short instructions only. The paper used single-sentence instructions, far simpler than the detailed multi-paragraph instructions given to human crowd-workers. Whether richer instructions would help further remains an open question.
-
Cost to serve. At 137B parameters, FLAN is expensive to deploy, limiting its practical accessibility despite its strong performance.
Why it mattered: the bridge to modern AI assistants
2020
GPT-3 — few-shot prompting
GPT-3 showed that large models can perform tasks from a few examples in context, but zero-shot performance remained weak — the model needed examples as format anchors.
2021
FLAN — instruction tuning
Showed that finetuning on diverse tasks described via instructions unlocks zero-shot generalization. A 137B model outperforms a larger 175B model on most tasks.
2021
T0 — multitask prompted training
Sanh et al. independently explored a similar idea with T5-11B, confirming that instruction-style finetuning improves zero-shot generalization at smaller scales too.
2022
InstructGPT — instruction tuning + RLHF
OpenAI combined instruction tuning with reinforcement learning from human feedback, creating models that follow instructions more safely and helpfully.
2022
ChatGPT — instruction-tuned models go mainstream
The conversational product that brought instruction-following models to 100 million users, built on the foundation that FLAN and InstructGPT established.
2023
FLAN-T5 & FLAN-PaLM — scaling instruction tuning
Google scaled instruction tuning to 1,800+ tasks, showing continued improvements and releasing open-source FLAN-T5 models that became widely adopted baselines.
FLAN showed that the path from "language model" to "instruction-following assistant" does not require a fundamentally new architecture — it requires teaching the model a new interface. Pretraining builds the knowledge; instruction tuning builds the ability to use that knowledge on demand. This decomposition — knowledge acquisition separate from instruction following — is the blueprint that every modern AI system follows.
CitationWei, Bosma, Zhao, Guu, Yu, Lester, Du, Dai, Le. Finetuned Language Models Are Zero-Shot Learners. ICLR, 2022.
Terms in this paper
- Instruction Tuningالضبط التعليمي
- Zero-Shot Learningالتعلّم بدون أمثلة
- Few-Shot Learningالتعلّم بأمثلة قليلة
- Task Clusterعنقود المهام
- Instruction Templateقالب تعليمات
- Cross-Task Generalizationالتعميم عبر المهام
- Multi-Task Learningالتعلّم متعدد المهام