Reinforcement Learning2023intermediate11 min read
LIMA: Less Is More for Alignment
LIMA: القليل يكفي للمواءمة
Zhou, C. · Liu, P. · Xu, P. · Iyer, S. · Sun, J. · Mao, Y. · Ma, X. · Efrat, A. · Yu, P. · Yu, L. · Zhang, S. · Ghosh, G. · Lewis, M. · Zettlemoyer, L. · Levy, O. — NeurIPS
The problem
By 2023, aligning large language models to follow instructions and match user preferences had become an expensive, multi-stage pipeline. First came large-scale on tens of thousands of examples, then (RLHF) using millions of annotator interactions. The conventional wisdom was that both stages were essential and that more data always meant better . No one had rigorously tested whether all that expense was truly necessary, or whether a simpler approach could suffice.
The contribution
The authors proposed the : almost all of a model's knowledge and capabilities are learned during , while alignment merely teaches the model which output style to use when interacting with users. To test this, they created LIMA — a 65B-parameter LLaMA model fine-tuned on only 1,000 carefully curated examples using standard supervised loss, with zero reinforcement learning. LIMA matched or outperformed models trained with RLHF on far more data, and ablations showed that and diversity matter far more than quantity for alignment.
The impact
LIMA reframed the alignment debate. It showed that expensive RLHF pipelines may be solving a much simpler problem than assumed — teaching style rather than knowledge. This insight accelerated research into efficient alignment methods like DPO and led teams to invest more in data curation than in scaling annotation. The Superficial Alignment Hypothesis became a widely cited framework for understanding why small, high-quality datasets can outperform massive but noisy ones.
Imagine training a world-class chef. You could spend years teaching them thousands of recipes, correcting every seasoning mistake, and having food critics grade every dish. Or — if the chef already has deep culinary intuition from years of tasting and studying — you could simply show them 1,000 beautifully plated dishes and say: "This is the style." The chef's existing knowledge does the rest.
LIMA proved that large language models are exactly this chef. The knowledge lives in pretraining. Alignment is just plating.
The problem: why was alignment so expensive?
By 2023, the standard recipe for building a capable AI assistant had two expensive stages after pretraining. First, instruction tuning: the model on tens or hundreds of thousands of -response pairs to teach it to follow instructions. Second, reinforcement learning from (RLHF): training a from human preferences, then using it to further optimize the language model's outputs. Projects like InstructGPT collected millions of human annotations. Alpaca used 52,000 distilled examples. The implicit assumption was clear: more data, more human feedback, more training — better alignment.
But this assumption had never been rigorously tested. Is all that data really teaching the model new knowledge, or is it merely teaching the model how to present knowledge it already has? This is the question LIMA set out to answer.
The core idea: the Superficial Alignment Hypothesis
The paper's central thesis is the Superficial Alignment Hypothesis: a model's knowledge and capabilities are learned almost entirely during pretraining, while alignment merely teaches the model which subdistribution of formats should be used when interacting with users. In other words, alignment is not about injecting knowledge — it's about unlocking a style.
Think of pretraining as building a massive library inside the model's weights — billions of facts, reasoning patterns, writing styles, code snippets, and world knowledge. Alignment, then, is not adding new books to the library. It's hiring a librarian who knows how to find the right book and present it clearly when someone asks a question.
If this hypothesis is correct, then a corollary follows: you should be able to align a pretrained model with a remarkably small number of examples, as long as those examples demonstrate the right style of interaction. This is exactly what LIMA tested.
The dataset: 1,000 examples, each one chosen with care
If alignment is about teaching style rather than knowledge, then the training data doesn't need to be massive — but it does need to be excellent. The authors curated exactly 1,000 prompt-response pairs, totaling roughly 750,000 tokens, from a carefully balanced mix of sources.
750 examples from community forums. 200 from STEM Stack Exchange, 200 from other Stack Exchange topics, and 200 from wikiHow articles, all filtered for quality (minimum length, no first-person language, no cross-references). An additional 150 creative writing examples came from Reddit's r/WritingPrompts. The filtering was strict: answers had to be self-contained, well-written, and stylistically consistent with what a helpful AI assistant would produce.
250 examples manually authored. The paper's authors wrote these themselves, covering diverse tasks — advice, coding help, creative writing, chitchat, and more. They maintained a uniform response style: acknowledge the question, then answer it clearly. 13 of these examples specifically addressed safety-sensitive prompts with careful refusals.
50 examples from Natural Instructions. These covered NLP tasks like summarization and paraphrasing, adding task diversity without sacrificing quality.
The total: 1,000 examples. Compare this to Alpaca's 52,000 or InstructGPT's millions.
Training LIMA: surprisingly standard
The training protocol is deliberately simple. Starting from LLaMA 65B, the authors fine-tuned on their 1,000 examples using standard supervised loss — no reward model, no PPO, no preference pairs. They introduced a special end-of-turn (EOT) to separate user and assistant turns, avoiding conflation with the pretrained model's existing EOS token.
The hyperparameters were standard: optimizer with , , of 0.1, starting at and linearly decaying to , of 32, and training for 15 epochs.
One notable technique was residual with linear increase: dropout started at at the bottom layer and linearly rose to at the top layer. This helped prevent on such a small dataset.
A surprising finding emerged: on held-out data negatively correlated with generation quality. Lower perplexity (normally considered better) actually produced worse outputs. The authors had to manually select checkpoints using a 50-example development set rather than relying on validation loss.
Results: 1,000 examples compete with millions
The authors compared LIMA against five strong baselines across 300 test prompts using human preference evaluation. The results were striking.
LIMA outperformed Alpaca 65B (trained on 52,000 examples) — humans preferred LIMA 53% of the time, with only 26% preferring Alpaca. It also outperformed DaVinci003 (OpenAI's RLHF-trained model) — LIMA won 44% vs 35%.
Against stronger baselines, LIMA remained competitive. Versus Bard, LIMA won or tied 58% of the time. Against Claude, 46%. Even against GPT-4 — widely considered the state of the art — LIMA produced equal or better responses 43% of the time. GPT-4 itself, when used as an annotator, preferred LIMA's own outputs over its own 19% of the time.
On an absolute scale, 50% of LIMA responses were rated excellent and 88% met prompt requirements. Only 12% failed to address the prompt adequately.
Why less is more: quality and diversity beat quantity
The authors ran three ablation experiments on a 7B LLaMA model to understand why a small dataset works so well, testing three factors: diversity, quality, and quantity.
Diversity matters. They compared 2,000 examples from Stack Exchange (diverse prompts across many topics) versus 2,000 examples from wikiHow (all "how-to" prompts). The diverse Stack Exchange data scored 3.83 on a 6-point quality scale, significantly outperforming wikiHow's 3.49. Diverse prompts teach the model to handle a wider range of user requests.
Quality matters. They compared 2,000 quality-filtered Stack Exchange examples versus 2,000 unfiltered ones (no length, style, or self-containedness filters). Filtered data scored 3.83 versus 3.33 for unfiltered — a 0.5-point gap from quality filtering alone.
Quantity surprisingly doesn't matter much. They trained on 2,000, 4,000, 8,000, 16,000, and 32,000 Stack Exchange examples. The result? Performance was essentially flat across this 16-fold increase in data size. Doubling or quadrupling the data provided no measurable gain. This is the strongest evidence for the Superficial Alignment Hypothesis: once the model has seen enough examples of the right style, more examples of the same style add nothing.
Bonus finding: dialogue from almost nothing
LIMA was trained entirely on single-turn interactions — one prompt, one response. Yet when tested in live multi-turn conversations, it showed surprisingly coherent dialogue ability, referencing information from previous turns. In 6 out of 10 test conversations, however, it failed within 3 interactions — a clear sign it was operating out of its training .
The fix was remarkably cheap: adding just 30 hand-crafted dialogue chains (10 written by authors, 20 adapted from Stack Exchange comment threads) to the training set. This tiny addition — 30 examples out of 1,030 total — transformed LIMA's dialogue performance: excellent responses jumped from 45.2% to 76.1%, and failures dropped from 15 per 42 turns to just 1 per 46 turns. The fine-tuned model was significantly better in 7 out of 10 conversations.
This result reinforces the Superficial Alignment Hypothesis from a different angle: the model already knew how to conduct dialogue from pretraining. It just needed a few examples to activate that capability in the right format.
Safety: a few examples go a long way, but not all the way
With only 13 safety-related training examples, LIMA responded safely to 80% of 30 potentially sensitive test prompts. When the malicious intent was explicit — like asking for a celebrity's home address — LIMA refused appropriately. However, when the harmful intent was implicit or disguised, LIMA was more likely to comply unsafely.
This is an important limitation. While the Superficial Alignment Hypothesis suggests that style can be learned from few examples, safety boundaries may require more robust training. The model had learned the format of a refusal ("I cannot help with that") but hadn't internalized the deeper reasoning about when refusals are necessary.
Context: the alignment efficiency timeline
2022
InstructGPT & RLHF at scale
OpenAI demonstrated that RLHF on top of supervised fine-tuning produces highly capable assistants, but at enormous annotation cost — millions of human comparisons.
2023
Alpaca — distillation at scale
Stanford showed that distilling GPT-3.5 outputs into 52,000 examples could train a competitive assistant. Cheaper than RLHF, but still assumed quantity matters.
2023
LIMA — 1,000 examples, no RLHF (this paper)
Proved that careful curation of just 1,000 examples can rival RLHF-trained models, introducing the Superficial Alignment Hypothesis and reshaping the field.
2023
DPO — alignment without RL
Direct Preference Optimization eliminated the reward model entirely, aligning models directly from preference pairs. Continued the trend LIMA started — simpler alignment.
2024
Data-centric alignment becomes mainstream
Teams across the industry shifted focus from annotation scale to data curation, directly influenced by LIMA's findings on quality over quantity.
CitationZhou, Liu, Xu, Iyer, Sun, Mao, Ma, Efrat, Yu, Yu, Zhang, Ghosh, Lewis, Zettlemoyer, Levy. LIMA: Less Is More for Alignment. NeurIPS, 2023.
Terms in this paper
- Superficial Alignment Hypothesisفرضية المواءمة السطحية
- Instruction Tuningالضبط التعليمي
- Fine-Tuningالضبط الدقيق
- Reinforcement Learning from Human Feedbackالتعلم بالتعزيز القائم على التقييم البشري
- Data Qualityجودة البيانات
- Pretrainingالتدريب المسبق
- Supervised Learningالتعلم الـمُوجّه (المصحوب ببيانات مرجعية)
- Promptالنص التوجيهي (الـمُحفّز)
- Alignmentالمحاذاة
- Ablation Studyدراسة الاستئصال
- Generalizationالتعميم
- Few-Shotالنمط القليل العيّنات
- Dropoutالإسقاط العشوائي للعصبونات
- Human Feedbackالتغذية الراجعة البشرية
- Benchmarkالمعيار المرجعي