Language Models2023intermediate11 min read

Llama 2: Open Foundation and Fine-Tuned Chat Models

Llama 2: نماذج أساسية مفتوحة الأوزان ونماذج محادثة مُحاذَاة

Touvron, H. · Martin, L. · Stone, K. · Albert, P. · Almahairi, A. · Babaei, Y. · et al. — Meta AI

The problem

By mid-2023, a gap had opened between open-source and closed-source language models. Models like ChatGPT and Claude were fine-tuned with human feedback to be helpful and safe, but their weights were proprietary. Open models like LLaMA 1, Falcon, and MPT offered strong pre-trained checkpoints but lacked the tuning that makes a model safe and useful in real conversations. No open model combined strong , rigorous RLHF alignment, and transparent documentation at scale.

The contribution

Llama 2: a family of pre-trained and fine-tuned LLMs from 7B to 70B parameters, released with . The pre-trained models use 40% more data than LLaMA 1 (2 trillion tokens), doubled context length (4096), and for efficient . The fine-tuned Llama 2-Chat models use a novel alignment pipeline combining , two separate reward models ( and safety), , , and for multi-turn consistency. On human evaluations for helpfulness and safety, Llama 2-Chat outperforms all open-source chat models and approaches closed models like ChatGPT.

The impact

Llama 2 proved that open-weights models can compete with proprietary chat systems in helpfulness and safety. It established the blueprint for transparent RLHF at scale — dual reward models, iterative rejection sampling, Ghost Attention — that influenced every subsequent open model. Its release catalyzed a wave of community fine-tuning, spawning thousands of derivative models and making advanced alignment techniques accessible to researchers worldwide. Llama 2 bridged the gap between open pre-training (LLaMA 1) and the alignment era (Llama 3), demonstrating that openness and safety are not opposing goals.

Imagine a brilliant chef who has tasted every cuisine in the world — she can identify any ingredient, reproduce any recipe, and improvise in any style. That is a pre-trained language model: immense knowledge, zero social awareness.

Now seat that chef in a restaurant. A customer says: "I'm allergic to peanuts." The chef ignores the warning and serves pad Thai — technically excellent, dangerously wrong. Llama 2's contribution is the training program that teaches the chef to listen: a team of tasters scores each dish on whether it is helpful (delicious) and safe (allergen-free). The chef gradually learns to cook meals that are both impressive and responsible — and then publishes her entire recipe book so every restaurant in the world can replicate the result.

The problem: open models can speak, but cannot converse

By early 2023, the landscape of large language models had split into two camps. Closed-source models — ChatGPT, Claude, Bard — had been carefully fine-tuned using (RLHF). They could hold multi-turn conversations, follow instructions safely, and refuse harmful requests. But their weights were proprietary: researchers could not study how alignment worked, and builders could not adapt them to new domains.

On the other side, open models like LLaMA 1, Falcon, and MPT released strong pre-trained checkpoints. These models could generate fluent text and score well on academic benchmarks. But without alignment tuning, they also generated toxic content, followed harmful instructions, and forgot system messages mid-conversation. They were engines without steering wheels.

The gap was not about raw capability — LLaMA 1-65B matched Chinchilla-70B on many benchmarks. The gap was about alignment: closed models had it, open models did not. Llama 2 set out to close this gap and publish the complete recipe.

Open in Lab
Open models matched closed models on knowledge benchmarks, but lagged far behind on alignment — helpfulness and safety. Llama 2-Chat closes this gap.
The demo wakes as you arrive…

Pre-training: a stronger foundation

Llama 2 builds on the same architecture as LLaMA 1, using pre-normalization with , the , and rotary position embeddings (). But the pre-training recipe was substantially upgraded in three ways:

1. More data. The training corpus grew by 40% to 2 trillion tokens from publicly available sources, with more robust data cleaning and an updated data mix. No data from Meta's own products or services was used.

2. Longer context. The doubled from 2048 to 4096 tokens, allowing the model to process and reason over longer passages.

3. Grouped Query Attention. The larger models (34B and 70B) adopted grouped-query attention (GQA), which shares key and value projections across multiple attention heads. GQA preserves quality close to standard while dramatically reducing the memory footprint of the KV cache during inference — a crucial optimization for serving large models at scale.

Open in Lab
Explore the Llama 2 model family — click any model size to see its architecture details and how it compares to LLaMA 1.
The demo wakes as you arrive…
GQA: Q∈Rn×d,  K,V∈Rn×(d/g)⇒each group of g heads shares one K,V pair\text{GQA: } Q \in \mathbb{R}^{n \times d},\; K, V \in \mathbb{R}^{n \times (d / g)} \quad\Rightarrow\quad \text{each group of } g \text{ heads shares one } K, V \text{ pair}
Grouped Query Attention — sharing KV across head groups — In standard multi-head attention each head has its own Key and Value projections. GQA groups g heads together and gives each group a single shared K and V, reducing KV cache memory by a factor of g while maintaining nearly the same quality.

The alignment pipeline: from raw model to safe chat assistant

The journey from Llama 2 (pre-trained) to Llama 2-Chat (aligned) has three stages, inspired by InstructGPT but extended with several innovations. Think of it as a three-step apprenticeship: first the model copies examples, then it learns from grades, and finally it practices until the grades improve.

Stage 1 — Supervised Fine-Tuning (SFT). The pre-trained model is trained on high-quality human demonstrations of helpful and safe conversations. The key insight: quality matters more than quantity. Meta found that a small set of 27,540 carefully curated vendor annotations outperformed millions of lower-quality examples. During SFT, the loss is computed only on the assistant's response tokens, not the user's prompt — this teaches the model what to say, not how to repeat the question.

Stage 2 — Reward Modeling. Two separate reward models are trained: one for helpfulness and one for safety. This dual-model approach avoids the tension where a single might sacrifice helpfulness to be safe, or vice versa. The reward models use binary comparison data: given a prompt and two responses, human annotators indicate which is better. The training objective uses a modified ranking loss with a margin term that rewards more decisive distinctions.

Stage 3 — RLHF with Rejection Sampling + PPO. Using the reward models as judges, the chat model iteratively improves through five rounds (RLHF-V1 through RLHF-V5). In each round, the 70B model generates K candidate responses per prompt, the best one is selected by the reward model (rejection sampling), and the model is fine-tuned on these top responses. Starting from RLHF-V5, PPO is applied on top for further optimization. Smaller models (7B, 13B) are also fine-tuned on rejection-sampled data from the 70B model — a form of distillation.

Open in Lab
Follow the three-stage alignment pipeline — from raw pre-trained model to safe chat assistant. Click each stage to see how it transforms the model.
The demo wakes as you arrive…
Lranking=−log⁡ σ ⁣(rθ(x,yc)−rθ(x,yr)−m(r))\mathcal{L}_{\text{ranking}} = -\log\,\sigma\!\bigl(r_\theta(x, y_c) - r_\theta(x, y_r) - m(r)\bigr)
Reward model training loss with margin — Given prompt x, chosen response y_c, and rejected response y_r, the reward model r_θ is trained so that the chosen response scores higher. The margin m(r) is proportional to the annotator's confidence: a clearly better response gets a larger margin, teaching the model to calibrate its scores, not just rank correctly.

Ghost Attention: remembering instructions across turns

In multi-turn conversations, a user might set a system instruction at the beginning — like "Always respond in French" or "Act as a travel guide." With standard RLHF training, the model tends to forget these instructions after a few turns, drifting back to its default behavior. This is the multi-turn consistency problem.

Llama 2 introduces Ghost Attention (GAtt), a technique inspired by . The idea is elegant: during fine-tuning, the system instruction is synthetically concatenated to every user message in the conversation, not just the first one. But in the training loss, only the last turn's response is scored — so the model learns to follow the instruction even though, at inference time, it only appears once at the start.

The "ghost" metaphor is apt: the instruction is invisibly present at every turn during training, teaching the model to maintain it in the actual conversation where it only appears once. After GAtt training, Llama 2-Chat maintains system instructions for over 20 turns — far beyond what was possible without it.

Open in Lab
Watch how Ghost Attention preserves system instructions across conversation turns — compare with and without GAtt.
The demo wakes as you arrive…

Safety: a first-class design goal

Unlike previous open models that treated safety as an afterthought, Llama 2 integrates safety at every stage. The approach operates on three fronts:

Safety data annotation. Annotators wrote prompts specifically designed to elicit unsafe responses — adversarial examples that probe the model's boundaries. Responses were rated on a five-point safety scale, and the safety reward model was trained on these labeled comparisons.

Safety RLHF. The safety reward model runs in parallel with the helpfulness reward model during RLHF. A key finding: early iterations showed a tension between helpfulness and safety. Being too cautious made the model refuse legitimate requests (over-refusal). The team discovered that increasing the safety training data proportion and using context distillation — generating safer responses with a safety preprompt, then training without the preprompt — significantly reduced this trade-off.

. Over 350 people with diverse expertise (cybersecurity, election fraud, social media manipulation, legal, government policy, and more) adversarially probed the model. Their findings directly informed the safety training data and the definition of risk categories. This is one of the most comprehensive red teaming efforts documented for an open model.

Open in Lab
Explore the tension between helpfulness and safety across RLHF iterations. Watch how later iterations resolve the trade-off.
The demo wakes as you arrive…

The idea in code

Rejection sampling — the core RLHF selection looppython

Simplified to show the idea — not the real implementation.

import numpy as np

def rejection_sampling(model, reward_model, prompt, K=64):
    """Generate K candidate responses; keep the highest-reward one."""
    candidates = []
    for _ in range(K):
        response = model.generate(prompt, temperature=0.7)
        score = reward_model.score(prompt, response)
        candidates.append((response, score))

    # Select the response with the highest reward
    best_response, best_score = max(candidates, key=lambda x: x[1])
    return best_response, best_score

def dual_reward_selection(prompt, response, help_rm, safety_rm):
    """Score with BOTH reward models — helpfulness and safety."""
    help_score  = help_rm.score(prompt, response)
    safety_score = safety_rm.score(prompt, response)

    # Safety acts as a gate: if safety score is below threshold,
    # the response is rejected regardless of helpfulness
    if safety_score < SAFETY_THRESHOLD:
        return -float('inf')  # reject unsafe responses
    return help_score         # rank safe responses by helpfulness

# Key insight: the 70B model generates candidates, and smaller models
# (7B, 13B) are fine-tuned on the 70B's best responses — distillation.

Results: closing the open-closed gap

Llama 2-Chat demonstrated strong results on both automated and human evaluations:

  • Helpfulness (human evaluation): Llama 2-Chat 70B achieved a 36% win rate against ChatGPT (GPT-3.5) in head-to-head comparisons — remarkable for an open model. It outperformed all existing open-source chat models including Falcon, MPT, and Vicuna.

  • Safety (human evaluation): Llama 2-Chat showed a single-turn safety violation rate of only 4% on adversarial prompts, compared to 7% for ChatGPT. On multi-turn adversarial conversations, the violation rate was 0.2%.

  • Pre-trained benchmarks: The pre-trained Llama 2 70B outperformed all open-source models available at the time on academic benchmarks including MMLU, HellaSwag, ARC, and WinoGrande. On coding (HumanEval) and math (GSM8K), it was competitive but not leading — a known gap addressed in later models.

  • Open vs closed: While Llama 2-Chat did not match GPT-4, it demonstrated that the gap between open and closed models could be dramatically narrowed through systematic alignment — and that the alignment recipe could be published transparently.

Open in Lab
Compare Llama 2-Chat against open and closed models across safety, helpfulness, and academic benchmarks.
The demo wakes as you arrive…

The Llama lineage

  1. 2023

    LLaMA 1 — Open pre-trained models

    Four decoder-only models (7B–65B) trained on public data. LLaMA-13B outperformed GPT-3 175B on most benchmarks, proving that smaller well-trained models can match giants. Released for research use only.

  2. 2023

    Llama 2 — Open foundation + aligned chat

    Three sizes (7B, 13B, 70B) with 40% more data, doubled context, GQA, and a full RLHF alignment pipeline. First open model to rival closed chat systems on safety and helpfulness. Released with a commercial license.

  3. 2023

    Community explosion

    Thousands of fine-tuned variants emerged within weeks — code specialists, multilingual adaptations, domain-specific models. Llama 2 became the de facto base for open alignment research.

  4. 2024

    Llama 3 — Scaling the recipe

    Expanded to 8B, 70B, and 405B parameters with 128K context, a 128K-token tokenizer, and multimodal capabilities. Built directly on Llama 2's alignment framework with more data and compute.

CitationTouvron, Martin, Stone, Albert, Almahairi, et al.. Llama 2: Open Foundation and Fine-Tuned Chat Models. Meta AI, 2023.

Terms in this paper