Language Models2024advanced13 min read
The Llama 3 Herd of Models
قطيع نماذج Llama 3
Grattafiori, A. · Dubey, A. · Jauhri, A. · Pandey, A. · Kadian, A. · et al. — arXiv
The problem
By mid-2024, the most capable language models — GPT-4, Gemini, Claude — were all closed-source. Researchers and developers could use them only through APIs with limited control, no visibility into data or methods, and no ability to fine-tune or deploy locally. Open-weight alternatives like Llama 2 existed but lagged significantly behind frontier models in reasoning, coding, multilingual capability, and long-context understanding. There was no open model that could genuinely compete with the best closed-source systems across the board.
The contribution
Llama 3: a family of dense language models at 8B, 70B, and 405B parameters, all open-weight. The 405B flagship was trained on 15.6 trillion tokens using a 128K , , , , and . Post-training uses multiple rounds of , , and . The result is an open model that matches GPT-4 level performance across benchmarks including MMLU (88.6%), HumanEval (89.0%), and multilingual MGSM (91.6%). Smaller models were trained beyond compute-optimal to maximize quality, and Llama Guard 3 provides input/output filtering.
The impact
Llama 3 proved that open-weight models can match frontier closed-source performance, reshaping the industry's assumptions about the value of openness. The 405B model became the most capable openly available LLM at its release, enabling researchers to study, fine-tune, and deploy frontier-quality models locally. Its training recipe — scaling data quality, training smaller models longer, and iterative post-training — became a widely adopted blueprint. Llama 3 accelerated the open-source ecosystem, influenced DeepSeek-V3 and other successors, and demonstrated that compute-optimal scaling with high-quality data matters more than architectural novelty.
Imagine a fleet of cargo ships: a speedboat (8B), a freighter (70B), and a supertanker (405B). Building the supertanker is the hard part — you need the best steel (data quality), the biggest shipyard (16,000 GPUs), and the most fuel (15 trillion tokens). But once the supertanker is sailing, you can teach the smaller ships its routes by watching and copying — a technique called .
Meta didn't just build the ships. They published the blueprints, opened the shipyard, and let anyone sail. That's what makes Llama 3 different: frontier capability, fully open.
Philosophy: three levers, not one
Most frontier labs chase architectural novelty — , state space models, new patterns. Meta took the opposite bet: use a standard dense Transformer and pour all effort into three levers:
-
Data quality — 15 trillion tokens, aggressively filtered and deduplicated. The filtering pipeline alone uses heuristic classifiers, fastText language identifiers, and model-based quality classifiers trained on Llama 2 itself.
-
Scale — 405 billion parameters trained on 3.8 × 10²⁵ FLOPs across 16,000 H100 GPUs. This is roughly compute-optimal for their training budget according to their own .
-
Managing complexity — by keeping the architecture simple, the team could iterate faster on data recipes, training stability, and post-training alignment without debugging novel components.
The philosophy is: don't complicate what works. Scale the recipe, not the architecture.
Data: quality over quantity (but also quantity)
Llama 3's training corpus is 15 trillion tokens — over 8× larger than Llama 2's 1.8 trillion. But the raw size is only half the story. The team built an elaborate multi-stage filtering pipeline:
Stage 1: Deduplication. URL-level, document-level, and line-level deduplication removes near-duplicate crawl data. MinHash is used for fuzzy document deduplication.
Stage 2: Heuristic filtering. Rules eliminate documents with too many dirty words, excessive special characters, or abnormally low counts.
Stage 3: Model-based quality scoring. A fastText classifier trained on curated high-quality pages (e.g. Wikipedia references) assigns quality scores. Only documents above a threshold pass through.
Stage 4: Domain-specific pipelines. Separate pipelines for code (syntax validation, decontamination), math (LaTeX extraction, problem filtering), and multilingual data (per-language quality classifiers).
The data mix itself is carefully tuned: roughly 50% English web data, 25% code, and 25% a mixture of multilingual, math, and curated knowledge sources. The mix shifts during training — in the final phase, they upsample the highest-quality sources.
Architecture: standard Transformer, smart choices
Llama 3 is a dense Transformer. The architecture deliberately avoids trendy additions like mixture of experts — simplicity is the point. But several carefully chosen modifications make it efficient at scale:
Grouped Query Attention (GQA): Instead of giving every query head its own key-value pair, 8 key-value heads are shared across all 128 query heads (in the 405B model). Think of it like a conference room: instead of 128 private secretaries taking notes, 8 shared note-takers serve everyone. This dramatically reduces the KV cache size during without hurting quality.
(RoPE): Positions are encoded by rotating the query and key vectors in 2D planes. The base frequency is set to 500,000, enabling stable context lengths up to 128K tokens. During , the context starts at 8K and is extended to 128K in a dedicated long-context training stage.
SwiGLU activation: The feed-forward sublayers use SwiGLU — a gated variant that multiplies two parallel projections, one through a Swish activation. This consistently outperforms ReLU and GELU in practice.
RMSNorm pre-normalization: Each sublayer is preceded by RMS normalization (not LayerNorm). This is simpler, computationally cheaper, and has become the standard for large-scale training.
Scaling: 16,000 GPUs and the art of not crashing
Training a 405B model at this scale is an engineering feat. The compute infrastructure used 16,384 NVIDIA H100 GPUs, each delivering roughly 400 TFLOPS of compute. Three parallelism strategies were combined:
(TP): Each layer's weight matrices are split across 8 GPUs within one node. Think of this as eight people each holding one column of a spreadsheet — they do their part and pass results around.
(PP): The 126 layers are divided into stages across 16 pipeline stages, so different GPUs process different layers of the model simultaneously — like stations on an assembly line.
(DP/FSDP): The same batch is split across replicas. With Fully Sharded Data Parallelism, optimizer states and gradients are sharded too, dramatically reducing per-GPU memory.
The total training compute was approximately 3.8 × 10²⁵ FLOPs. To put this in perspective: if a single H100 ran continuously, it would take over 30 million years to complete the same training run. The team achieved 38–43% Model FLOPs Utilization (MFU), meaning they extracted nearly half of the theoretical GPU throughput — a strong result at this scale. During training they experienced 466 job interruptions, of which 47 were planned, demonstrating the need for robust checkpointing and automatic recovery.
Post-training: from raw model to helpful assistant
Pre-training produces a next-token predictor. Post-training turns it into an assistant that follows instructions, reasons carefully, and refuses harmful requests. Llama 3's post-training is an iterative loop with three stages:
1. Rejection Sampling: Generate multiple candidate responses using the current model, score them with a , and keep only the best. This produces high-quality synthetic training data that pushes quality higher than human-written examples alone.
2. Supervised Fine-Tuning (SFT): Fine-tune on a curated mix of human-written and synthetic demonstrations. The data covers dialogue, coding, math reasoning, , multilingual interaction, and long-context tasks.
3. Direct Preference Optimization (DPO): Instead of training a separate reward model and running PPO (as in Llama 2), DPO directly optimizes the policy by contrasting chosen vs rejected response pairs. This is simpler, more stable, and the team found it performed as well as or better than PPO.
This loop runs multiple rounds — each round uses the improved model to generate better rejection samples for the next round, creating a self-improving cycle. The key insight: by the later rounds, the model generates training data that exceeds what human annotators can produce on complex tasks like coding and math.
Tokenizer: 128K vocabulary for the world
Llama 3 uses a with a of 128,256 tokens — 4× larger than Llama 2's 32K vocabulary. This larger vocabulary achieves a compression ratio of 3.94 characters per token (vs ~4.5 for Llama 2), meaning the same text requires fewer tokens and the model can process more text within its context window.
Critically, 28,000 tokens were specifically added for non-English languages. This dramatically improved compression for languages like Arabic, Hindi, Thai, and others that were poorly served by the smaller vocabulary. Better compression means the model can process longer passages in these languages and training efficiency improves.
Safety: Llama Guard and the safety pipeline
Safety is woven into every stage. The post-training mix includes safety-focused data: adversarial prompts with safe refusals, borderline examples where the model should engage rather than over-refuse, and red-teaming data to handle edge cases.
Llama Guard 3 is a separately trained classifier that acts as an input/output filter. It evaluates both user prompts and model responses against a taxonomy of harmful categories (violence, sexual content, criminal advice, etc.) and can block or flag unsafe content. This modular design means developers can adjust safety thresholds without retraining the main model.
The team reports achieving a much better balance between helpfulness and harmlessness compared to Llama 2, which was criticized for excessive refusals. They specifically optimized to reduce false refusals while maintaining safety.
Results: open meets frontier
Llama 3 405B matches or closely approaches GPT-4 across major benchmarks:
- MMLU (general knowledge): 88.6% — competitive with GPT-4's ~86.4%
- HumanEval (): 89.0% — matching frontier code models
- GSM8K (math reasoning): 96.8% — near-ceiling performance
- MGSM (multilingual math): 91.6% — strong across 8 languages
- MATH (competition math): 73.8% — substantial improvement over Llama 2
The smaller models are equally impressive relative to their class. The 8B model outperforms Mistral 7B and Gemma 7B across most tasks. The 70B model competes with much larger models including Mixtral 8x22B. These smaller models were intentionally over-trained — trained far beyond compute-optimal on more tokens — sacrificing training efficiency for better final quality.
Multimodal experiments: images, video, and speech
The paper also presents experimental results for multimodal extensions of Llama 3, using a compositional approach — specialized encoders are connected to the language model via adapter layers, without modifying the core Transformer.
Vision: A -H/14 image (630M parameters) produces patch embeddings that are injected into the text model through cross-attention layers every four Transformer blocks. This adds roughly 100B parameters to the 405B model.
Video: Frames are encoded by the same ViT, then a temporal aggregator (Perceiver Resampler) merges every 32 frames into summary representations.
Speech: A 1B-parameter Conformer encoder processes audio spectrograms, and its outputs are integrated through the same adapter architecture.
These multimodal versions performed competitively with state-of-the-art models but were not publicly released as they were still under development at the time of publication.
The idea in code
Simplified to show the idea — not the real implementation.
import torch
import torch.nn as nn
import torch.nn.functional as F
class RMSNorm(nn.Module):
"""RMSNorm — simpler and faster than LayerNorm."""
def __init__(self, dim, eps=1e-6):
super().__init__()
self.weight = nn.Parameter(torch.ones(dim))
self.eps = eps
def forward(self, x):
rms = torch.sqrt(x.pow(2).mean(-1, keepdim=True) + self.eps)
return x / rms * self.weight
class SwiGLU(nn.Module):
"""SwiGLU: gate = swish(W1·x), output = gate ⊙ (W2·x)."""
def __init__(self, dim, ffn_dim):
super().__init__()
self.w1 = nn.Linear(dim, ffn_dim, bias=False) # gate projection
self.w2 = nn.Linear(dim, ffn_dim, bias=False) # value projection
self.w3 = nn.Linear(ffn_dim, dim, bias=False) # down projection
def forward(self, x):
return self.w3(F.silu(self.w1(x)) * self.w2(x))
class GroupedQueryAttention(nn.Module):
"""GQA: 128 query heads share 8 KV heads → 16x smaller KV cache."""
def __init__(self, dim, n_heads=128, n_kv_heads=8):
super().__init__()
self.n_heads = n_heads
self.n_kv_heads = n_kv_heads
self.head_dim = dim // n_heads
self.repeats = n_heads // n_kv_heads # how many Q heads per KV head
self.wq = nn.Linear(dim, n_heads * self.head_dim, bias=False)
self.wk = nn.Linear(dim, n_kv_heads * self.head_dim, bias=False)
self.wv = nn.Linear(dim, n_kv_heads * self.head_dim, bias=False)
self.wo = nn.Linear(n_heads * self.head_dim, dim, bias=False)
def forward(self, x):
B, L, _ = x.shape
q = self.wq(x).view(B, L, self.n_heads, self.head_dim)
k = self.wk(x).view(B, L, self.n_kv_heads, self.head_dim)
v = self.wv(x).view(B, L, self.n_kv_heads, self.head_dim)
# Repeat KV heads to match query head count
k = k.repeat_interleave(self.repeats, dim=2)
v = v.repeat_interleave(self.repeats, dim=2)
# Apply RoPE here (omitted for brevity)
# Standard scaled dot-product attention
q, k, v = [t.transpose(1, 2) for t in (q, k, v)]
attn = F.scaled_dot_product_attention(q, k, v, is_causal=True)
return self.wo(attn.transpose(1, 2).reshape(B, L, -1))
# Full Llama 3 block: Pre-Norm → GQA → residual → Pre-Norm → SwiGLU → residual
# 405B: 126 of these blocks, dim=16384, 128 Q heads, 8 KV heads, ffn=53248Scaling laws: predicting the future from small experiments
One of the most practically impactful contributions is Llama 3's approach to scaling laws. Before committing billions of GPU-hours, the team trained hundreds of small models (40M to 16B parameters) and used them to predict the performance of the 405B model.
Their key finding: the standard Chinchilla scaling law (which predicts compute-optimal model size) works well for determining the 405B flagship size, but smaller models benefit from being trained far beyond compute-optimal. The 8B model, for instance, was trained on tokens far exceeding what Chinchilla would recommend — because at inference time, compute cost depends on model size, not training tokens. A smaller model trained longer is cheaper to deploy than a compute-optimal larger model.
They also developed a two-stage scaling law for predicting performance: first predict the from compute, then predict the benchmark score from the log-likelihood. This pipeline accurately predicted the 405B model's benchmark scores from small-scale experiments.
The Llama lineage
2023
Llama 1
The first open-weight LLM from Meta. Models from 7B to 65B, trained on 1.4T tokens. Proved that open models with enough training data could rival closed-source systems.
2023
Llama 2
Extended training to 2T tokens. Introduced GQA for inference efficiency and RLHF with rejection sampling + PPO. First open model with a dedicated chat variant.
2024
Llama 3 / 3.1
The flagship 405B model trained on 15T tokens. 128K context, DPO replaces PPO, 128K vocabulary, Llama Guard 3 for safety. First open model to match GPT-4 quality.
2024
DeepSeek-V3
Built on lessons from the Llama recipe, DeepSeek-V3 introduced a mixture-of-experts architecture with innovative training efficiency, becoming the most capable open model shortly after Llama 3.
CitationGrattafiori, Dubey, Jauhri, Pandey, Kadian, et al. (530+ authors). The Llama 3 Herd of Models. arXiv, 2024.
Terms in this paper
- Grouped Query Attentionانتباه الاستعلام المُجمَّع
- Rotary Position Embedding (RoPE)التضمين الموضعي الدوار
- RMSNormتسوية الجذر التربيعي للمتوسط
- SwiGLUSwiGLU
- Scaling Lawقانون التحجيم
- Direct Preference Optimizationالتحسين المباشر للتفضيلات
- Rejection Samplingالترشيح بالرفض
- Supervised Fine-Tuningالضبط الدقيق الخاضع للإشراف
- Context Windowنافذة السياق
- Byte Pair Encoding (BPE)ترميز زوج البايت
- Knowledge Distillationتقطير المعرفة
- Safetyأمان الأنظمة الذكية
- Tool Useاستخدام الأدوات
- Foundation Modelنموذج أساسي
- Open Weightsالأوزان المفتوحة