Language Models2023intermediate11 min read
LLaMA: Open and Efficient Foundation Language Models
LLaMA: نماذج أساسية مفتوحة وعالية الكفاءة للغة
Touvron, H. · Lavril, T. · Izacard, G. · Martinet, X. · Lachaux, M.-A. · Lacroix, T. · Rozière, B. · Goyal, N. · Hambro, E. · Azhar, F. · Rodriguez, A. · Joulin, A. · Grave, E. · Lample, G. — arXiv
The problem
By early 2023, the strongest language models (GPT-3 175B, Chinchilla 70B, PaLM 540B) were trained on proprietary data behind closed doors. The prevailing scaling recipe from Chinchilla optimized for cost — pick the model size and count that minimize for a fixed compute budget. But this ignored : a model trained optimally for compute may still be too large to serve affordably. Meanwhile, no competitive open-weight model existed for the research community to study, fine-tune, or deploy.
The contribution
LLaMA: a family of foundation models (7B, 13B, 33B, 65B parameters) trained exclusively on publicly available data. The insight: for a target budget, train a smaller model on far more tokens than Chinchilla-optimal. LLaMA-13B outperforms GPT-3 (175B) on most benchmarks despite being 10× smaller. LLaMA-65B is competitive with Chinchilla-70B and PaLM-540B. The architecture refines the decoder with three proven modifications: pre-normalization, activation, and (). All weights were released to the research community.
The impact
LLaMA cracked open the era of open-weight large language models. Within weeks of its release, the research community built Alpaca, Vicuna, and dozens of fine-tuned variants — proving that a strong base model plus community effort could rival proprietary systems. LLaMA's architecture choices (RMSNorm, SwiGLU, RoPE) became the de facto standard for subsequent open models including LLaMA 2, Mistral, and Qwen. Its scaling philosophy — prioritize inference efficiency by training smaller models longer — reshaped how the field thinks about the compute-performance tradeoff.
Imagine two students preparing for a medical exam. Student A has a huge brain but crams from one expensive textbook the night before. Student B has a normal brain but reads every free textbook in the library over months.
On exam day, Student B scores higher — and can share every textbook with classmates.
LLaMA is Student B. It showed the world that a smaller model trained on more publicly available data can outperform models 10× its size — and then it gave away all its notes.
The insight: train smaller, train longer
In 2022, the Chinchilla paper showed that most large language models were undertrained — you should scale data and model size equally for optimal training loss. But Chinchilla optimized for a fixed training budget, ignoring inference cost.
LLaMA asked a different question: given a target performance level, which model is cheapest to actually run? The answer: a smaller model trained on far more data. A 7B model trained on 1 trillion tokens can match a 10B model trained on 200 billion tokens — and at inference time, the 7B model is 1.4× faster. A 13B model trained on 1 trillion tokens outperforms GPT-3's 175B on most benchmarks.
This reframing — optimize for the model people will deploy, not the one that finishes training fastest — changed the economics of open AI research.
Training data: publicly available, nothing proprietary
A core principle of LLaMA: every byte of training data comes from publicly available sources. This was deliberate — the authors wanted to prove that competitive performance doesn't require proprietary datasets. The training mix spans 1.4 trillion tokens across seven sources, each chosen for a specific strength. CommonCrawl (67%) provides broad web coverage; C4 (15%) adds deduplicated, quality-filtered web text; GitHub (4.5%) contributes code; Wikipedia (4.5%) offers structured factual knowledge in 20 languages; Books (4.5%) bring long-form reasoning; ArXiv (2.5%) adds scientific writing; and StackExchange (2%) provides structured question-answer pairs.
The is SentencePiece with . Notably, LLaMA splits all numbers into individual digits — "2023" becomes four tokens — which improves arithmetic reasoning.
Architecture: a refined Transformer decoder
LLaMA doesn't invent a new architecture — it assembles the best proven improvements on top of the standard Transformer decoder. Think of it as upgrading a car's engine, suspension, and tires instead of designing a new car from scratch. Three specific modifications make the difference:
1. Pre-normalization with RMSNorm. Instead of normalizing after each sub-layer (post-normalization), LLaMA normalizes the input to each sub-layer. And instead of LayerNorm, which centers and rescales, it uses RMSNorm — which only rescales by the root mean square, skipping the centering step. This is simpler, faster, and empirically just as effective at stabilizing training.
2. SwiGLU . The feed-forward network inside each Transformer block replaces with SwiGLU — a gated activation that combines the smooth function with a learned gate. The gate lets the network control which information flows through, and the smooth curve avoids the dead- problem of ReLU. To keep the count constant, the hidden dimension is adjusted to ⅔ × 4d instead of the standard 4d.
3. Rotary Position Embeddings (RoPE). Instead of adding absolute position vectors once at the input (like the original Transformer), LLaMA applies a rotation to and key vectors at every layer. The rotation angle depends on position, so the between two tokens naturally encodes their relative distance. This means the model doesn't need to learn position from scratch — relative position is baked into the geometry of the computation.
RMSNorm: simpler normalization, same stability
LayerNorm does two things: (1) center the vector by subtracting the mean, then (2) rescale by dividing by the standard deviation. RMSNorm asks: do we really need step 1? Empirically, the answer is no — the rescaling alone is enough to keep activations stable. By dropping the mean subtraction, RMSNorm saves compute and simplifies the flow.
SwiGLU: a smarter activation with a learned gate
In the standard Transformer, each feed-forward block applies: (x) = ReLU(xW₁)W₂. ReLU sets all negatives to zero — simple but brutal. Any neuron that lands negative during training is "dead" — it produces zero gradient and never recovers.
SwiGLU replaces this with a gated mechanism: SwiGLU(x) = (Swish(xW₁) ⊙ xV)W₂. The input is projected twice — once through the smooth Swish function (which allows small negative values), and once through a gate matrix V. The elementwise product (⊙) lets the gate decide what information passes through. The smooth Swish curve and the learnable gate together give the network finer control over its nonlinearities.
RoPE: position through rotation
The original Transformer added a fixed sinusoidal vector to each — absolute position encoding. This tells the model "I am at position 5" but not directly "I am 3 positions after you."
RoPE takes a different approach: instead of adding position information, it rotates the query and key vectors by an angle proportional to their position. When two tokens compute their dot product in attention, the angle between their rotated vectors depends only on their relative position — not their absolute positions. It's as if each token stands on a spinning wheel, and the attention score depends on how far apart they are on the wheel, not where the wheel started.
This has a practical advantage: the model generalizes better to sequence lengths it hasn't seen during training, because relative position is an intrinsic geometric property.
The LLaMA family: four sizes, one recipe
LLaMA ships four models that share the same architecture and differ only in scale. All use a 2048-token , the AdamW with , and the implementations (memory-efficient attention and ) to reduce memory and runtime.
The 7B and 13B models train on 1.0 trillion tokens. The 33B and 65B models train on 1.4 trillion tokens — well beyond the Chinchilla-optimal point, deliberately spending extra compute to squeeze more performance out of each parameter. Even after 1 trillion tokens the training loss was still falling, suggesting even more data would have helped.
Efficient training at scale
Training a 65B-parameter model on 1.4 trillion tokens requires serious engineering. LLaMA uses several optimizations to make this feasible:
- Efficient attention. The causal uses an implementation that avoids storing the full n × n attention matrix, reducing memory from O(n²) to O(n). Combined with FlashAttention, this cuts both memory and runtime.
- . Instead of storing all intermediate activations for the backward pass, the network recomputes them during . This trades compute for memory, enabling larger batch sizes.
- Model and . The model is sharded across GPUs, and the data pipeline distributes batches across nodes. The 65B model trained on 2048 A100 GPUs for approximately 21 days.
The total training cost for LLaMA-65B was approximately 1,022,362 GPU-hours on A100-80GB hardware. For comparison, Chinchilla-70B used a similar compute budget but resulted in a model that is harder to serve due to its reliance on proprietary data and infrastructure.
The LLaMA block in code
Simplified to show the idea — not the real implementation.
import numpy as np
def rms_norm(x, gamma, eps=1e-6):
"""RMSNorm: rescale without centering."""
rms = np.sqrt(np.mean(x ** 2, axis=-1, keepdims=True) + eps)
return (x / rms) * gamma
def swish(x):
"""Swish: smooth, self-gated activation."""
return x * (1 / (1 + np.exp(-x))) # x · σ(x)
def swiglu_ffn(x, W1, V, W2):
"""SwiGLU feed-forward: gated activation with smooth curve."""
return (swish(x @ W1) * (x @ V)) @ W2
def rope(x, positions, d, base=10000):
"""Rotary Position Embedding: rotate Q/K by position-dependent angles."""
freqs = 1.0 / (base ** (np.arange(0, d, 2) / d))
angles = np.outer(positions, freqs)
cos_a, sin_a = np.cos(angles), np.sin(angles)
x1, x2 = x[..., ::2], x[..., 1::2]
return np.stack([x1 * cos_a - x2 * sin_a,
x1 * sin_a + x2 * cos_a], axis=-1).reshape(x.shape)
def llama_block(x, attn_weights, ffn_weights, gamma1, gamma2):
"""One LLaMA Transformer block:
Pre-norm → Attention → Residual → Pre-norm → SwiGLU FFN → Residual
"""
# 1. Pre-norm + multi-head attention + residual
h = rms_norm(x, gamma1)
h = multi_head_attention(h, **attn_weights) # with RoPE inside
x = x + h # residual connection
# 2. Pre-norm + SwiGLU FFN + residual
h = rms_norm(x, gamma2)
h = swiglu_ffn(h, **ffn_weights)
x = x + h # residual connection
return x
# Stack N of these blocks (N = 32 for 7B, 80 for 65B),
# and you have LLaMA. That's the entire architecture.Results: smaller can beat bigger
Key results:
- Common Sense Reasoning (BoolQ, PIQA, SIQA, HellaSwag, WinoGrande, ARC, OBQA): LLaMA-13B outperforms GPT-3 175B on all benchmarks except BoolQ. LLaMA-65B outperforms state-of-the-art models of similar size.
- Closed-book Question Answering (NaturalQuestions, TriviaQA): LLaMA-65B achieves state-of-the-art among existing models, despite training on fewer tokens.
- Code Generation (HumanEval, MBPP): LLaMA-13B outperforms most models that weren't specifically trained on code, despite having only 4.5% code in its training mix.
- Mathematical Reasoning (MATH, GSM8k): LLaMA-65B outperforms Minerva-62B on GSM8k despite not being fine-tuned on mathematical data.
- (Massive Multitask Language Understanding): LLaMA-65B lags behind Chinchilla-70B and PaLM-540B — the one area where more parameters still help. After instruction , LLaMA-65B (MMLU 63.4%) closes the gap significantly.
Why it mattered — opening the gate
2020
GPT-3 — 175B parameters, API only
Demonstrated emergent abilities at scale but locked behind a paid API. The weights were never released. Research was limited to prompt engineering.
2022
Chinchilla — scaling laws revisited
Showed previous models were undertrained. The optimal recipe: scale data and model size equally. But optimized for training cost, not inference cost.
2023
LLaMA — open, efficient, competitive
Proved that smaller models trained on more public data beat larger proprietary ones. Released weights to the research community. Launched the open-model movement.
2023
Alpaca & Vicuna — community fine-tuning
Stanford's Alpaca showed that instruction-tuning LLaMA-7B costs just \$600 and produces GPT-3.5-level chat. Vicuna used ShareGPT conversations. The open ecosystem exploded.
2023
LLaMA 2 — full open release
Meta released LLaMA 2 with a permissive license. 7B, 13B, 70B models trained on 2T tokens. Added Grouped Query Attention and 4096-token context. Chat versions included.
2024
Mistral, Qwen, and the open frontier
LLaMA's architecture recipe (RMSNorm + SwiGLU + RoPE) became the standard for open models. Mistral-7B, Qwen, and others pushed the frontier further.
LLaMA's deepest contribution is not any single architectural trick — it's the demonstration that openness and competitiveness reinforce each other. By training on public data and releasing the weights, Meta enabled a Cambrian explosion of research that no single company could have produced alone.
CitationTouvron, Lavril, Izacard, Martinet, Lachaux, Lacroix, Rozière, Goyal, Hambro, Azhar, Rodriguez, Joulin, Grave, Lample. LLaMA: Open and Efficient Foundation Language Models. arXiv, 2023.
Terms in this paper
- RMSNormتسوية الجذر التربيعي للمتوسط
- SwiGLUSwiGLU
- Rotary Position Embeddingsالتضمينات الموضعية الدوّارة
- Scaling Lawقانون التحجيم
- Foundation Modelنموذج أساسي
- Inference Costتكلفة الاستدلال
- Pre-Normالتسوية المسبقة
- Open Weightsالأوزان المفتوحة
- FlashAttentionFlashAttention
- Byte-Pair Encodingترميز أزواج البايت
- Activation Checkpointingنقاط فحص التنشيطات
- Cosine Learning Rate Scheduleجدول معدل التعلم الجيب تمامي
- SwishSwish
- Dead Neuronالخلايا العصبية الميتة
- Efficient Attentionالانتباه عالي الكفاءة