Language Models2024intermediate12 min read
Gemma: Open Models Based on Gemini Research and Technology
Gemma: نماذج مفتوحة الأوزان مبنية على أبحاث وتقنيات Gemini
Gemma Team · Mesnard, T. · Hardin, C. · Dadashi, R. · Sifre, L. · Rivière, M. · Kale, M. S. — arXiv
The problem
By early 2024, frontier language models like GPT-4, Gemini, and Claude were achieving remarkable performance — but remained locked behind APIs and corporate walls. The open-source community had strong models like LLaMA-2 and Mistral, yet these lacked the recipes, data quality, and safety rigor of their closed-source counterparts. Researchers needed open- models that combined state-of-the-art performance at small scale with responsible safety practices and transparent development.
The contribution
Gemma: a family of lightweight, openly available models in two sizes — 2 billion and 7 billion parameters — built from the same research and technology behind Gemini. Gemma provides both pre-trained and instruction-tuned checkpoints, uses a decoder architecture enhanced with Multi-Query , Rotary Position Embeddings (), GeGLU activations, and . Trained on up to 6 trillion tokens, Gemma 7B outperforms similarly sized open models on 11 out of 18 benchmarks, achieving 64.3% on MMLU and 46.4% on GSM8K. The paper also presents comprehensive safety evaluations and memorization analyses.
The impact
Gemma demonstrated that frontier-grade training recipes can produce state-of-the-art open models at small scale. It set a new standard for responsible open model releases by combining performance, safety evaluations, and memorization analysis. Gemma spawned Gemma 2 and CodeGemma, and showed the industry that open-weight models with careful safety practices can democratize access to powerful AI without compromising on responsibility.
Imagine a world-class restaurant that developed a legendary menu over years of research — the finest ingredients, the most refined techniques, the most meticulous food safety. Now imagine the head chef opens a cooking school and hands every student a scaled-down version of that menu: same techniques, same ingredient quality, same safety standards — just portioned for a home kitchen instead of a banquet hall.
That is what Google did with Gemma. They took the architecture, training data quality, and safety rigor behind their flagship Gemini models and distilled them into two compact, openly available models that anyone can download, fine-tune, and deploy — on a GPU, a , or even a laptop CPU.
The big picture: why open-weight models matter
By early 2024, the most powerful language models were closed: accessible only through APIs, with no visibility into weights, training data, or safety processes. Open alternatives like LLaMA-2 (7B, 13B, 70B) and Mistral 7B offered strong performance, but lacked the training infrastructure, data curation, and systematic safety evaluation that characterize frontier labs.
Google's response was Gemma — not just another open model, but a deliberate transfer of frontier methodology to the open ecosystem. Gemma models are built from the same research pipeline as Gemini: same architecture family, similar data processing, same safety evaluation framework. The goal is to give researchers and developers access to models that are simultaneously high-performing, small enough to run locally, and rigorously evaluated for safety.
The name "Gemma" comes from the Latin for "gem" — a small, precious stone. Two sizes were released: a 7 billion parameter model for GPU and TPU deployment, and a 2 billion parameter model designed for CPU and on-device applications. Both come as pre-trained checkpoints and instruction-tuned variants.
Architecture: the transformer decoder, upgraded
Gemma uses a decoder-only transformer — the same fundamental architecture behind GPT and Gemini. But the original 2017 transformer has been upgraded with four key improvements that emerged from the community over the following years. Think of these as engine upgrades on a proven car design: the chassis is the same, but each component has been refined for better efficiency and performance.
1. Multi-Query Attention (MQA) vs (MHA). Standard multi-head attention maintains separate Key and Value projections for every attention head. Multi-Query Attention shares a single Key-Value pair across all heads, dramatically reducing memory during . Gemma 2B uses MQA (1 KV head), while Gemma 7B uses standard MHA (16 KV heads) — a choice made through ablation studies showing each variant performed best at its respective scale.
2. Rotary Position Embeddings (RoPE). Instead of adding absolute position numbers to tokens, RoPE encodes position through rotation in the space. Each position rotates the query and key vectors by an angle proportional to its index. This gives the model a natural sense of distance: nearby tokens have similar rotations, distant tokens diverge. RoPE generalizes better to sequence lengths not seen during training.
3. GeGLU Activations. The standard ReLU activation ("pass positive values, block negatives") is replaced by GeGLU — a gated variant of that multiplies the activation by a learned gate. Think of it as a smart valve: instead of simply blocking or passing, the network learns how much to let through, giving finer control over information flow.
4. RMSNorm. Instead of LayerNorm (which centers and scales), Gemma uses RMSNorm — normalization by root-mean-square alone, without the centering step. This is computationally cheaper and empirically works just as well for transformers, applied to the input of each attention and feedforward sub-layer.
Two sizes, one recipe
Both Gemma models share the same architectural blueprint but differ in scale. The 7B model has 28 transformer layers with a model dimension of 3072 and 16 attention heads. The 2B model has 18 layers with a model dimension of 2048 and 8 attention heads. Both use a head size of 256 and a shared vocabulary of 256,128 tokens via a SentencePiece inherited from Gemini.
The feedforward hidden dimension is intentionally large relative to the model dimension — 49,152 for 7B and 32,768 for 2B — because GeGLU activations split this dimension in half internally (one half for the gate, one half for the value), so the effective width is half the listed number. Both models share their input and output embeddings to reduce parameter count.
Training: infrastructure, data, and process
Gemma models are trained on Google's TPUv5e hardware. The 7B model trains across 16 pods (4,096 TPUv5e chips total), using 16-way and 16-way within each pod, with optimizer state sharding similar to ZeRO-3. The 2B model uses 2 pods (512 chips) with simple 256-way data replication. Cross-pod communication uses the Pathways system.
The training data is primarily English, sourced from web documents, mathematics, and code. The 2B model trains on 2 trillion tokens, while the 7B model trains on 6 trillion tokens. Notably, Gemma is not multimodal — unlike Gemini, it processes text only.
Data filtering is thorough: heuristics and model-based classifiers remove harmful or low-quality content, personal information is filtered out, evaluation sets are excluded from training data to prevent contamination, and the data mixture is staged — meaning the proportion of high-quality data increases toward the end of training, a technique shared with Gemini.
Instruction tuning: from raw model to helpful assistant
A pre-trained language model can complete text, but it does not know how to follow instructions or hold a conversation. Gemma's instruction-tuned (IT) variants go through two stages of alignment.
Stage 1: (SFT). The model is trained on a curated mix of English-only, text-only prompt-response pairs — both human-written and synthetic. The data mixture is selected using LM-based side-by-side evaluations: given held-out prompts, a high-capability judge model compares responses from the candidate and a baseline, expressing a preference. This selects for helpfulness, factuality, creativity, and safety simultaneously.
Stage 2: (RLHF). Human raters compare pairs of model responses. A is trained under the Bradley-Terry framework to predict which response humans prefer. The policy is then optimized using a variant of REINFORCE with a KL-divergence regularization term toward the SFT model — this prevents the model from drifting too far from its supervised baseline while maximizing the reward signal.
Gemma also defines a specific conversation format using control tokens. Every user turn is wrapped with <start_of_turn>user and <end_of_turn>, and every model turn with <start_of_turn>model and <end_of_turn>. This formatting is required at inference time — generating without it is possible but produces worse results because the model is out of distribution.
Simplified to show the idea — not the real implementation.
<start_of_turn>user
What is the capital of France?<end_of_turn>
<start_of_turn>model
The capital of France is Paris.<end_of_turn>
<start_of_turn>user
Tell me more about it.<end_of_turn>
<start_of_turn>model
Paris is the most populous city in France...<end_of_turn>
Benchmarks: how Gemma stacks up
Gemma was evaluated on 18 academic benchmarks spanning reasoning, knowledge, mathematics, coding, and safety. The headline result: Gemma 7B outperforms comparably sized open models on 11 of 18 tasks, and even surpasses the larger LLaMA-2 13B on several benchmarks.
On MMLU (broad knowledge), Gemma 7B scores 64.3% — higher than Mistral 7B (62.5%) and far above LLaMA-2 7B (45.3%). On GSM8K (grade school math), Gemma 7B achieves 46.4%, beating Mistral 7B (35.4%) by over 10 points. On HumanEval (code generation), Gemma 7B scores 32.3% — outperforming even the code-specialized CodeLLaMA 7B (which scores 41.4% on MBPP where Gemma scores 44.4%).
The 2B model holds its own at its smaller scale, scoring 42.3% on MMLU and 22.0% on HumanEval — competitive for a model designed to run on CPUs and mobile devices.
In human preference evaluations against Mistral 7B v0.2 Instruct, Gemma 7B IT achieved a 51.7% win rate on instruction following and a 58% win rate on safety. Even Gemma 2B IT achieved a 56.5% safety win rate against the much larger Mistral 7B — showing that safety is not purely a function of model size.
Safety: beyond benchmarks
Safety is not an afterthought in Gemma — it is embedded at every stage. The paper presents evaluations on 10 safety benchmarks including RealToxicity (measuring toxic text generation), BOLD and CrowS-Pairs (social bias), BBQ (ambiguous bias questions), Winogender and Winobias (gender bias), TruthfulQA (factual accuracy), and Toxigen (adversarial toxicity prompts).
Three layers of safety defense are applied. First, data filtering: training data is scrubbed using heuristic and model-based classifiers to remove harmful, toxic, and personal content. Second, for safety: the SFT and RLHF stages explicitly optimize for safety alongside helpfulness, with data subsets that encourage hedging, refusals, and attribution. Third, memorization evaluation: the paper tests whether the model can regurgitate training data verbatim, checking both exact and approximate memorization, and specifically measures leakage of personal and sensitive information.
Memorization: does the model store training data?
A critical safety concern for any language model is whether it memorizes and can reproduce training data — especially personal information. The Gemma team tested this by sampling 10,000 documents, using the first 50 tokens as a prompt, and checking if the model generates the next 50 tokens verbatim.
The key findings: Gemma memorizes data at rates comparable to PaLM models of similar size. No sensitive personal data was found to be memorized. Some potentially personal data was detected, but the automated detection tools used (Google Cloud Sensitive Data Protection) are known to produce many false positives. Approximate memorization (using a 10% edit distance threshold) showed roughly 50% more data is approximately memorized than exactly — consistent across all dataset categories.
Context: the open-weight model landscape
2023
LLaMA (Meta)
Meta released LLaMA, a family of foundation models from 7B to 65B parameters that demonstrated competitive performance against closed models. Initially restricted to researchers, later versions opened wider access.
2023
LLaMA-2 (Meta)
LLaMA-2 expanded to 70B and included chat-tuned variants, released with a more permissive license for commercial use.
2023
Mistral 7B
Mistral 7B matched LLaMA-2 13B performance with half the parameters, introducing Sliding Window Attention and setting a new efficiency benchmark for open models.
2024
Gemma (Google DeepMind)
Google released Gemma — open-weight 2B and 7B models built from Gemini technology. Outperformed all similarly-sized open models on most benchmarks, with comprehensive safety evaluations and memorization analysis.
2024
Gemma 2 & CodeGemma
Google followed up with Gemma 2 (expanded sizes up to 27B) and CodeGemma (code-specialized variants), continuing the strategy of transferring frontier research into open models.
Gemma's lasting contribution is not just performance numbers — it is the demonstration that responsible open releases are feasible at the frontier. By combining high-quality data curation, modern architectural improvements, systematic safety evaluation, and memorization testing, Google showed that "open" and "responsible" are not contradictory. The recipe book is available; the community decides what to cook with it.
CitationGemma Team, Mesnard, Hardin, Dadashi, Sifre, Rivière, Kale, et al.. Gemma: Open Models Based on Gemini Research and Technology. arXiv, 2024.
Terms in this paper
- Open Weightsالأوزان المفتوحة
- Decoder-Only Modelنموذج فكّ الترميز فقط
- Multi-Head Attentionالانتباه المتعدد المسارات
- Rotary Position Embedding (RoPE)التضمين الموضعي الدوار
- RMSNormتسوية الجذر التربيعي للمتوسط
- GELUوحدة الخطأ الخطية الغاوسية (GELU)
- Instruction Tuningالضبط التعليمي
- Reinforcement Learning from Human Feedbackالتعلم بالتعزيز القائم على التقييم البشري
- Supervised Fine-Tuningالضبط الدقيق الخاضع للإشراف
- Knowledge Distillationتقطير المعرفة
- Pre-trainingالتدريب المسبق
- Fine-Tuningالضبط الدقيق
- Benchmarkالمعيار المرجعي
- AI Safetyسلامة الذكاء الاصطناعي
- Reward Modelنموذج المكافأة