Language Models2022intermediate11 min read
OPT: Open Pre-Trained Transformer Language Models
OPT: نماذج لغوية مفتوحة مُدرَّبة مسبقاً بمعمارية المُحوِّل
Zhang, S. · Roller, S. · Goyal, N. · Artetxe, M. · Chen, M. · Chen, S. · Dewan, C. · Diab, M. · Li, X. · Lin, X. V. · Mihaylov, T. · Ott, M. · Shleifer, S. · Shuster, K. · Simig, D. · Koura, P. S. · Sridhar, A. · Wang, T. · Zettlemoyer, L. — arXiv
The problem
By 2022, large language models like GPT-3 had demonstrated remarkable and few-shot capabilities, but they remained locked behind corporate APIs. No one outside the originating lab could inspect the weights, reproduce the , or conduct independent safety audits. Training a 175-billion-parameter model costs millions of dollars in compute, so academic labs and smaller organizations were entirely excluded from the frontier of language model research. The field needed an open alternative — not just open weights, but open training code, open data recipes, and transparent documentation of what actually happens when you train at this scale.
The contribution
Open Pre-trained Transformers (OPT): a suite of decoder-only language models ranging from 125M to 175B parameters, released with full weights, training code, and an unprecedented training logbook. OPT-175B matches GPT-3 on standard NLP benchmarks while requiring only 1/7th the to develop. The logbook documents every hardware failure, every learning-rate adjustment, and every divergence across two months of training on 992 A100 GPUs — providing the first public window into the messy reality of training models at this scale.
The impact
OPT broke the norm that frontier language models must stay closed. By releasing weights, code, and training logs, it catalyzed the open-source LLM movement that produced LLaMA, BLOOM, Falcon, and dozens of community models. The logbook set a precedent for training transparency, and the efficiency improvements — 17% better GPU utilization than prior NVIDIA benchmarks — became available for the entire community to build upon.
Imagine a pharmaceutical company discovers a life-saving drug but publishes only the clinical trial results — not the formula, not the manufacturing process, not the lab notebook with every failed batch and adjusted temperature. Other scientists can see that the drug works, but they cannot verify how it works, improve it, or check for hidden side effects.
OPT is like a competing lab that synthesizes the same drug, matches its effectiveness, and then publishes everything: the molecular formula, the step-by-step synthesis protocol, and a brutally honest lab diary documenting every contaminated batch, every equipment malfunction, and every "2 AM restart." Now the entire scientific community can replicate, audit, and improve.
The problem: frontier models behind closed doors
GPT-3 demonstrated in 2020 that a single 175-billion-parameter language model could perform dozens of tasks with zero or few examples — translation, question answering, arithmetic, code generation — simply by conditioning on a text prompt. The results were extraordinary, but the model was accessible only through a paid API. The weights were never released, the training data was described but not shared, and the training process was summarized in a handful of paragraphs.
This created a two-tier research ecosystem. Labs with hundreds of millions of dollars could train and study frontier models. Everyone else could only observe from outside, guessing at architectural details, unable to run ablation studies, safety audits, or fairness evaluations on the actual weights. Reproducibility — a cornerstone of science — was impossible.
OPT was Meta AI's response: replicate GPT-3's capability at the 175B scale, then release everything — model weights across eight sizes (125M, 350M, 1.3B, 2.7B, 6.7B, 13B, 30B, 175B), the full training code, and a detailed logbook of the training process. The goal was not to advance the state of the art but to democratize access to it.
Architecture: following GPT-3 with deliberate simplicity
OPT deliberately mirrors the GPT-3 architecture to enable fair comparison. Each model is a decoder-only — a stack of layers where every can attend only to previous tokens through a causal attention mask. This autoregressive design means the model generates text one token at a time, each prediction conditioned on all tokens that came before.
The architecture choices closely follow GPT-3 with a few notable differences. OPT uses ReLU activation instead of GeLU, and applies from the Megatron-LM codebase. All models use a sequence length of 2048 tokens and are tokenized with the GPT-2 byte-level tokenizer. Batch sizes range from 0.5M to 4M tokens depending on model size.
The training data is a combination of publicly available datasets: the text corpora used for RoBERTa (BookCorpus, CCNews, CC-Stories), a filtered subset of The Pile (including CommonCrawl, OpenWebText2, Wikipedia, and several others), and PushShift.io Reddit conversations. After deduplication — which reduced The Pile by roughly 66% — the final corpus contains approximately 180 billion tokens (about 800 GB of text). All data is tokenized with the GPT-2 byte-level BPE tokenizer for consistency with prior open language models.
Training at scale: hardware failures, loss spikes, and honest logging
Training OPT-175B required 992 NVIDIA A100 GPUs (80 GB each), organized using Fully Sharded Data Parallel (FSDP) combined with Megatron-LM . This combination achieved up to 147 TFLOP/s per GPU — 17% more efficient than prior NVIDIA-published benchmarks. The optimizer was used with FP32 states and FP16 model weights, along with and dynamic loss scaling to maintain numerical stability.
But the true contribution is not the infrastructure recipe — it is the logbook. Over the course of two months, the team documented every significant event: at least 35 manual restarts, over 100 hosts cycled due to hardware failures, and multiple loss divergences that required lowering the and rolling back to earlier checkpoints.
The logbook reveals a pattern: loss divergences were detected by monitoring the dynamic loss scalar crashing to zero and the L2-norm of final-layer activations spiking. When this happened, the team rolled back to a where the loss scalar was still healthy (≥1.0), lowered the learning rate, and resumed training. This mid-flight adjustment was not a one-time fix — it happened repeatedly throughout the training run.
This level of transparency was unprecedented. Before OPT, the published account of training a 175B model was essentially "we trained it and it worked." The logbook showed that training at this scale is fragile, iterative, and requires constant human intervention — a reality that the polished final papers had been hiding.
Efficiency: 1/7th the carbon footprint
A key claim of the paper is that OPT-175B was developed with approximately 1/7th the carbon footprint of GPT-3. This estimate accounts for the compute used during the successful training run. The authors emphasize that most published carbon footprint numbers exclude the experimental and research phases — the failed runs, the hyperparameter searches, the debugging — which can dwarf the final training cost. OPT's logbook makes at least some of this hidden cost visible.
The improved efficiency comes from combining FSDP with Megatron-LM tensor parallelism, which achieved 17% better GPU utilization than previously published results. These infrastructure improvements were released alongside the model, meaning the entire community benefits from Meta's engineering investment — not just Meta.
Evaluation: comparable to GPT-3, not identical
The authors evaluated OPT across 16 standard NLP benchmarks in zero-shot and few-shot settings, following the GPT-3 evaluation protocol. The results show that OPT-175B achieves performance that is, on average, comparable to GPT-3 Davinci across model sizes. On some tasks OPT performs slightly better; on others, slightly worse.
However, the authors note an important caveat: OPT consistently underperforms GPT-3 on a subset of tasks. They hypothesize that differences in the exact one-shot and few-shot evaluation setup — prompt formatting, example selection, nuances — may account for some of this gap. This highlights a broader problem in LLM evaluation: small differences in how you format a prompt can meaningfully change scores.
OPT was also evaluated on dialogue tasks using five benchmark datasets: ConvAI2, Wizard of Wikipedia, Empathetic Dialogues, Blended Skill Talk, and Wizard of the Internet. In a fully unsupervised setting — no on dialogue data — OPT-175B performed competitively against models that were fully supervised on these tasks. This demonstrates the strong emergent conversational ability that comes from large-scale on diverse text including Reddit conversations.
Bias and toxicity: the cost of open data
The authors conducted responsible AI evaluations using three established benchmarks. On hate speech detection (ETHOS dataset), OPT-175B considerably outperformed GPT-3 Davinci in all zero-shot through few-shot settings — a notable strength.
However, on CrowS-Pairs — which measures stereotypical biases across nine categories including gender, race, religion, and socioeconomic status — OPT-175B performed worse than GPT-3 in most categories (lower is better, indicating less ). The overall bias score was 69.5% for OPT versus 67.2% for GPT-3. The authors attribute this to the PushShift.io Reddit corpus, which has a higher incidence of stereotypical and discriminatory text compared to other corpora.
On StereoSet, which measures stereotype awareness across professions, gender, religion, and race, OPT-175B and GPT-3 performed similarly. The authors are transparent about these limitations, noting that opening the model allows the community to study and mitigate these biases — something impossible with a closed model.
Training objective: predicting the next token
Like GPT-3, OPT is trained with the standard causal language modeling objective. The intuition is simple: given a sequence of tokens, predict what comes next. The model reads tokens left-to-right, and at each position it produces a probability distribution over the entire for the next token. Training maximizes the probability of the correct next token across the entire training corpus.
Using OPT: from download to generation
Simplified to show the idea — not the real implementation.
from transformers import AutoModelForCausalLM, AutoTokenizer
# Load OPT-1.3B (fits on a single consumer GPU)
model_name = "facebook/opt-1.3b"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name)
# Generate text autoregressively
prompt = "The future of open-source AI is"
inputs = tokenizer(prompt, return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=50)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Impact: the open-source LLM movement
2020
GPT-3 — Closed Frontier
175B parameters, remarkable few-shot learning, but weights never released. Access only through paid API.
2022
OPT — Breaking the Barrier
125M to 175B parameters, open weights, open code, open logbook. Matched GPT-3 at 1/7th the carbon footprint.
2022
BLOOM — Community-Built Open LLM
176B multilingual model trained by BigScience consortium. The first truly community-driven frontier LLM.
2023
LLaMA — Efficient Open Models
Meta's next step: smaller, more compute-efficient models (7B–65B) that match much larger ones. Sparked a massive open ecosystem.
2023
LLaMA 2 — Open for Commercial Use
Released with commercial license. Fine-tuned chat variants included. Open-source LLMs became viable for production.
OPT's deepest legacy is not its benchmark numbers — it is the norm it established. Before OPT, the default for frontier language models was secrecy. After OPT, openness became a competitive advantage and a research expectation. Meta followed OPT with LLaMA, which pushed the efficiency frontier further. The BigScience consortium released BLOOM, a 176B multilingual model trained by over a thousand researchers. Falcon, Mistral, and dozens of others followed.
The logbook tradition also took root. Later open training runs published increasingly detailed accounts of training dynamics, loss spikes, and infrastructure challenges. Training a large model was no longer presented as a solved problem but as an active engineering challenge — and OPT was the paper that forced that honesty.
CitationZhang, Roller, Goyal, Artetxe, Chen, Chen, Dewan, Diab, Li, Lin, Mihaylov, Ott, Shleifer, Shuster, Simig, Koura, Sridhar, Wang, Zettlemoyer. OPT: Open Pre-trained Transformer Language Models. arXiv, 2022.
Terms in this paper
- Decoder-Only Modelنموذج فكّ الترميز فقط
- Autoregressive Modelالنموذج التوليدي التراجعي
- Causal Language Modelالنموذج اللغوي السببي
- Pretrainingالتدريب المسبق
- Fine-Tuningالضبط الدقيق
- Transfer Learningنقل التعلم
- Zero-Shot Learningالتعلّم بدون أمثلة
- Few-Shot Learningالتعلّم بأمثلة قليلة
- Data Parallelismتوازي البيانات
- Tensor Parallelismتوازي مصفوفات الموتّرات
- Distributed Trainingالتدريب الموزَّع
- Dropoutالإسقاط العشوائي للعصبونات
- Weight Decayاضمحلال الأوزان
- Batch Sizeحجم الدفعة الحسابية
- Learning Rateمعدل التعلم