Language Models2022intermediate11 min read

OPT: Open Pre-Trained Transformer Language Models

OPT: نماذج لغوية مفتوحة مُدرَّبة مسبقاً بمعمارية المُحوِّل

Zhang, S. · Roller, S. · Goyal, N. · Artetxe, M. · Chen, M. · Chen, S. · Dewan, C. · Diab, M. · Li, X. · Lin, X. V. · Mihaylov, T. · Ott, M. · Shleifer, S. · Shuster, K. · Simig, D. · Koura, P. S. · Sridhar, A. · Wang, T. · Zettlemoyer, L. — arXiv

The problem

By 2022, large language models like GPT-3 had demonstrated remarkable and few-shot capabilities, but they remained locked behind corporate APIs. No one outside the originating lab could inspect the weights, reproduce the , or conduct independent safety audits. Training a 175-billion-parameter model costs millions of dollars in compute, so academic labs and smaller organizations were entirely excluded from the frontier of language model research. The field needed an open alternative — not just open weights, but open training code, open data recipes, and transparent documentation of what actually happens when you train at this scale.

The contribution

Open Pre-trained Transformers (OPT): a suite of decoder-only language models ranging from 125M to 175B parameters, released with full weights, training code, and an unprecedented training logbook. OPT-175B matches GPT-3 on standard NLP benchmarks while requiring only 1/7th the to develop. The logbook documents every hardware failure, every learning-rate adjustment, and every divergence across two months of training on 992 A100 GPUs — providing the first public window into the messy reality of training models at this scale.

The impact

OPT broke the norm that frontier language models must stay closed. By releasing weights, code, and training logs, it catalyzed the open-source LLM movement that produced LLaMA, BLOOM, Falcon, and dozens of community models. The logbook set a precedent for training transparency, and the efficiency improvements — 17% better GPU utilization than prior NVIDIA benchmarks — became available for the entire community to build upon.

Imagine a pharmaceutical company discovers a life-saving drug but publishes only the clinical trial results — not the formula, not the manufacturing process, not the lab notebook with every failed batch and adjusted temperature. Other scientists can see that the drug works, but they cannot verify how it works, improve it, or check for hidden side effects.

OPT is like a competing lab that synthesizes the same drug, matches its effectiveness, and then publishes everything: the molecular formula, the step-by-step synthesis protocol, and a brutally honest lab diary documenting every contaminated batch, every equipment malfunction, and every "2 AM restart." Now the entire scientific community can replicate, audit, and improve.

The problem: frontier models behind closed doors

GPT-3 demonstrated in 2020 that a single 175-billion-parameter language model could perform dozens of tasks with zero or few examples — translation, question answering, arithmetic, code generation — simply by conditioning on a text prompt. The results were extraordinary, but the model was accessible only through a paid API. The weights were never released, the training data was described but not shared, and the training process was summarized in a handful of paragraphs.

This created a two-tier research ecosystem. Labs with hundreds of millions of dollars could train and study frontier models. Everyone else could only observe from outside, guessing at architectural details, unable to run ablation studies, safety audits, or fairness evaluations on the actual weights. Reproducibility — a cornerstone of science — was impossible.

OPT was Meta AI's response: replicate GPT-3's capability at the 175B scale, then release everything — model weights across eight sizes (125M, 350M, 1.3B, 2.7B, 6.7B, 13B, 30B, 175B), the full training code, and a detailed logbook of the training process. The goal was not to advance the state of the art but to democratize access to it.

Open in Lab
Click each model size to explore its architecture — layers, attention heads, embedding dimensions, and training compute.
The demo wakes as you arrive…

Architecture: following GPT-3 with deliberate simplicity

OPT deliberately mirrors the GPT-3 architecture to enable fair comparison. Each model is a decoder-only — a stack of layers where every can attend only to previous tokens through a causal attention mask. This autoregressive design means the model generates text one token at a time, each prediction conditioned on all tokens that came before.

The architecture choices closely follow GPT-3 with a few notable differences. OPT uses ReLU activation instead of GeLU, and applies from the Megatron-LM codebase. All models use a sequence length of 2048 tokens and are tokenized with the GPT-2 byte-level tokenizer. Batch sizes range from 0.5M to 4M tokens depending on model size.

The training data is a combination of publicly available datasets: the text corpora used for RoBERTa (BookCorpus, CCNews, CC-Stories), a filtered subset of The Pile (including CommonCrawl, OpenWebText2, Wikipedia, and several others), and PushShift.io Reddit conversations. After deduplication — which reduced The Pile by roughly 66% — the final corpus contains approximately 180 billion tokens (about 800 GB of text). All data is tokenized with the GPT-2 byte-level BPE tokenizer for consistency with prior open language models.

Training at scale: hardware failures, loss spikes, and honest logging

Training OPT-175B required 992 NVIDIA A100 GPUs (80 GB each), organized using Fully Sharded Data Parallel (FSDP) combined with Megatron-LM . This combination achieved up to 147 TFLOP/s per GPU — 17% more efficient than prior NVIDIA-published benchmarks. The optimizer was used with FP32 states and FP16 model weights, along with and dynamic loss scaling to maintain numerical stability.

But the true contribution is not the infrastructure recipe — it is the logbook. Over the course of two months, the team documented every significant event: at least 35 manual restarts, over 100 hosts cycled due to hardware failures, and multiple loss divergences that required lowering the and rolling back to earlier checkpoints.

Open in Lab
Explore the OPT-175B training timeline: hardware failures, learning-rate reductions, and loss spikes across two months of training.
The demo wakes as you arrive…

The logbook reveals a pattern: loss divergences were detected by monitoring the dynamic loss scalar crashing to zero and the L2-norm of final-layer activations spiking. When this happened, the team rolled back to a where the loss scalar was still healthy (≥1.0), lowered the learning rate, and resumed training. This mid-flight adjustment was not a one-time fix — it happened repeatedly throughout the training run.

This level of transparency was unprecedented. Before OPT, the published account of training a 175B model was essentially "we trained it and it worked." The logbook showed that training at this scale is fragile, iterative, and requires constant human intervention — a reality that the polished final papers had been hiding.

Efficiency: 1/7th the carbon footprint

A key claim of the paper is that OPT-175B was developed with approximately 1/7th the carbon footprint of GPT-3. This estimate accounts for the compute used during the successful training run. The authors emphasize that most published carbon footprint numbers exclude the experimental and research phases — the failed runs, the hyperparameter searches, the debugging — which can dwarf the final training cost. OPT's logbook makes at least some of this hidden cost visible.

The improved efficiency comes from combining FSDP with Megatron-LM tensor parallelism, which achieved 17% better GPU utilization than previously published results. These infrastructure improvements were released alongside the model, meaning the entire community benefits from Meta's engineering investment — not just Meta.

Open in Lab
Compare the estimated carbon footprint and compute cost of OPT-175B vs GPT-3.
The demo wakes as you arrive…

Evaluation: comparable to GPT-3, not identical

The authors evaluated OPT across 16 standard NLP benchmarks in zero-shot and few-shot settings, following the GPT-3 evaluation protocol. The results show that OPT-175B achieves performance that is, on average, comparable to GPT-3 Davinci across model sizes. On some tasks OPT performs slightly better; on others, slightly worse.

However, the authors note an important caveat: OPT consistently underperforms GPT-3 on a subset of tasks. They hypothesize that differences in the exact one-shot and few-shot evaluation setup — prompt formatting, example selection, nuances — may account for some of this gap. This highlights a broader problem in LLM evaluation: small differences in how you format a prompt can meaningfully change scores.

Open in Lab
Compare OPT-175B against GPT-3 Davinci across NLP benchmark categories. Toggle between zero-shot and few-shot.
The demo wakes as you arrive…

OPT was also evaluated on dialogue tasks using five benchmark datasets: ConvAI2, Wizard of Wikipedia, Empathetic Dialogues, Blended Skill Talk, and Wizard of the Internet. In a fully unsupervised setting — no on dialogue data — OPT-175B performed competitively against models that were fully supervised on these tasks. This demonstrates the strong emergent conversational ability that comes from large-scale on diverse text including Reddit conversations.

Bias and toxicity: the cost of open data

The authors conducted responsible AI evaluations using three established benchmarks. On hate speech detection (ETHOS dataset), OPT-175B considerably outperformed GPT-3 Davinci in all zero-shot through few-shot settings — a notable strength.

However, on CrowS-Pairs — which measures stereotypical biases across nine categories including gender, race, religion, and socioeconomic status — OPT-175B performed worse than GPT-3 in most categories (lower is better, indicating less ). The overall bias score was 69.5% for OPT versus 67.2% for GPT-3. The authors attribute this to the PushShift.io Reddit corpus, which has a higher incidence of stereotypical and discriminatory text compared to other corpora.

On StereoSet, which measures stereotype awareness across professions, gender, religion, and race, OPT-175B and GPT-3 performed similarly. The authors are transparent about these limitations, noting that opening the model allows the community to study and mitigate these biases — something impossible with a closed model.

Open in Lab
Compare CrowS-Pairs bias scores between OPT-175B and GPT-3 Davinci across categories. Lower is better (less bias).
The demo wakes as you arrive…

Training objective: predicting the next token

Like GPT-3, OPT is trained with the standard causal language modeling objective. The intuition is simple: given a sequence of tokens, predict what comes next. The model reads tokens left-to-right, and at each position it produces a probability distribution over the entire for the next token. Training maximizes the probability of the correct next token across the entire training corpus.

L(θ)=−∑t=1Tlog⁡Pθ(xt∣x1,x2,…,xt−1)\mathcal{L}(\theta) = -\sum_{t=1}^{T} \log P_\theta(x_t \mid x_1, x_2, \ldots, x_{t-1})
Causal language modeling loss — the heart of GPT-style training — At each position tt, the model predicts the probability of token xtx_t given all preceding tokens x1,…,xt−1x_1, \ldots, x_{t-1}. Training minimizes the negative log-likelihood over the entire sequence. This is the same objective used by GPT-2 and GPT-3 — OPT changes nothing about the training goal, only about what gets shared afterward.

Using OPT: from download to generation

Loading and using OPT with Hugging Face Transformerspython

Simplified to show the idea — not the real implementation.

from transformers import AutoModelForCausalLM, AutoTokenizer

# Load OPT-1.3B (fits on a single consumer GPU)
model_name = "facebook/opt-1.3b"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name)

# Generate text autoregressively
prompt = "The future of open-source AI is"
inputs = tokenizer(prompt, return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=50)

print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Impact: the open-source LLM movement

  1. 2020

    GPT-3 — Closed Frontier

    175B parameters, remarkable few-shot learning, but weights never released. Access only through paid API.

  2. 2022

    OPT — Breaking the Barrier

    125M to 175B parameters, open weights, open code, open logbook. Matched GPT-3 at 1/7th the carbon footprint.

  3. 2022

    BLOOM — Community-Built Open LLM

    176B multilingual model trained by BigScience consortium. The first truly community-driven frontier LLM.

  4. 2023

    LLaMA — Efficient Open Models

    Meta's next step: smaller, more compute-efficient models (7B–65B) that match much larger ones. Sparked a massive open ecosystem.

  5. 2023

    LLaMA 2 — Open for Commercial Use

    Released with commercial license. Fine-tuned chat variants included. Open-source LLMs became viable for production.

OPT's deepest legacy is not its benchmark numbers — it is the norm it established. Before OPT, the default for frontier language models was secrecy. After OPT, openness became a competitive advantage and a research expectation. Meta followed OPT with LLaMA, which pushed the efficiency frontier further. The BigScience consortium released BLOOM, a 176B multilingual model trained by over a thousand researchers. Falcon, Mistral, and dozens of others followed.

The logbook tradition also took root. Later open training runs published increasingly detailed accounts of training dynamics, loss spikes, and infrastructure challenges. Training a large model was no longer presented as a solved problem but as an active engineering challenge — and OPT was the paper that forced that honesty.

CitationZhang, Roller, Goyal, Artetxe, Chen, Chen, Dewan, Diab, Li, Lin, Mihaylov, Ott, Shleifer, Shuster, Simig, Koura, Sridhar, Wang, Zettlemoyer. OPT: Open Pre-trained Transformer Language Models. arXiv, 2022.

Terms in this paper