Language Models2023intermediate10 min read

GPT-4 Technical Report

التقرير التقني لـ GPT-4

OpenAI — arXiv

The problem

By 2022, large language models like GPT-3.5 had shown impressive abilities but remained unpredictable at scale: you couldn't know how well a trillion-dollar training run would perform until it finished. Models also struggled with factual accuracy, safety, and could not process images alongside text. The gap between what LLMs could do and what professionals needed was still wide.

The contribution

GPT-4: a large-scale model that accepts both text and image inputs and produces text outputs. Three core contributions: (1) — infrastructure and optimization methods that let engineers accurately predict GPT-4's performance from models trained with 1/1,000th the compute using power-law scaling laws. (2) Human-level exam performance — GPT-4 passes the bar exam in the top 10% and achieves 86.4% on , beating prior state-of-the-art across 57 subjects. (3) A model-assisted using (RBRMs) and that reduced disallowed content by 82% over GPT-3.5.

The impact

GPT-4 demonstrated that large language models could achieve professional-grade performance across law, medicine, and STEM — transforming how society views AI capability. Its predictable scaling methodology showed that expensive training runs need not be gambles. The safety pipeline established a template adopted industry-wide. GPT-4 powers ChatGPT, Microsoft Copilot, and hundreds of enterprise applications, marking the moment LLMs transitioned from research curiosity to critical infrastructure.

Training a is like planning a skyscraper: the taller you build, the more unknowns you face — will the foundation hold? Will the cost spiral? Traditionally, you build the whole thing and pray.

GPT-4's breakthrough is like discovering that a scale model in a wind tunnel can accurately predict how the real skyscraper will sway. By testing tiny models first, the engineers knew what to expect from the full giant — before spending the compute to build it.

The skyscraper also has a security system: a model-assisted safety pipeline that watches for harmful outputs the way a smart alarm watches for intruders — using the building's own intelligence to protect its occupants.

What is GPT-4?

GPT-4 is a large-scale built on the Transformer architecture. Like its predecessors, it is pre-trained to predict the next in a document, then fine-tuned using (RLHF) to follow instructions and behave safely. But GPT-4 introduces two leaps: it can accept images alongside text as input, and its builders developed methods to predict its performance before training finished.

The report deliberately withholds architecture details — model size, hardware, training compute, and dataset construction — citing competitive and safety reasons. What it does reveal is how GPT-4 performs, how its performance was predicted, and how its safety was improved.

Predictable scaling: knowing the answer before the exam

The central engineering insight of GPT-4 is that performance at massive scale can be predicted from small-scale experiments. Training GPT-4 required enormous compute — far too expensive for trial and error. Instead, the team trained much smaller models using the same methodology and fitted a to predict the final .

The intuition: imagine plotting the fuel efficiency of toy cars at different sizes. If the data follows a smooth curve, you can predict what a full-sized car will achieve without building one. The key formula is a with an term that captures the floor below which no amount of compute can push.

The prediction matched GPT-4's actual final loss with remarkable accuracy — this was done before the full training run completed, using models trained with at most 1/10,000th the compute.

L(C)=a⋅Cb+cL(C) = a \cdot C^{b} + c
Scaling law with irreducible loss — This scaling law describes how model performance improves as more computation is invested in training. The loss typically decreases in a highly predictable pattern: large increases in compute continue to produce improvements, but each additional unit of compute yields smaller gains than the one before. The curve eventually approaches a performance floor that cannot be surpassed merely by adding more compute. Because this relationship is remarkably consistent across model sizes, researchers can estimate the performance of much larger models by observing trends from smaller training runs.
Open in Lab
Drag the compute slider to see how loss decreases predictably. The red star shows GPT-4's actual performance matching the prediction.
The demo wakes as you arrive…

Human-level exam performance

GPT-4 was tested on exams designed for humans — bar exams, SATs, AP tests, medical knowledge assessments, and coding challenges — with no exam-specific training. The results were striking:

  • Uniform Bar Exam: scored ~298/400, placing in the top 10% of test takers (GPT-3.5 scored in the bottom 10%).
  • SAT Math: 700/800 (~89th percentile).
  • MMLU (57 academic subjects): 86.4% accuracy, beating all prior models and most benchmark-specialized systems.
  • GRE Verbal: 169/170 (~99th percentile).

These capabilities stem primarily from pre-training, not from RLHF — the base model and the RLHF model performed nearly identically on multiple-choice exams (73.7% vs 74.0% averaged across all exams).

Open in Lab
Click an exam category to compare GPT-4 vs GPT-3.5 performance.
The demo wakes as you arrive…

Beyond English: multilingual mastery

To test GPT-4 in languages beyond English, the team translated the entire MMLU benchmark into 26 languages using Azure Translate. GPT-4 outperformed the English-language scores of both GPT-3.5 and prior models (Chinchilla and PaLM) in 24 of 26 languages — including low-resource languages such as Latvian, Welsh, and Swahili.

This is remarkable because GPT-4 was not specifically trained on multilingual exam data. The model's understanding of concepts transfers across language boundaries. For Arabic, GPT-4 achieved approximately 80% accuracy on MMLU — compared to 70% for GPT-3.5 in English.

Open in Lab
Hover over a language to see GPT-4's MMLU accuracy compared to prior English baselines.
The demo wakes as you arrive…

Seeing the world: multimodal vision

GPT-4 accepts prompts consisting of both images and text — a capability absent from prior GPT models. The model can interpret diagrams, photographs, screenshots, and memes, generating text outputs that demonstrate genuine visual understanding. Standard techniques like few-shot prompting and chain-of-thought reasoning work just as well with visual inputs.

Think of it as giving the model eyes: where GPT-3.5 could only read the description of a chart, GPT-4 can look at the chart itself and answer questions about it. The model can read handwritten text, explain the humor in a meme, solve physics problems from diagrams, and summarize academic papers by looking at their figures.

Open in Lab
Click a category to see examples of GPT-4's visual reasoning capabilities.
The demo wakes as you arrive…

Still not perfect: known limitations

Despite its capabilities, GPT-4 has the same fundamental limitations as earlier GPT models:

  • Hallucinations: the model can generate fluent text that is factually wrong, sometimes with high confidence. GPT-4 improved factual accuracy by 19 percentage points over GPT-3.5 on adversarial factuality tests, but hallucinations remain.
  • Limited : the model processes a fixed-length input and cannot consult external sources mid-generation.
  • No learning from experience: unlike humans, GPT-4 does not learn from its interactions after training is complete.
  • loss after RLHF: the pre-trained model is highly calibrated (its confidence matches its accuracy), but RLHF reduces this calibration — making the model sometimes overconfident or underconfident.
Open in Lab
Compare pre-training calibration (left, nearly perfect diagonal) vs post-RLHF calibration (right, degraded). Drag the slider to see how RLHF affects confidence.
The demo wakes as you arrive…

Making it safe: the model-assisted safety pipeline

GPT-4 introduced a two-pronged approach to safety that uses the model itself as a tool for its own .

The first prong is adversarial testing by domain experts: over 50 specialists in cybersecurity, biosecurity, and international security probed the model for dangerous capabilities. Their findings informed targeted data collection — for example, teaching GPT-4 to refuse requests about synthesizing dangerous chemicals.

The second prong is rule-based reward models (RBRMs): zero-shot GPT-4 classifiers that provide an additional reward signal during RLHF training. Given a prompt, a model response, and a human-written rubric (e.g., "classify as: desired refusal / undesired refusal / disallowed content / safe response"), the RBRM classifies the output. This lets the system reward correct refusals on dangerous prompts and penalize over-cautious refusals on safe prompts.

The result: an 82% reduction in responses to disallowed content compared to GPT-3.5, and a 29% improvement in responding appropriately to sensitive requests.

Open in Lab
Step through the safety pipeline to see how a prompt flows through expert testing, RLHF, and RBRM classification.
The demo wakes as you arrive…

The complete GPT-4 pipeline

GPT-4's training pipeline has three major stages, each building on the one before it:

Stage 1 — Pre-training: the model learns to predict the next token from a massive corpus of text and images from the internet and licensed sources. This is where the model acquires its broad knowledge of language, reasoning, facts, and code. The scaling laws described earlier allow engineers to predict the outcome of this stage.

Stage 2 — (SFT): human labelers write ideal responses to diverse prompts, and the model learns to follow instructions by imitating these examples.

Stage 3 — RLHF with RBRMs: human labelers rank multiple model responses, training a that captures human preferences. The policy model is then optimized against this reward model, with additional rule-based reward signals (RBRMs) to steer safety behavior at a fine-grained level.

Open in Lab
Click each stage to explore what happens at every step of the GPT-4 training pipeline.
The demo wakes as you arrive…

The scaling law in code

Fitting a scaling law to predict large model performancepython

Simplified to show the idea — not the real implementation.

import numpy as np
from scipy.optimize import curve_fit

def scaling_law(C, a, b, c):
    """Power law with irreducible loss: L(C) = a * C^b + c"""
    return a * np.power(C, b) + c

# Compute budgets (normalized) and final losses for small models
small_compute = np.array([1e-4, 3e-4, 1e-3, 3e-3, 0.01, 0.03, 0.1])
small_losses  = np.array([5.8,  5.2,  4.5,  4.0,  3.6,  3.2,  2.8])

# Fit the scaling law to small models only
params, _ = curve_fit(scaling_law, small_compute, small_losses,
                      p0=[5.0, -0.1, 1.5], maxfev=10000)
a, b, c = params

# Predict GPT-4's performance (compute = 1.0, normalized)
gpt4_predicted_loss = scaling_law(1.0, a, b, c)
gpt4_actual_loss    = 1.8  # actual final loss

print(f"Predicted: {gpt4_predicted_loss:.2f}")
print(f"Actual:    {gpt4_actual_loss}")
print(f"Irreducible loss floor: {c:.2f}")
# The prediction closely matches — that's predictable scaling.

Why it mattered

  1. 2018

    GPT-1 — 117M parameters

    Proved that unsupervised pre-training on a Transformer decoder transfers to downstream tasks. The seed of the "predict the next word" paradigm.

  2. 2019

    GPT-2 — 1.5B parameters

    Scaling 10× showed emergent zero-shot abilities. OpenAI initially withheld the full model citing misuse concerns — the first major AI safety debate.

  3. 2020

    GPT-3 — 175B parameters

    Few-shot learning without fine-tuning. The model could translate, write code, and do arithmetic from examples alone. Demonstrated that size unlocks emergent abilities.

  4. 2022

    ChatGPT (GPT-3.5) — RLHF revolution

    Instruction-tuning and RLHF turned a language model into a conversational product. 100 million users in two months. Proved alignment techniques work at scale.

  5. 2023

    GPT-4 — multimodal, predictable, safer

    Added vision, predictable scaling, and a model-assisted safety pipeline. Passed the bar exam in the top 10%. Transformed LLMs from research tools to professional infrastructure.

  6. 2024

    GPT-4o — omni-modal

    Extended multimodal capabilities to speech, enabling real-time voice conversation with reasoning across text, image, and audio natively.

CitationOpenAI. GPT-4 Technical Report. arXiv, 2023.

Terms in this paper