Language Models2023intermediate10 min read
GPT-4 Technical Report
التقرير التقني لـ GPT-4
OpenAI — arXiv
The problem
By 2022, large language models like GPT-3.5 had shown impressive abilities but remained unpredictable at scale: you couldn't know how well a trillion-dollar training run would perform until it finished. Models also struggled with factual accuracy, safety, and could not process images alongside text. The gap between what LLMs could do and what professionals needed was still wide.
The contribution
GPT-4: a large-scale model that accepts both text and image inputs and produces text outputs. Three core contributions: (1) — infrastructure and optimization methods that let engineers accurately predict GPT-4's performance from models trained with 1/1,000th the compute using power-law scaling laws. (2) Human-level exam performance — GPT-4 passes the bar exam in the top 10% and achieves 86.4% on , beating prior state-of-the-art across 57 subjects. (3) A model-assisted using (RBRMs) and that reduced disallowed content by 82% over GPT-3.5.
The impact
GPT-4 demonstrated that large language models could achieve professional-grade performance across law, medicine, and STEM — transforming how society views AI capability. Its predictable scaling methodology showed that expensive training runs need not be gambles. The safety pipeline established a template adopted industry-wide. GPT-4 powers ChatGPT, Microsoft Copilot, and hundreds of enterprise applications, marking the moment LLMs transitioned from research curiosity to critical infrastructure.
Training a is like planning a skyscraper: the taller you build, the more unknowns you face — will the foundation hold? Will the cost spiral? Traditionally, you build the whole thing and pray.
GPT-4's breakthrough is like discovering that a scale model in a wind tunnel can accurately predict how the real skyscraper will sway. By testing tiny models first, the engineers knew what to expect from the full giant — before spending the compute to build it.
The skyscraper also has a security system: a model-assisted safety pipeline that watches for harmful outputs the way a smart alarm watches for intruders — using the building's own intelligence to protect its occupants.
What is GPT-4?
GPT-4 is a large-scale built on the Transformer architecture. Like its predecessors, it is pre-trained to predict the next in a document, then fine-tuned using (RLHF) to follow instructions and behave safely. But GPT-4 introduces two leaps: it can accept images alongside text as input, and its builders developed methods to predict its performance before training finished.
The report deliberately withholds architecture details — model size, hardware, training compute, and dataset construction — citing competitive and safety reasons. What it does reveal is how GPT-4 performs, how its performance was predicted, and how its safety was improved.
Predictable scaling: knowing the answer before the exam
The central engineering insight of GPT-4 is that performance at massive scale can be predicted from small-scale experiments. Training GPT-4 required enormous compute — far too expensive for trial and error. Instead, the team trained much smaller models using the same methodology and fitted a to predict the final .
The intuition: imagine plotting the fuel efficiency of toy cars at different sizes. If the data follows a smooth curve, you can predict what a full-sized car will achieve without building one. The key formula is a with an term that captures the floor below which no amount of compute can push.
The prediction matched GPT-4's actual final loss with remarkable accuracy — this was done before the full training run completed, using models trained with at most 1/10,000th the compute.
Human-level exam performance
GPT-4 was tested on exams designed for humans — bar exams, SATs, AP tests, medical knowledge assessments, and coding challenges — with no exam-specific training. The results were striking:
- Uniform Bar Exam: scored ~298/400, placing in the top 10% of test takers (GPT-3.5 scored in the bottom 10%).
- SAT Math: 700/800 (~89th percentile).
- MMLU (57 academic subjects): 86.4% accuracy, beating all prior models and most benchmark-specialized systems.
- GRE Verbal: 169/170 (~99th percentile).
These capabilities stem primarily from pre-training, not from RLHF — the base model and the RLHF model performed nearly identically on multiple-choice exams (73.7% vs 74.0% averaged across all exams).
Beyond English: multilingual mastery
To test GPT-4 in languages beyond English, the team translated the entire MMLU benchmark into 26 languages using Azure Translate. GPT-4 outperformed the English-language scores of both GPT-3.5 and prior models (Chinchilla and PaLM) in 24 of 26 languages — including low-resource languages such as Latvian, Welsh, and Swahili.
This is remarkable because GPT-4 was not specifically trained on multilingual exam data. The model's understanding of concepts transfers across language boundaries. For Arabic, GPT-4 achieved approximately 80% accuracy on MMLU — compared to 70% for GPT-3.5 in English.
Seeing the world: multimodal vision
GPT-4 accepts prompts consisting of both images and text — a capability absent from prior GPT models. The model can interpret diagrams, photographs, screenshots, and memes, generating text outputs that demonstrate genuine visual understanding. Standard techniques like few-shot prompting and chain-of-thought reasoning work just as well with visual inputs.
Think of it as giving the model eyes: where GPT-3.5 could only read the description of a chart, GPT-4 can look at the chart itself and answer questions about it. The model can read handwritten text, explain the humor in a meme, solve physics problems from diagrams, and summarize academic papers by looking at their figures.
Still not perfect: known limitations
Despite its capabilities, GPT-4 has the same fundamental limitations as earlier GPT models:
- Hallucinations: the model can generate fluent text that is factually wrong, sometimes with high confidence. GPT-4 improved factual accuracy by 19 percentage points over GPT-3.5 on adversarial factuality tests, but hallucinations remain.
- Limited : the model processes a fixed-length input and cannot consult external sources mid-generation.
- No learning from experience: unlike humans, GPT-4 does not learn from its interactions after training is complete.
- loss after RLHF: the pre-trained model is highly calibrated (its confidence matches its accuracy), but RLHF reduces this calibration — making the model sometimes overconfident or underconfident.
Making it safe: the model-assisted safety pipeline
GPT-4 introduced a two-pronged approach to safety that uses the model itself as a tool for its own .
The first prong is adversarial testing by domain experts: over 50 specialists in cybersecurity, biosecurity, and international security probed the model for dangerous capabilities. Their findings informed targeted data collection — for example, teaching GPT-4 to refuse requests about synthesizing dangerous chemicals.
The second prong is rule-based reward models (RBRMs): zero-shot GPT-4 classifiers that provide an additional reward signal during RLHF training. Given a prompt, a model response, and a human-written rubric (e.g., "classify as: desired refusal / undesired refusal / disallowed content / safe response"), the RBRM classifies the output. This lets the system reward correct refusals on dangerous prompts and penalize over-cautious refusals on safe prompts.
The result: an 82% reduction in responses to disallowed content compared to GPT-3.5, and a 29% improvement in responding appropriately to sensitive requests.
The complete GPT-4 pipeline
GPT-4's training pipeline has three major stages, each building on the one before it:
Stage 1 — Pre-training: the model learns to predict the next token from a massive corpus of text and images from the internet and licensed sources. This is where the model acquires its broad knowledge of language, reasoning, facts, and code. The scaling laws described earlier allow engineers to predict the outcome of this stage.
Stage 2 — (SFT): human labelers write ideal responses to diverse prompts, and the model learns to follow instructions by imitating these examples.
Stage 3 — RLHF with RBRMs: human labelers rank multiple model responses, training a that captures human preferences. The policy model is then optimized against this reward model, with additional rule-based reward signals (RBRMs) to steer safety behavior at a fine-grained level.
The scaling law in code
Simplified to show the idea — not the real implementation.
import numpy as np
from scipy.optimize import curve_fit
def scaling_law(C, a, b, c):
"""Power law with irreducible loss: L(C) = a * C^b + c"""
return a * np.power(C, b) + c
# Compute budgets (normalized) and final losses for small models
small_compute = np.array([1e-4, 3e-4, 1e-3, 3e-3, 0.01, 0.03, 0.1])
small_losses = np.array([5.8, 5.2, 4.5, 4.0, 3.6, 3.2, 2.8])
# Fit the scaling law to small models only
params, _ = curve_fit(scaling_law, small_compute, small_losses,
p0=[5.0, -0.1, 1.5], maxfev=10000)
a, b, c = params
# Predict GPT-4's performance (compute = 1.0, normalized)
gpt4_predicted_loss = scaling_law(1.0, a, b, c)
gpt4_actual_loss = 1.8 # actual final loss
print(f"Predicted: {gpt4_predicted_loss:.2f}")
print(f"Actual: {gpt4_actual_loss}")
print(f"Irreducible loss floor: {c:.2f}")
# The prediction closely matches — that's predictable scaling.Why it mattered
2018
GPT-1 — 117M parameters
Proved that unsupervised pre-training on a Transformer decoder transfers to downstream tasks. The seed of the "predict the next word" paradigm.
2019
GPT-2 — 1.5B parameters
Scaling 10× showed emergent zero-shot abilities. OpenAI initially withheld the full model citing misuse concerns — the first major AI safety debate.
2020
GPT-3 — 175B parameters
Few-shot learning without fine-tuning. The model could translate, write code, and do arithmetic from examples alone. Demonstrated that size unlocks emergent abilities.
2022
ChatGPT (GPT-3.5) — RLHF revolution
Instruction-tuning and RLHF turned a language model into a conversational product. 100 million users in two months. Proved alignment techniques work at scale.
2023
GPT-4 — multimodal, predictable, safer
Added vision, predictable scaling, and a model-assisted safety pipeline. Passed the bar exam in the top 10%. Transformed LLMs from research tools to professional infrastructure.
2024
GPT-4o — omni-modal
Extended multimodal capabilities to speech, enabling real-time voice conversation with reasoning across text, image, and audio natively.
CitationOpenAI. GPT-4 Technical Report. arXiv, 2023.
Terms in this paper
- Scaling Lawقانون التحجيم
- Predictable Scalingالتوسُّع القابل للتنبؤ
- Irreducible Lossالفقد غير القابل للاختزال
- Power Lawقانون القدرة
- RLHFالتعلُّم المعزز من التغذية الراجعة البشرية
- Rule-Based Reward Modelsنماذج المكافأة القائمة على القواعد
- Red Teamingالاختبار العدائي
- Alignmentالمحاذاة
- Safety Pipelineخط أنابيب السلامة
- Multimodalمتعدد الوسائط
- Hallucinationالهلوسة الرقمية
- Calibrationالمعايرة
- MMLUMMLU
- Supervised Fine-Tuningالضبط الدقيق الخاضع للإشراف
- Reward Modelنموذج المكافأة