Language Models2023advanced12 min read

Sparks of Artificial General Intelligence: Early Experiments with GPT-4

شرارات الذكاء الاصطناعي العام: تجارب مبكرة مع GPT-4

Bubeck, S. · Chandrasekaran, V. · Eldan, R. · Gehrke, J. · Horvitz, E. · Kamar, E. · Lee, P. · Lee, Y. T. · Li, Y. · Lundberg, S. · Nori, H. · Palangi, H. · Ribeiro, M. T. · Zhang, Y. — arXiv

The problem

By early 2023, large language models had shown impressive performance on individual benchmarks, but these benchmarks tested narrow skills in isolation. No systematic study had explored whether a single model could exhibit broad, flexible intelligence across many domains simultaneously — the kind of generality that approaches . Moreover, traditional NLP benchmarks risked data contamination, making it unclear whether strong results reflected true understanding or memorization.

The contribution

A 154-page empirical investigation of an early version of GPT-4, testing it across mathematics, coding, vision (via text), medicine, law, psychology, and creative tasks. The authors designed novel prompts to minimize contamination risk and showed that GPT-4 performs at or near human level across these domains — a breadth no prior model achieved. They also systematically documented limitations: hallucinations, arithmetic errors, lack of , inability to backtrack, and poor . The paper adopted a psychology-style evaluation methodology rather than standard ML benchmarking.

The impact

The "Sparks" paper ignited a fierce debate on whether scaling language models leads toward AGI or merely produces sophisticated pattern matching. It shifted the evaluation paradigm from leaderboard benchmarks to open-ended qualitative testing. It directly influenced the wave of capability evaluations, red-teaming frameworks, and the broader discourse on AI safety and governance. The paper remains one of the most cited AI publications of 2023 and a landmark reference in the AGI debate.

Imagine a student who has never attended a single class at your university, yet walks into every final exam — law, medicine, mathematics, computer science, psychology — and scores at or near the top of the class. No one taught them these subjects directly; they learned by reading the entire library.

You would be astonished — and suspicious. Does this student truly understand, or are they an extraordinary mimic? This paper puts GPT-4 through exactly that gauntlet: hundreds of novel exams across dozens of fields, designed to test whether the model is genuinely reasoning or just performing a dazzling parlor trick.

A new lens: testing intelligence, not benchmarks

Traditional machine learning evaluates models on fixed benchmarks: accuracy on ImageNet, BLEU scores on translation, F1 on question answering. But the authors argue that these benchmarks are insufficient for GPT-4. First, the model may have seen the test data during (data contamination). Second, narrow benchmarks miss the breadth of capabilities that distinguishes general intelligence from narrow expertise.

Instead, the authors adopt an approach inspired by psychology: design novel, creative tasks that probe reasoning, abstraction, and cross-domain transfer — tasks GPT-4 could not have memorized. They reference the 1994 definition of intelligence from a panel of 52 psychologists: "A very general mental capability that, among other things, involves the ability to reason, plan, solve problems, think abstractly, comprehend complex ideas, learn quickly and learn from experience."

This framing is deliberate and provocative. By using a human-intelligence definition as their yardstick, the authors invite direct comparison — and direct controversy.

The evidence: GPT-4 across domains

The core argument rests on GPT-4's performance across a remarkable range of domains. The authors test it in coding, mathematics, medicine, law, psychology, vision (via text descriptions), creative writing, and more. In each case, they compare GPT-4 against ChatGPT (GPT-3.5) to highlight the leap in capability. What makes this convincing is not any single result, but the consistency across domains — the same model excels at tasks that normally require years of specialized training.

Open in Lab
GPT-4 vs ChatGPT across domains. Click a domain to see example tasks and results from the paper.
The demo wakes as you arrive…

Coding. GPT-4 can write functional code from natural language descriptions, handle complex data structures, and even solve problems from competitive programming contests. It passes mock technical interviews at the level of a strong software engineering candidate. When given LeetCode-style problems it had never seen, it produced correct solutions with clear explanations.

Mathematics. The model solves novel math problems requiring multi-step reasoning, proves theorems, and even writes proofs in verse. However, arithmetic remains a weakness — it can reason about mathematical concepts but stumbles on raw computation like multiplying large numbers.

Medicine. Given clinical vignettes, GPT-4 produces differential diagnoses that match expert reasoning. It passed the USMLE medical licensing exam at a level competitive with medical students, demonstrating not just factual recall but clinical reasoning.

Law. GPT-4 scored in the top percentiles on the Uniform Bar Exam, demonstrating legal reasoning, case analysis, and argumentation — tasks requiring integration of rules, precedent, and factual application.

Understanding humans: theory of mind

One of the most surprising findings is GPT-4's apparent capacity for — the ability to reason about what other people believe, intend, and feel, even when those beliefs differ from reality. This is a hallmark of human social cognition, typically tested with "false belief" tasks in developmental psychology.

The authors designed scenarios where GPT-4 must track what different characters know and don't know. For example: Alice saves a file to a shared folder, then Bob moves it without telling her. When asked "Where will Alice look for the file?", GPT-4 correctly answers the original location — demonstrating that it can model Alice's false belief rather than simply reporting the file's actual location.

This is significant because theory of mind was long considered beyond the reach of language models. It requires maintaining separate mental models for different agents and reasoning about the gap between belief and reality.

Open in Lab
Follow a false-belief scenario step by step. See how GPT-4 tracks each character's knowledge state separately.
The demo wakes as you arrive…

Tools and agency: beyond text generation

Perhaps the most forward-looking section of the paper examines GPT-4's ability to use tools. The authors show that with a simple prompt instructing it that certain functions are available (search engine, calculator, API calls), GPT-4 can autonomously decide when to invoke them, compose the correct function calls, and integrate the results back into its reasoning.

This is profound because it transforms the model from a text generator into an . It can search the web for current information, call a calculator for precise arithmetic (compensating for its own weakness), and chain multiple tool calls to solve complex tasks. The authors note that is one of the hallmarks of human intelligence and was historically considered beyond AI's reach.

GPT-4 can even use itself as a tool — for instance, generating a plan, then critiquing that plan in a separate prompt, then revising based on the critique. This recursive self-improvement loop hints at a form of metacognition.

The cracks: where GPT-4 fails

The authors are careful not to claim that GPT-4 is AGI. They document four systematic failure modes that reveal fundamental limitations of the autoregressive architecture.

. GPT-4 confidently generates false information — invented citations, incorrect facts, fabricated statistics. It cannot distinguish what it knows from what it is guessing. This is not a minor bug but a structural feature: the model is trained to produce plausible text, not true text. There is no internal fact-checking mechanism.

Arithmetic and calculation. Despite strong mathematical reasoning, GPT-4 fails at basic arithmetic. It can set up complex equations but miscalculate 7×4+8×87 \times 4 + 8 \times 8. The limitation is revealing: the model processes numbers as tokens, not as quantities with mathematical structure.

Planning and . GPT-4 struggles with tasks that require maintaining a complex state over many steps. Merging two lists, playing chess at an expert level, or writing a story with strict structural constraints all degrade as complexity grows.

No backtracking. The autoregressive architecture generates tokens left-to-right with no ability to revise. If the model starts down a wrong path, it cannot backtrack. Each is committed the moment it is produced. This makes the model fragile on tasks where early decisions constrain later ones.

Open in Lab
See why autoregressive generation cannot backtrack. The model commits to each token permanently — a wrong early choice cascades through the rest.
The demo wakes as you arrive…

Calibration: does GPT-4 know what it knows?

A well-calibrated model should be uncertain when it is likely wrong and confident when it is likely right. The authors find that GPT-4's calibration is mixed. On some domains it shows reasonable calibration — when it gives a high-confidence answer, it is usually correct. But on others, especially factual recall, it is confidently wrong at alarming rates.

This is a critical safety issue. A model that sounds certain even when wrong is more dangerous than one that hedges, because users learn to trust the confident tone. The authors argue that improving calibration is one of the most important challenges for future models.

Open in Lab
Explore GPT-4's calibration across domains. Well-calibrated means the bar heights match the accuracy — gaps reveal overconfidence.
The demo wakes as you arrive…

The controversy: is this really AGI?

The paper's most provocative claim — that GPT-4 shows "sparks" of AGI — drew fierce criticism. Critics raised several objections. First, the evaluation is largely qualitative: the authors select impressive examples rather than running controlled experiments with statistical significance. Cherry-picking successes while acknowledging failures is not the same as rigorous evaluation. Second, there is no agreed definition of AGI, and using a 1994 psychology definition as the benchmark is contested — it was never intended for AI systems.

Third, and most fundamentally: does passing exams mean understanding? A model that produces correct answers might be performing sophisticated pattern matching without any internal representation of meaning. The philosopher's zombie objection applies: the outputs look intelligent, but is anything "home" inside?

The authors anticipate these critiques and respond: they ask whether the distinction between "true understanding" and "very convincing improvisation" is even meaningful. If GPT-4 can solve novel problems it has never seen, in domains it was never explicitly trained for, using reasoning steps that mirror human cognition — at what point does the distinction collapse?

Open in Lab
Explore the arguments for and against the AGI interpretation. Each card presents a claim and its strongest counterargument.
The demo wakes as you arrive…

Methodological questions

The paper's methodology breaks from ML conventions in ways that are both its strength and its weakness. On one hand, the creative open-ended testing is precisely what reveals GPT-4's breadth — standard benchmarks would miss most of what makes it interesting. On the other hand, the lack of systematic controls, statistical testing, and blind evaluation makes it hard to distinguish genuine capability from selection bias.

The authors tested an early version of GPT-4 during active development, meaning the model they evaluated may differ from the publicly released version. They also note that prompt sensitivity — small changes in wording producing large differences in output — makes reproducibility challenging. The same model can appear brilliant or mediocre depending on how you ask the question.

Despite these concerns, the paper's impact is undeniable. By shifting the conversation from "how well does it score?" to "what can it do?", it opened a richer and more productive discourse about AI capabilities.

What comes next: the path forward

The authors conclude with challenges that must be addressed for future progress. They suggest that the next-token prediction paradigm may have inherent ceilings — perhaps a fundamentally different architecture is needed for true planning, reliable factual grounding, and continuous learning.

They also raise societal concerns: a model this capable but this unreliable creates risks in high-stakes domains like medicine, law, and education. The potential for misuse — persuasion, manipulation, disinformation — scales with capability. They call for new evaluation frameworks, safety mechanisms, and governance structures.

  1. 2020

    GPT-3 — In-context learning emerges

    175B parameters. Few-shot prompting works without fine-tuning. The scale-capabilities connection becomes clear, but performance is inconsistent and domain-specific.

  2. 2022

    ChatGPT — Instruction following for the masses

    Fine-tuned GPT-3.5 with RLHF. Impressive dialogue but weak on reasoning, math, and factual accuracy. Establishes the conversational AI paradigm.

  3. 2023

    GPT-4 — The "Sparks" moment

    A leap across domains: coding, math, medicine, law, and creative tasks all approach human-level. The Sparks paper systematically documents both the capabilities and the cracks.

  4. 2023

    The counter-arguments — "Are Emergent Abilities a Mirage?"

    Schaeffer et al. argue that emergent abilities are an artifact of evaluation metrics, not sudden phase transitions. The debate between "sparks of AGI" and "sophisticated pattern matching" intensifies.

Code: simulating a capability probe

The paper's evaluation style — probing with creative tasks — can be distilled into a simple prompt-and-judge loop. Below is a simplified version of how one might systematically test a model across domains, score it, and compute per-domain accuracy.

Multi-domain capability evaluation looppython

Simplified to show the idea — not the real implementation.

# Simplified multi-domain capability evaluation
# inspired by the "Sparks of AGI" methodology.

domains = {
    "math":    ["Prove √2 is irrational", "Solve x²+3x-10=0"],
    "coding":  ["Write quicksort in Python", "Parse a JSON API response"],
    "law":     ["Explain res judicata", "Draft a cease-and-desist"],
    "medicine":["Diagnose: chest pain + dyspnea + JVD", "Explain ACE-inhibitor mechanism"],
}

results = {}
for domain, tasks in domains.items():
    correct = 0
    for task in tasks:
        response = model.generate(task)
        score = expert_judge(response, task)  # human or model judge
        correct += score
    accuracy = correct / len(tasks)
    results[domain] = accuracy
    print(f"{domain}: {accuracy:.0%}")

# Key insight: AGI = consistent high accuracy ACROSS domains,
# not just excellence in one. Narrow AI excels in one domain;
# general intelligence excels broadly.
avg = sum(results.values()) / len(results)
print(f"\nCross-domain average: {avg:.0%}")

CitationBubeck, Chandrasekaran, Eldan, Gehrke, Horvitz, Kamar, Lee, Lee, Li, Lundberg, Nori, Palangi, Ribeiro, Zhang. Sparks of Artificial General Intelligence: Early Experiments with GPT-4. arXiv, 2023.

Terms in this paper