Language Models2021intermediate13 min read

WebGPT: Browser-Assisted Question-Answering with Human Feedback

WebGPT: الإجابة عن الأسئلة بمساعدة متصفّح ويب وتغذية راجعة بشرية

Nakano, R. · Hilton, J. · Balaji, S. · Wu, J. · Ouyang, L. · Kim, C. · Hesse, C. · Jain, S. · Kosaraju, V. · Saunders, W. · Jiang, X. · Cobbe, K. · Eloundou, T. · Krueger, G. · Button, K. · Knight, M. · Chess, B. · Schulman, J. — arXiv

The problem

By 2021, large language models like GPT-3 could generate fluent, convincing text, but they had two fundamental weaknesses. First, they answered questions purely from memory — whatever patterns they memorized during pre- — with no ability to look up or verify facts. This meant answers could be confidently wrong, a problem known as . Second, there was no way for users to check where information came from, since answers had no citations. Long-form question-answering (LFQA) required both accurate retrieval and coherent synthesis, and existing systems tackled these separately, achieving results well below human quality.

The contribution

WebGPT introduces a text-based web-browsing environment that GPT-3 can interact with — searching Bing, clicking links, scrolling pages, and quoting passages as references. The model is first trained via on human demonstrations, then a is trained on human preference comparisons, and finally answer quality is optimized using (best-of-n) against the reward model. The best configuration (175B best-of-64) produces answers preferred by humans 56% of the time over human demonstrators and 69% over Reddit's highest-voted answers, while also collecting inline citations that make fact-checking straightforward.

The impact

WebGPT demonstrated that language models can learn to use external tools — specifically a web browser — through , pioneering the tool-use paradigm that would later define systems like ChatGPT's browsing mode. It also established the pattern of combining with reward modeling that became central to RLHF-based alignment. The idea that models should cite their sources influenced the design of systems. Together with InstructGPT, WebGPT laid the groundwork for the modern approach to building AI assistants that are both helpful and grounded in verifiable evidence.

Imagine a brilliant student who has memorized an entire encyclopedia. Ask her anything and she will give a confident, articulate answer — but sometimes she invents facts, and you have no way to check. Now imagine giving that student a laptop with a web browser, teaching her to search, read, and take notes with page numbers. Suddenly her answers come with footnotes, and you can verify every claim.

But here is the twist: we did not just teach her how to browse — we let thousands of people vote on which of her answers they preferred, and used those votes to make her even better. The result? An assistant whose answers humans prefer over answers written by other humans.

That is WebGPT: GPT-3 with a browser, trained by watching humans, then refined by human preferences.

The browsing environment: giving GPT-3 a web browser

The core innovation is deceptively simple: instead of trying to improve document retrieval with a new neural architecture, the authors gave GPT-3 access to a familiar, powerful existing tool — the Bing search engine — through a text-based interface. The model sees a text summary of the current browser state: the question, the visible portion of the current web page, a scrollbar position indicator, and a list of past actions. It then issues one of a fixed set of commands — search, click a link, scroll, quote a passage, or end browsing — and the cycle repeats.

Think of it as a command-line browser: there are no images, no JavaScript rendering, just cleaned-up text from each page. The model has no memory between steps — the only record of what it has done lives in a compact summary appended to each new prompt. This stateless design is critical: it means the entire browsing session can be represented as a sequence of text completions, making it compatible with standard .

While browsing, the model can quote passages from the current page. Each quote is saved with its page title and URL as a reference. When the model finishes browsing, it is prompted with the question and all collected quotes, and composes a final answer with inline citations. This two-phase design — browse then answer — forces the model to ground every claim in evidence it found.

Open in Lab
Step through a simulated browsing session — watch the model search, click, scroll, and quote before composing a cited answer.
The demo wakes as you arrive…

The training pipeline: from imitation to optimization

WebGPT's training follows a four-stage pipeline, each building on the previous one. The key insight from Learning to Summarize is that imitating human behavior (behavior cloning) gets you a working baseline, but optimizing for human preferences using a reward model takes you further — potentially beyond human-level performance.

Stage 1 — Behavior Cloning (BC). Human demonstrators use a graphical browser interface to answer questions from the ELI5 dataset. They search, navigate, quote references, and compose answers. Around 6,000 of these demonstrations are collected. GPT-3 is then fine-tuned on these demonstrations using , with the human's commands as labels. This gives the model a working ability to use the browser.

Stage 2 — Reward Modeling (RM). The BC model generates pairs of answers to the same question. Human labelers compare each pair and indicate which answer is better based on factual accuracy, coherence, and usefulness. Around 21,500 comparisons are collected. A reward model is trained to predict these human preferences, outputting a scalar score that represents an Elo-style quality rating.

Stage 3 — (RL). The BC model is further fine-tuned using PPO (), with the reward model providing the environment reward at the end of each browsing episode. A KL penalty from the BC model is added at each to prevent the from drifting too far and overoptimizing the reward model.

Stage 4 — Rejection Sampling (Best-of-n). Multiple answers are sampled from the BC or RL model, and the one scoring highest on the reward model is selected. This is a simple but powerful alternative to RL that requires no additional training — only more -time compute. The best configuration samples 64 answers and picks the best one.

Open in Lab
Click each stage to see how data flows through the four-stage training pipeline — from human demonstrations to the final optimized model.
The demo wakes as you arrive…

The reward model: learning what humans prefer

The reward model is the engine that drives quality beyond imitation. It takes a question, an answer, and the collected references as input, and outputs a single scalar — an Elo score that represents how much a human labeler would prefer this answer.

Training uses the comparison data: given two answers A and B to the same question, the reward model learns to assign a higher score to whichever answer the human preferred. The difference between two reward scores represents the log-odds of one being preferred over the other. If answer A has a reward 1 point higher than answer B, then A would be preferred about 73% of the time.

The reward model is initialized from the BC model with the final unembedding layer removed, meaning it inherits the same reading comprehension and language understanding. It is then trained with a cross-entropy loss on the comparison labels, with ties treated as soft 50% labels. The labelers evaluated answers on multiple dimensions — unsupported claims, relevance, coherence, citation quality — but only the final overall preference was used for training, keeping the objective simple and robust.

loss(θ)=−E(q,aw,al)[log⁡σ(rθ(q,aw)−rθ(q,al))]\text{loss}(\theta) = -\mathbb{E}_{(q,a_w,a_l)} \left[ \log \sigma \bigl( r_\theta(q, a_w) - r_\theta(q, a_l) \bigr) \right]
Reward model loss — the model learns to score winning answers higher than losing ones — Given a question qq, a preferred answer awa_w and a dispreferred answer ala_l, the reward model rθr_\theta is trained to produce a higher scalar score for awa_w. The sigmoid σ\sigma converts the score difference to a probability. This is equivalent to modeling human preference as a Bradley-Terry model over Elo ratings.
Open in Lab
See how the reward model scores two answers to the same question — adjusting reference quality and accuracy shifts which answer wins.
The demo wakes as you arrive…

Rejection sampling: the surprisingly effective shortcut

One of the most striking findings in the paper is that rejection sampling — simply generating many answers and picking the best one according to the reward model — works better than reinforcement learning. The 175B best-of-64 model (behavior cloning + rejection sampling) is preferred 68% of the time over the plain BC model, while RL alone gets only 58%.

Why does this outperform RL? The authors offer several explanations. First, rejection sampling gives the model many independent browsing attempts — it can visit different websites each time and evaluate the information with hindsight. Second, the reward model was trained primarily on data from BC and rejection sampling policies, making it more robust to overoptimization by rejection sampling than by RL. Third, RL requires careful tuning, while rejection sampling is essentially tuning-free.

Think of it as a job interview: RL is like coaching one candidate to give better answers, while rejection sampling is like interviewing 64 candidates and hiring the best one. When the "interviewer" (reward model) is reliable, selecting the best from a large pool can outperform intensive coaching of a single candidate.

Open in Lab
Sample multiple answers and watch the reward model select the best one. Increase the number of samples (n) to see how quality improves.
The demo wakes as you arrive…

Results: surpassing human demonstrators

The paper evaluates WebGPT in three ways, each revealing a different aspect of its capabilities.

ELI5 vs. human demonstrators. The 175B best-of-64 model is preferred 56% of the time over answers written by the human demonstrators who trained it. This is a meaningful threshold: you would not expect to exceed 50% by imitation alone, so the reward model optimization is clearly providing additional value.

ELI5 vs. Reddit. The same model is preferred 69% of the time over Reddit's highest-voted answers. For fairness, citations were stripped from the model's answers for this comparison, and new labelers with minimal instructions were used. Even without citation formatting advantages, the model substantially outperformed community-generated answers.

. On an adversarially constructed dataset of short-form questions designed to elicit common misconceptions, WebGPT answered truthfully 75% of the time and was both truthful and informative 54% of the time — substantially better than base GPT-3, though still below human performance. Importantly, truthfulness increased with model size for WebGPT, unlike for the base GPT-3 model where larger models tend to produce more confident falsehoods.

Open in Lab
Compare WebGPT evaluation results across different model sizes and baselines.
The demo wakes as you arrive…

Truthfulness: two kinds of falsehood

The paper introduces a useful distinction between two categories of false statements that language models make.

Imitative falsehoods are false claims that the model is incentivized to produce by its training objective — even with infinite data and compute. For example, reproducing common misconceptions that appear frequently in the training data. WebGPT reduces these because it can look up facts rather than relying on memorized patterns, and because it is incentivized to prefer reliable sources.

Non-imitative falsehoods are false claims that result from the model failing at its training objective — most notably hallucinations, where the model generates plausible- sounding but fabricated information. WebGPT reduces these because retrieval-augmented generation inherently reduces hallucination, and because the model must ground its answer in collected references.

However, neither category is eliminated. On out-of-distribution questions (like TruthfulQA's adversarial prompts), WebGPT sometimes quotes from highly unreliable sources. And while non-imitative falsehoods are reduced, the model still occasionally makes mistakes when paraphrasing or synthesizing information from its references.

Scaling behavior: data, parameters, and samples

The paper provides detailed analysis across three dimensions.

For data scaling, doubling the number of demonstrations increased the policy's reward model score by about 0.13, and doubling the number of comparisons increased the reward model accuracy by about 1.8%. Returns are diminishing but far from exhausted.

For parameter scaling, doubling the number of parameters in the policy increased its reward model score by roughly 0.09, and doubling the reward model parameters increased accuracy by roughly 0.4%. The trends are noisier than data scaling.

For rejection sampling scaling, the authors analyzed compute-efficient trade-offs between model size and number of samples. The Pareto frontier favors some rejection sampling at every model size: the 760M best-of-4, 13B best-of-16, and 175B best-of-64 configurations are the compute-efficient models. Using too many samples provides diminishing returns because the reward model can be overoptimized.

Risks and broader implications

The paper raises several important concerns about deploying web-browsing language models.

Bias reinforcement. WebGPT inherits GPT-3's biases and amplifies them through search and synthesis. The model tends to accept the implicit assumptions of questions, which could reinforce users' confirmation bias. When asked "What does a wedding look like?" it overwhelmingly assumes a Western, American perspective.

Live web access risks. While WebGPT's actions are limited to Bing searches and link following, the authors note that more capable models with web access could potentially exploit real-world side effects — like editing Wikipedia to create reliable-looking references. They argue that as model capabilities increase, so should the safety burden for granting web access.

Reference cherry-picking. The model is incentivized to find references that labelers will find convincing, not references that reflect a fair assessment of evidence. This is analogous to a lawyer building a one-sided case. The authors suggest methods like debate — where models are trained to find evidence both for and against claims — as a mitigation.

Question stance sensitivity. When questions were phrased to affirm misconceptions (e.g., "Why did the government fake the moon landing?"), the model was more likely to produce inaccurate answers that reinforce the false belief, compared to neutrally or skeptically framed versions of the same question.

Legacy: from browsing to tool-using agents

  1. 2020

    Learning to Summarize (Stiennon et al.)

    Established the pattern of behavior cloning → reward modeling → RL optimization using human feedback for text generation tasks. WebGPT builds directly on this.

  2. 2021

    WebGPT (this paper)

    Gave GPT-3 a text-based web browser, trained it with human demonstrations and preferences, and achieved human-level question-answering with inline citations.

  3. 2022

    InstructGPT (Ouyang et al.)

    Applied RLHF to instruction-following, training GPT-3 to be helpful, harmless, and honest. Shared several authors and methods with WebGPT.

  4. 2022

    ReAct (Yao et al.)

    Combined reasoning and acting in language models, interleaving thought traces with tool-use actions. Extended WebGPT's tool-use paradigm to multi-step reasoning.

  5. 2023

    Toolformer (Schick et al.)

    Taught language models to self-supervisedly learn when and how to call external APIs (search, calculator, translator). Generalized WebGPT's single-tool approach to arbitrary tool use.

  6. 2023

    WebArena (Zhou et al.)

    Created a realistic web environment benchmark for evaluating autonomous web agents, building on WebGPT's vision of language models as web users.

WebGPT stands at a pivotal junction in AI history. Looking backward, it drew on GPT-3's language understanding and Learning to Summarize's RLHF pipeline. Looking forward, it pioneered the tool-use paradigm that defines modern AI assistants: the idea that a language model should not try to know everything, but should know how to find everything and present it with verifiable evidence.

CitationNakano, Hilton, Balaji, Wu, Ouyang, Kim, Hesse, Jain, Kosaraju, Saunders, Jiang, Cobbe, Eloundou, Krueger, Button, Knight, Chess, Schulman. WebGPT: Browser-Assisted Question-Answering with Human Feedback. arXiv, 2021.

Terms in this paper