Speech Recognition2022intermediate14 min read

Robust Speech Recognition via Large-Scale Weak Supervision

التعرّف على الكلام بمتانة عالية من خلال إشراف ضعيف على نطاق واسع

Radford, A. · Kim, J. W. · Xu, T. · Brockman, G. · McLeavey, C. · Sutskever, I. — ICML

The problem

By 2022, the dominant paradigm in was unsupervised pre-training (like Wav2Vec 2.0) followed by on a narrow supervised dataset. While this achieved impressive in-distribution results, it produced brittle models: a system trained on LibriSpeech could score "superhuman" on that dataset yet stumble badly on podcasts, phone calls, or accented speech. Fine-tuning made models memorize dataset quirks rather than learn robust speech understanding. Additionally, each deployment required its own fine-tuning stage — a complex process demanding expert knowledge and clean labeled data for every new domain.

The contribution

Whisper: a family of models trained on 680,000 hours of weakly supervised multilingual audio-transcript pairs scraped from the internet. Instead of self-supervised pre-training plus fine-tuning, Whisper uses a single supervised training stage on noisy but massively diverse data. A multitask format lets one model perform speech recognition, translation, , and . In zero-shot evaluation — with no fine-tuning on any target dataset — Whisper approaches human-level accuracy and robustness, achieving a 55.2% relative error reduction over comparably-performing fine-tuned models on datasets.

The impact

Whisper proved that scaling weakly supervised data is a viable — and arguably simpler — alternative to the self-supervised pre-training pipeline that dominated speech recognition. Its open-source release catalyzed an ecosystem of downstream tools: real-time transcription apps, podcast search engines, multilingual subtitle generators, and medical dictation systems. Whisper became the default speech backbone for many applications, and its approach of trading data quality for data quantity at scale influenced subsequent work across modalities.

Traditional speech recognition is like training a sommelier exclusively in one vineyard: she can identify every bottle from that estate with uncanny precision, but hand her a wine from another region and she's lost.

Whisper takes a different approach — it's a traveler who has tasted wines in every country, from fine restaurants to street markets. The labels are often smudged and the pairings are imperfect, but after 680,000 hours of tasting, this traveler can identify almost any wine you pour, in any glass, under any lighting. That breadth of exposure, not the perfection of any single tasting, is what makes the palate robust.

The problem: narrow training breeds brittle models

Before Whisper, the path to state-of-the-art speech recognition followed a two-stage recipe:

  • Stage 1: Unsupervised pre-training. Train an audio (like Wav2Vec 2.0) on up to 1,000,000 hours of raw unlabeled audio. The encoder learns rich representations of speech sounds, but it cannot produce transcripts on its own — it has no .

  • Stage 2: Supervised fine-tuning. Add a decoder and fine-tune the whole system on a small, carefully labeled dataset (often just ~1,000 hours). This step teaches the model to output words.

The problem? Fine-tuning optimizes for the specific quirks of the training dataset. A model fine-tuned on LibriSpeech — clean audiobook recordings — can score a 1.4% on LibriSpeech yet make twice as many errors as a human on podcasts, phone calls, or noisy recordings. The model learned the style of LibriSpeech, not the skill of understanding speech.

Open in Lab
Toggle between a fine-tuned model and Whisper. Notice how the fine-tuned model excels on its home dataset but collapses on others, while Whisper maintains steady performance everywhere.
The demo wakes as you arrive…

The idea: trade data quality for data quantity

Whisper's insight is radical in its simplicity: instead of carefully curating a small, gold-standard dataset, gather 680,000 hours of audio paired with whatever transcripts already exist on the internet — subtitles, captions, translations — and train a single supervised model on all of it.

The transcripts are noisy: some are machine-generated, some are poorly aligned, some are in the wrong language. But the sheer scale and diversity of the data more than compensate for individual errors. The model sees so many accents, recording conditions, topics, and languages that it cannot overfit to any single distribution. It must learn general speech understanding.

This is at scale: the labels are imperfect, but there are so many of them — and they span such variety — that the signal overwhelms the . The result is a model that works reliably out of the box, with no fine-tuning needed for any specific deployment.

Building the dataset: 680,000 hours from the wild

The dataset is constructed from audio paired with transcripts found on the internet — a vast and diverse source that spans different environments, recording setups, speakers, and languages. However, diversity in audio quality helps robustness while diversity in transcript quality hurts it. Several automated filters address this:

  • Machine-generated transcript detection. Many internet transcripts are outputs of existing ASR systems, not human-written. Training on these would teach the model to produce "transcript-ese" — text with missing punctuation, no capitalization, and no style. Heuristics detect and remove these (e.g., all-uppercase or all-lowercase text is almost certainly machine-generated).

  • Language verification. An audio language detector ensures the spoken language matches the transcript language. Mismatches are excluded from speech recognition training — but if the transcript is in English, the pair is repurposed as training data.

  • Deduplication. Fuzzy transcript deduplication reduces repeated and auto-generated content. After training an initial model, high-error-rate sources are manually inspected and removed.

All audio is broken into 30-second segments paired with the corresponding transcript subset. Segments with no speech are included (at sub-sampled rate) as voice activity detection training data.

Open in Lab
Follow the data from raw internet audio to clean training pairs. Click each stage to see what gets filtered and why.
The demo wakes as you arrive…

The model: an encoder-decoder Transformer

Whisper deliberately uses an off-the-shelf architecture — an encoder-decoder Transformer — so that its results reflect the power of the data and training approach, not architectural novelty. The design choice is intentional: by keeping the architecture standard, any improvement must come from the data.

The encoder processes audio features. All audio is resampled to 16 kHz and converted into an 80-channel log-magnitude (25 ms windows, 10 ms stride). Two 1D layers with GELU activation (the second with stride 2) form a stem, followed by sinusoidal and standard Transformer encoder blocks with pre-activation residual connections.

The decoder is an autoregressive Transformer that attends to the encoder output via . It uses learned positional embeddings and tied input-output token representations. Both encoder and decoder have the same width and number of blocks. The is byte-level (the same as GPT-2 for English-only models, refitted for multilingual models).

Open in Lab
Click any layer to see what it does — from raw audio to predicted tokens.
The demo wakes as you arrive…

The model family spans five sizes, from 39M to 1.55B parameters:

ModelLayersWidthHeadsParametersTiny4384639MBase6512874MSmall1276812244MMedium24102416769MLarge321280201550M\begin{array}{lcccc} \textbf{Model} & \textbf{Layers} & \textbf{Width} & \textbf{Heads} & \textbf{Parameters} \\ \hline \text{Tiny} & 4 & 384 & 6 & 39\text{M} \\ \text{Base} & 6 & 512 & 8 & 74\text{M} \\ \text{Small} & 12 & 768 & 12 & 244\text{M} \\ \text{Medium} & 24 & 1024 & 16 & 769\text{M} \\ \text{Large} & 32 & 1280 & 20 & 1550\text{M} \\ \end{array}
Whisper model family — All sizes share the same architecture; only depth and width differ. Encoder and decoder have identical dimensions. The Large V2 model was trained 2.5× longer with additional regularization (SpecAugment, stochastic depth, BPE dropout).

One model, many tasks: the multitask token format

A traditional speech pipeline chains separate models for each subtask — voice activity detection, language identification, transcription, translation, timestamp alignment. Whisper replaces this entire pipeline with a single decoder that reads a sequence of special tokens specifying what to do.

The decoder is fundamentally an audio-conditional : given audio features from the encoder and a prefix of special tokens, it autoregressively generates the output. The token sequence acts as a task specification:

<|startoftranscript|> → Language tag (e.g. <|en|>) → Task tag (<|transcribe|> or <|translate|>) → Timestamps toggle (<|notimestamps|> or begin-time tokens) → Output text → <|endoftranscript|>

This design means adding a new task requires only adding new special tokens — no architectural change. The same model handles English transcription, non-English transcription, any-to-English translation, voice activity detection (via the <|nospeech|> token), and timestamp-aligned transcription.

Open in Lab
Click a task to see how the special token sequence changes — the same model handles all of them.
The demo wakes as you arrive…

Training: simple recipe, massive scale

The training recipe is deliberately simple. Models are trained with using mixed precision and AdamW optimizer. A of 256 segments and linear decay from a warm start are used. The models train for approximately 2–3 passes over the full dataset.

Because the dataset is so large and diverse, no or is needed — the diversity is the regularization. is not a concern at this scale. For the improved Large V2 model, three additional techniques were added: SpecAugment (frequency/time masking of spectrograms), stochastic depth (randomly dropping layers), and BPE dropout (randomly segmenting subwords differently).

One practical fix: Whisper models initially tried to predict speaker names from transcripts that included them — but speaker identity is rarely inferable from just 30 seconds of audio. A brief fine-tuning on transcripts without speaker annotations removed this behavior.

Results: robust where others are brittle

The core finding of the paper is not about achieving the lowest possible error rate on any single dataset — it's about consistent performance across many datasets. The key metric is : how much does performance degrade when moving from the reference dataset (LibriSpeech) to other, out-of-distribution datasets?

A zero-shot Whisper Large V2 model scores 2.7% WER on LibriSpeech test-clean — roughly matching a wav2vec 2.0 model fine-tuned on LibriSpeech. But on 12 other datasets (podcasts, phone calls, meetings, accented speech), Whisper makes 55.2% fewer errors on average. The fine-tuned model's "superhuman" LibriSpeech performance was an illusion — it had memorized the dataset, not learned speech.

When compared to human transcribers on a diverse set of recordings, Whisper's performance is within a fraction of a percentage point of professional human accuracy.

Open in Lab
Each dot is a model. X-axis: LibriSpeech WER. Y-axis: average WER on other datasets. The diagonal is ideal robustness (equal performance everywhere). Notice how Whisper models approach the human frontier while fine-tuned models fall far below it.
The demo wakes as you arrive…

Noise robustness: Whisper shines in the real world

Real-world audio is rarely clean. To test noise robustness, the authors added white noise and pub noise (ambient restaurant/bar chatter) at various signal-to-noise ratios to LibriSpeech test-clean and compared Whisper with 14 LibriSpeech-trained models.

At high SNR (low noise, 40 dB), many fine-tuned models outperform Whisper — unsurprising since they were trained on clean LibriSpeech audio. But as noise increases, these models degrade rapidly. Below 10 dB SNR (moderately noisy), Whisper outperforms every compared model. The pub noise results are especially striking: Whisper has heard so many noisy recordings during training that restaurant chatter barely affects it.

Open in Lab
Drag the SNR slider from clean (40 dB) to noisy (-10 dB). Watch how fine-tuned models collapse while Whisper degrades gracefully.
The demo wakes as you arrive…

Multilingual and translation capabilities

Of the 680,000 hours in the dataset, 117,000 hours cover 96 languages beyond English and 125,000 hours are any-to-English translation data. This makes Whisper naturally multilingual.

A striking finding: the word error rate for a given language is highly predictable from the amount of training data in that language. The correlation coefficient is 0.83 on a log-log scale, and WER roughly halves for every 16× increase in training data. Languages with unique scripts (Chinese, Korean, Hebrew) or those distantly related to the Indo-European languages dominating the dataset are outliers with higher-than-expected WER.

For speech translation (any language → English), Whisper achieves state-of-the-art BLEU of 29.1 on CoVoST2 zero-shot, outperforming prior supervised work on low- and mid-resource languages thanks to its 68,000 hours of translation data — vastly more than the 861 hours in CoVoST2's training set.

Open in Lab
Each dot is a language. Hover to see the language name and WER. Notice the strong log-log linear trend — more data means lower error, predictably.
The demo wakes as you arrive…

Scaling properties

Two scaling axes matter: model size and dataset size.

Model scaling: Performance improves reliably with model size across all tasks except English speech recognition, which shows diminishing returns — likely because the largest models are approaching human-level performance (a ceiling effect).

Dataset scaling: Training on just 0.5% of the data (3,400 hours) yields 30.5% English WER. Scaling to the full 680,000 hours drops this to 9.9% — and every intermediate step shows improvement. Multilingual performance follows a power-law trend up to 54,000 hours then shows diminishing returns, suggesting the largest models may be under-trained relative to dataset size.

Multitask transfer: For small models, joint multilingual-multitask training hurts English performance (negative transfer). But at larger scales, joint training helps — languages and tasks share useful representations, and the largest joint models outperform English-only models even when adjusting for compute.

Long-form transcription and decoding strategies

Whisper processes 30-second chunks, but real-world audio can last hours. The solution is buffered transcription: transcribe each 30-second window, then shift forward based on predicted timestamps. Several heuristics prevent failure modes:

  • (5 beams) reduces repetition loops that occur in .
  • scheduling: Start at temperature 0 (deterministic). If the average log probability is below −1 or the text compression rate exceeds 2.4 (signs of or repetition), increase temperature by 0.2 up to 1.0.
  • Previous-text conditioning: Feed the preceding transcript as context when temperature is below 0.5, improving coherence across segments.
  • Voice activity detection: Combine the <|nospeech|> token probability (threshold 0.6) with the average log-probability threshold (−1) for reliable silence detection.
  • Initial timestamp constraint: Force the first timestamp between 0.0 and 1.0 seconds to prevent the model from skipping the beginning of segments.

On seven long-form datasets — TED talks, late-night TV, podcasts, earnings calls — Whisper outperforms the best open-source model and most commercial ASR services.

Limitations and future directions

Despite its strengths, Whisper has several acknowledged limitations:

  • Decoding failures. Seq2seq models can hallucinate, get stuck in repetition loops, or skip words at segment boundaries. These are non-perceptual errors that don't decrease smoothly with scale.
  • Low-resource languages. Performance on many languages remains poor due to English-heavy training data. The clear log-linear scaling trend suggests a targeted data collection effort could dramatically improve these.
  • No fine-tuning study. The paper focuses on zero-shot evaluation; fine-tuning on domain-specific data would likely improve results further but was not systematically studied.
  • Language model effects unclear. It's unknown how much of Whisper's robustness comes from its encoder vs. its decoder (which acts as a language model). Ablating these components could clarify the contribution of each.

Why it mattered

  1. 2020

    Wav2Vec 2.0

    Self-supervised pre-training on raw audio. Learns speech representations without labels but requires fine-tuning for any downstream task — setting the stage for Whisper's critique of the fine-tuning paradigm.

  2. 2021

    BigSSL — 1M hours of unsupervised audio

    Scaled unsupervised speech pre-training to 1,000,000 hours. Showed the value of scale but still required fine-tuning — the decoder remained the bottleneck.

  3. 2021

    SpeechStew — 5,140 hours of mixed supervision

    Mixed seven high-quality supervised datasets to improve robustness. Showed that multi-domain supervised training helps, but at a far smaller scale than Whisper.

  4. 2022

    Whisper — 680,000 hours of weak supervision

    Closed the gap by scaling weakly supervised data by an order of magnitude. Zero-shot performance approaching human accuracy without any fine-tuning. Open-sourced.

  5. 2023

    Whisper Large V3 — 5M hours

    Trained on 1M hours of labeled data + 4M hours of pseudo-labeled data, with 128 Mel bins instead of 80. Achieved 10–20% error reduction over V2 across languages.

The idea in code

Whisper inference pipeline — simplifiedpython

Simplified to show the idea — not the real implementation.

import whisper
import numpy as np

# Load the model — no fine-tuning needed
model = whisper.load_model("large-v2")

# Transcribe: zero-shot, any language
result = model.transcribe("audio.mp3")
print(result["text"])

# Under the hood:
# 1. Audio → 16 kHz → 80-channel log-Mel spectrogram (30s chunks)
# 2. Spectrogram → Transformer encoder → audio features
# 3. Special tokens → Transformer decoder → predicted text
#    [SOT] [EN] [TRANSCRIBE] [NOTIMESTAMPS] → "The quick brown..."
#
# The decoder is an audio-conditional language model:
# P(token_t | audio, token_1, ..., token_{t-1})
#
# That's the entire pipeline. No separate language detector,
# no voice activity detector, no inverse text normalizer.
# One model. One forward pass per 30-second chunk.

CitationRadford, Kim, Xu, Brockman, McLeavey, Sutskever. Robust Speech Recognition via Large-Scale Weak Supervision. ICML, 2023.

Terms in this paper