NLP Evaluation2002beginner9 min read

BLEU: A Method for Automatic Evaluation of Machine Translation

BLEU: طريقة للتقييم التلقائي للترجمة الآلية

Papineni, K. · Roukos, S. · Ward, T. · Zhu, W.-J. — ACL

The problem

Before 2002, evaluating required expensive human judges. Every time a researcher changed a , they had to wait days or weeks for bilingual experts to rate the output. This made rapid experimentation impossible — you couldn't try 50 model tweaks in a weekend because each one needed . The field needed a fast, cheap, automatic metric that correlated with human judgment.

The contribution

BLEU (Bilingual Evaluation Understudy): an automatic metric that scores machine translation by counting matches between the candidate translation and one or more reference translations. It uses modified (capping each n-gram's count at its maximum occurrence in any reference), combines 1-gram through 4-gram precisions via geometric mean, and applies a brevity penalty to discourage short translations. BLEU achieved a 0.99 Pearson with human judgments on a level.

The impact

BLEU became the standard metric for machine translation for over a decade and remains widely used today. It accelerated the MT research cycle from months to hours, enabling the rapid iteration that produced statistical MT and later neural MT. Its ideas — n-gram overlap, modified precision, brevity penalty — influenced every subsequent metric (, ROUGE, CIDEr, BERTScore). BLEU showed that cheap automatic evaluation can substitute for expensive human evaluation at the system level.

Imagine you're a teacher grading a student's translation of a French passage. You have three answer keys from different professional translators. You don't read for style — you lay the student's paper next to each key and count matching phrases: first single words, then pairs, then triplets, then four-word chunks.

More matching phrases = better translation. But there's a catch: if a lazy student writes just "the" — a word that appears in every answer key — they'd score 100% on single-word matches. So you cap each word's count at how many times it actually appears in the best answer key.

One more rule: if the student's answer is suspiciously short (they only translated half the passage), you apply a length penalty. That's BLEU in a nutshell: capped phrase matching + a shortness penalty.

The bottleneck: human evaluation is slow and expensive

In the early 2000s, the only trusted way to evaluate machine translation was to hire bilingual human judges. Each evaluation round required assembling a panel, waiting for them to read and rate hundreds of translations, then aggregating the scores. This process had three fatal flaws:

  • Cost. A single evaluation campaign could cost thousands of dollars and take weeks to complete.
  • Inconsistency. Different judges applied different standards. Even the same judge might rate the same translation differently on different days.
  • Speed. Researchers couldn't iterate quickly. Every model change required a new round of human evaluation, turning the research cycle into a crawl.

The field was stuck: you couldn't improve what you couldn't measure fast enough.

Open in Lab
Human evaluation takes days per experiment. BLEU runs in seconds, enabling rapid iteration.
The demo wakes as you arrive…

The core idea: counting matching word sequences

BLEU's insight is deceptively simple: a good translation should share many short phrases with a professional human translation. If your system outputs "The cat sat on the mat" and the reference is "The cat is on the mat", they share many of the same unigrams (single words), bigrams (word pairs), and so on.

An n-gram is simply a contiguous sequence of n words. "The cat" is a 2-gram (bigram), "cat sat on" is a 3-gram (trigram). BLEU computes the precision of n-grams: what fraction of the n-grams in the candidate translation also appear in the reference?

Open in Lab
Toggle n-gram sizes to see which phrases the candidate shares with the reference.
The demo wakes as you arrive…

The cheating problem: why naive precision fails

Naive precision has a fatal flaw. Consider this degenerate candidate:

Candidate: the the the the the the the

Reference 1: The cat is on the mat. Reference 2: There is a cat on the mat.

The word "the" appears in both references, so every word in the candidate matches. Naive unigram precision = 7/7 = 100%. But this is clearly a terrible translation!

The paper calls this the "gaming" problem: a system could achieve perfect precision by repeating a single common word. The solution is modified precision: cap each n-gram's count in the candidate at its maximum count in any single reference. "the" appears at most twice in Reference 2, so the modified count is capped at 2, giving a modified precision of 2/7 ≈ 28.6%.

Open in Lab
Compare naive vs modified precision. Try the "gaming" candidate to see how capping prevents cheating.
The demo wakes as you arrive…
pn=∑C∈Candidates∑n-gram∈CCountclip(n-gram)∑C′∈Candidates∑n-gram′∈C′Count(n-gram′)p_n = \frac{\displaystyle\sum_{C \in \text{Candidates}} \sum_{\text{n-gram} \in C} \text{Count}_{\text{clip}}(\text{n-gram})} {\displaystyle\sum_{C' \in \text{Candidates}} \sum_{\text{n-gram}' \in C'} \text{Count}(\text{n-gram}')}
Modified n-gram precision — Count_clip caps each n-gram's count at its maximum occurrence in any single reference. The numerator sums these clipped counts over all candidate sentences; the denominator sums the total n-gram counts. This prevents gaming via repetition.

The brevity penalty: punishing short translations

Modified precision solves the repetition problem, but there's another trick a system could play: output only the words it's most confident about and omit everything else. A very short but highly precise translation would score well on precision alone.

Unlike (which measures how much of the reference you captured), precision doesn't penalize omissions. The paper's elegant solution is the brevity penalty (BP): if the candidate translation is shorter than the reference, BLEU multiplies the score by an exponentially decaying factor. The shorter your translation compared to the reference, the steeper the penalty.

Think of it as a completeness check: "Did you translate the whole passage, or just the easy parts?"

BP={1if c>re(1−r/c)if c≤rBP = \begin{cases} 1 & \text{if } c > r \\ e^{(1 - r/c)} & \text{if } c \le r \end{cases}
Brevity Penalty — c = candidate length (total words), r = effective reference length (closest reference length). If the candidate is at least as long as the reference, no penalty (BP = 1). If shorter, the penalty decays exponentially — halving the candidate length roughly halves the score.
Open in Lab
Drag the candidate length to see how the brevity penalty kicks in below the reference length.
The demo wakes as you arrive…

The full formula: combining everything

Now we can assemble the complete . The key insight is that different n-gram sizes measure different qualities: unigrams (n=1) capture word choice (adequacy), while higher n-grams capture fluency and grammatical correctness. "cat the on sat mat" has the same unigrams as "the cat sat on the mat" but terrible bigrams.

BLEU combines all n-gram precisions using a geometric mean — which means if any single n-gram precision is zero, the entire score collapses to zero. In practice, BLEU uses n-grams up to N=4 with uniform weights (w_n = 1/4 each).

BLEU=BP⋅exp⁡ ⁣(∑n=1Nwnlog⁡pn)\text{BLEU} = BP \cdot \exp\!\left(\sum_{n=1}^{N} w_n \log p_n\right)
The BLEU score — the complete formula — BP = brevity penalty · p_n = modified n-gram precision for each n · w_n = weights (typically 1/4 each for n=1…4) · The exp of sum-of-logs is equivalent to the geometric mean of the precisions, scaled by the brevity penalty.
Open in Lab
Enter candidate and reference sentences to see BLEU computed step by step.
The demo wakes as you arrive…

The same idea in code

BLEU score from scratchpython

Simplified to show the idea — not the real implementation.

import math
from collections import Counter

def ngrams(tokens, n):
    """Extract all n-grams from a list of tokens."""
    return [tuple(tokens[i:i+n]) for i in range(len(tokens)-n+1)]

def modified_precision(candidate, references, n):
    """Clipped n-gram precision: cap each n-gram count at its max in any reference."""
    cand_ngrams = Counter(ngrams(candidate, n))
    max_ref = Counter()
    for ref in references:
        ref_ngrams = Counter(ngrams(ref, n))
        for ng in ref_ngrams:
            max_ref[ng] = max(max_ref[ng], ref_ngrams[ng])
    clipped = {ng: min(count, max_ref[ng]) for ng, count in cand_ngrams.items()}
    return sum(clipped.values()), max(sum(cand_ngrams.values()), 1)

def bleu(candidate, references, max_n=4):
    """Compute corpus-level BLEU (simplified for one sentence)."""
    c = len(candidate)
    r = min((len(ref) for ref in references), key=lambda l: abs(l - c))
    bp = 1.0 if c >= r else math.exp(1 - r / c)      # brevity penalty

    log_avg = 0.0
    for n in range(1, max_n + 1):
        num, den = modified_precision(candidate, references, n)
        if num == 0:
            return 0.0                                 # geometric mean → 0
        log_avg += (1 / max_n) * math.log(num / den)

    return bp * math.exp(log_avg)

# Example:
cand = "the cat sat on the mat".split()
refs = ["the cat is on the mat".split(),
        "there is a cat on the mat".split()]
print(f"BLEU = {bleu(cand, refs):.4f}")               # ≈ 0.4868

Does it actually work? Correlation with human judges

The paper's strongest claim is empirical: BLEU scores correlate with human judgments at r = 0.99 (Pearson) when averaged over a sufficiently large test corpus. The authors compared five Chinese-to-English MT systems, rating each with both human judges and BLEU. The system rankings matched almost perfectly.

There's an important caveat: BLEU correlates well at the corpus level (hundreds or thousands of sentences) but poorly at the sentence level. A single sentence can get a misleading BLEU score because the sample size is too small for n-gram statistics to stabilize. This is why BLEU is always reported as an aggregate metric over a test set, never as a score for individual translations.

Open in Lab
Each dot is an MT system. Notice how tightly BLEU tracks the human ranking.
The demo wakes as you arrive…

Limitations: what BLEU misses

BLEU opened the door to automatic evaluation, but it has well-known blind spots:

  • Synonyms are invisible. If the reference says "automobile" and the candidate says "car", BLEU counts zero matches — even though the meaning is identical. BLEU only sees exact surface-form matches.
  • Word order doesn't matter enough. BLEU treats n-grams as a bag: "the dog bit the man" and "the man bit the dog" share all the same bigrams. Higher n-grams help, but only partially.
  • No semantic understanding. BLEU can't tell if a translation is factually wrong, inappropriate, or missing critical nuance. Two translations with the same BLEU score can differ dramatically in meaning.
  • Sentence-level unreliability. On individual sentences, BLEU is noisy and unreliable. A valid paraphrase of the reference can score near zero.

These limitations motivated a generation of successors: METEOR (which handles synonyms via ), ROUGE (which focuses on recall for ), TER (which counts edits), and modern learned metrics like BERTScore and COMET.

The legacy: an evaluation revolution

BLEU's deepest contribution wasn't the specific formula — it was the idea that automatic evaluation of generated text is possible and useful. Before BLEU, the notion of scoring translations by counting word overlaps seemed too crude. After BLEU, every text generation field — summarization, dialogue, image captioning — adopted its own automatic metrics.

The Google Neural MT system that revolutionized translation in 2016 was benchmarked with BLEU. So were the -based models that came before it, and the models that came after. BLEU gave researchers a common yardstick, enabling the rapid progress from statistical MT to neural MT to large language models.

  1. 2002

    BLEU published

    Papineni et al. introduce BLEU at ACL 2002. For the first time, MT researchers can evaluate systems in seconds instead of weeks.

  2. 2004

    ROUGE for summarization

    Lin adapts the n-gram overlap idea for text summarization, focusing on recall instead of precision. ROUGE becomes the standard metric for summarization tasks.

  3. 2005

    METEOR adds synonyms

    Banerjee and Lavie create METEOR, which goes beyond exact matching to handle synonyms, stems, and paraphrases using WordNet.

  4. 2014

    Neural MT benchmarked with BLEU

    Sutskever et al. and Bahdanau et al. benchmark sequence-to-sequence and attention-based neural MT using BLEU, establishing it as the neural era's yardstick.

  5. 2016

    Google NMT achieves near-human BLEU

    Google's Neural Machine Translation system narrows the gap with human translators as measured by BLEU, marking a milestone for production MT.

  6. 2020

    BERTScore and learned metrics

    Learned metrics like BERTScore use contextual embeddings to compare meaning rather than surface forms, addressing BLEU's synonym blindness.

CitationPapineni, Roukos, Ward, Zhu. Bleu: a Method for Automatic Evaluation of Machine Translation. ACL, 2002.

Terms in this paper

  • BLEU Scoreمعيار بلو الإحصائي لتقييم الترجمة
  • N-gramالسلاسل الرمزية المتجاورة
  • Machine Translationالترجمة الآلية
  • Evaluation Metricمعيار قياس الأداء
  • Precisionمدى المحركية (الدقة المحددة لفئة)
  • Recallالقدرة على الاستدعاء الشامل
  • Smoothingالتمهيد الاحتمالي
  • ROUGE Scoreمعيار روج لتقييم التلخيص
  • METEORمعيار ميتيور المطور