Language Models2020intermediate11 min read

ELECTRA: Pre-Training Text Encoders as Discriminators Rather Than Generators

ELECTRA: تدريب مُرمِّزات النصوص مسبقاً بالتمييز بدلاً من التوليد

Clark, K. · Luong, M.-T. · Le, Q. V. · Manning, C. D. — ICLR

The problem

, as used in BERT, replaces 15% of input tokens with [MASK] and trains the model to reconstruct them. This works well but is deeply inefficient: 85% of tokens contribute nothing to the training signal. Every forward pass processes the full sequence, yet the only learns from the masked subset. To match state-of-the-art results, BERT-style models require enormous compute — hundreds of GPU-days. Can we design a task that extracts a learning signal from every token, not just the masked ones?

The contribution

ELECTRA introduces "" (RTD), a new pre-training objective. A small network (a small MLM) proposes plausible replacement tokens. A larger network then classifies every token as "original" or "replaced." Because the discriminator receives a training signal from all input positions — not just 15% — it learns far more efficiently. ELECTRA-Small, trainable on a single GPU in 4 days, outperforms GPT on GLUE. ELECTRA-Base matches BERT-Large with 1/4 the compute. ELECTRA-Large sets a new state-of-the-art on 2.0.

The impact

ELECTRA proved that the pre-training task matters as much as scale. By replacing BERT's generative objective with a discriminative one, it achieved better results with far less compute — democratizing pre-training for teams without massive GPU clusters. The generator-discriminator framework influenced DeBERTa and subsequent efficient pre-training methods. ELECTRA showed that the -inspired idea of "detect fakes" transfers powerfully to NLP, opening a new axis of research beyond masked language modeling.

Imagine a language exam with two formats. In the first (BERT's approach), the examiner whites out 15% of the words in a paragraph and asks: "Fill in the blanks." The student works hard on those blanks but completely ignores the other 85% — those words are freebies.

ELECTRA redesigns the exam. A mischievous assistant first replaces some words with plausible-but-wrong alternatives — "The chef ate a delicious meal" instead of "The chef cooked a delicious meal." Now the student must play detective: read every single word and mark each one as genuine or tampered. No word gets a free pass. The result? The student learns far more from the same paragraph.

The inefficiency problem: learning from only 15% of the input

BERT introduced masked language modeling (MLM): randomly mask 15% of tokens, then predict the original token at each masked position. This was a breakthrough — it enabled deep pre-training for the first time.

But there is a structural inefficiency hiding in plain sight. Consider a sentence with 100 tokens. BERT masks 15 of them. The model processes all 100 tokens through the full , but the loss function only computes gradients for those 15 masked positions. The remaining 85 tokens pass through the network, consume compute, yet generate zero training signal. It's as if a student reads an entire textbook but is only tested on every seventh page.

The consequence is clear: to learn the same amount, BERT needs far more data passes — and therefore far more compute. RoBERTa showed that BERT was undertrained; the solution was simply more data and more compute. But this is treating the symptom, not the cause.

Open in Lab
Compare how BERT (MLM) and ELECTRA (RTD) extract training signal from the same sentence. Colored tokens = learning signal.
The demo wakes as you arrive…

The idea: detect fakes instead of filling blanks

ELECTRA's insight is to replace the generative MLM task with a discriminative one called replaced token detection (RTD). The setup uses two networks that work together:

The Generator is a small masked . It takes the original input, masks 15% of tokens (just like BERT), and predicts replacement tokens for each masked position. These replacements are plausible — they come from the generator's learned vocabulary distribution, not random noise. If the original word was "cooked," the generator might propose "ate" or "prepared" — words that fit grammatically but aren't the original.

The Discriminator is the main model — this is what becomes the pre-trained . It receives the corrupted sequence (original tokens with generator substitutions spliced in) and must classify every single token as "original" or "replaced." This is a binary classification at each position — a simple over each token's representation.

The key difference: BERT computes loss on 15% of tokens. ELECTRA computes loss on 100% of tokens. Every position either confirms "yes, this is real" or flags "no, this was swapped." The discriminator gets a training signal from everywhere.

Open in Lab
Step through the ELECTRA pipeline: the generator proposes replacements, then the discriminator judges every token.
The demo wakes as you arrive…

The training objective: two losses, one forward pass

ELECTRA trains the generator and discriminator jointly. The generator uses the standard MLM cross-entropy loss on masked positions — identical to BERT. The discriminator uses a loss on all positions, predicting whether each token is original or replaced.

The intuition for combining these two losses: the generator's job is to make the task hard enough for the discriminator. If replacements were random noise ("the" → "xkq"), the discriminator's job would be trivially easy and it would learn nothing useful. By training a generator that produces plausible replacements, the discriminator faces a genuinely challenging task — distinguishing "ate" from "cooked" requires real language understanding.

LGen=∑t∈M−log⁡ pG(xt∣x\M)\mathcal{L}_{\text{Gen}} = \sum_{t \in \mathcal{M}} -\log\, p_G(x_t \mid \mathbf{x}_{\backslash \mathcal{M}})
Generator loss — standard MLM cross-entropy on masked positions — The generator is trained exactly like BERT: given a sentence with some tokens masked, it predicts the original token at each masked position. The set M contains the masked positions. This loss trains the generator to produce realistic replacement tokens.
LDisc=∑t=1n−[1(x~t=xt)log⁡D(x~,t)+1(x~t≠xt)log⁡(1−D(x~,t))]\mathcal{L}_{\text{Disc}} = \sum_{t=1}^{n} -\Big[\mathbb{1}(\tilde{x}_t = x_t)\log D(\mathbf{\tilde{x}}, t) + \mathbb{1}(\tilde{x}_t \neq x_t)\log(1 - D(\mathbf{\tilde{x}}, t))\Big]
Discriminator loss — binary cross-entropy on ALL positions — The discriminator examines *every* token in the corrupted sequence and outputs a probability (via sigmoid) that the token is the original. The loss is standard binary cross-entropy. Crucially, this sums over all n positions — not just the 15% that were masked. This is why ELECTRA is more sample-efficient than BERT.
LELECTRA=LGen+λ LDisc\mathcal{L}_{\text{ELECTRA}} = \mathcal{L}_{\text{Gen}} + \lambda\,\mathcal{L}_{\text{Disc}}
Combined ELECTRA loss — The total loss is the sum of both objectives weighted by λ (typically 50). The large weight on the discriminator loss reflects that the discriminator is the primary model we care about. After pre-training, the generator is discarded and only the discriminator is used for fine-tuning.

Architecture: a small generator, a large discriminator

Both the generator and discriminator are Transformer encoders — the same architecture as BERT. The critical design choice is their relative size. The authors found that the generator should be substantially smaller than the discriminator — typically 1/4 to 1/3 the size. Why?

If the generator is as large and powerful as the discriminator, it generates near-perfect replacements. The discriminator's task becomes too easy (most replacements match the original), and it learns less. A smaller generator makes enough mistakes that the discriminator faces a genuinely challenging task.

ELECTRA model sizes:

  • ELECTRA-Small: Discriminator has 12 layers, 256 hidden dim, 4 heads (14M params). Generator is 12 layers but with a smaller hidden size.
  • ELECTRA-Base: Discriminator matches BERT-Base (12 layers, 768 hidden dim, 12 heads, 110M params). Generator is 12 layers with 256 hidden dim — 1/3 the discriminator's width.
  • ELECTRA-Large: Discriminator matches BERT-Large (24 layers, 1024 hidden dim, 16 heads, 330M params). Generator is 1/4 the discriminator's size.

A key technique: the generator and discriminator share their weights and weights. This is efficient — the embeddings are a large fraction of the generator's parameters — and it allows the generator's learned vocabulary knowledge to flow into the discriminator.

Open in Lab
Explore the ELECTRA architecture — click any component to learn its role.
The demo wakes as you arrive…

The idea in code

ELECTRA replaced token detection — the core training looppython

Simplified to show the idea — not the real implementation.

import numpy as np

def electra_pretrain_step(tokens, generator, discriminator, mask_prob=0.15):
    """One ELECTRA pre-training step: generator proposes, discriminator judges."""
    n = len(tokens)

    # Step 1: Mask 15% of positions (same as BERT)
    mask = np.random.random(n) < mask_prob
    masked_input = tokens.copy()
    masked_input[mask] = MASK_TOKEN_ID

    # Step 2: Generator predicts replacements for masked positions
    gen_logits = generator(masked_input)           # (n, vocab_size)
    replacements = sample_from_softmax(gen_logits)  # sample, not argmax

    # Step 3: Build corrupted input — swap masked tokens with generator samples
    corrupted = tokens.copy()
    corrupted[mask] = replacements[mask]

    # Step 4: Discriminator classifies EVERY token: original or replaced?
    disc_logits = discriminator(corrupted)          # (n,) — one sigmoid per token
    labels = (corrupted != tokens).astype(float)    # 1 = replaced, 0 = original

    # Step 5: Compute losses
    gen_loss = cross_entropy(gen_logits[mask], tokens[mask])      # MLM on 15%
    disc_loss = binary_cross_entropy(disc_logits, labels)         # BCE on 100%

    total_loss = gen_loss + 50 * disc_loss  # λ = 50

    return total_loss
    # After pre-training: discard generator, fine-tune discriminator only.

Design decisions: what the authors tried and why

The ELECTRA paper is notable for its thorough . The authors systematically tested alternative designs and reported what didn't work — insights that are just as valuable as the final method:

Adversarial training (GAN-style): Training the generator to fool the discriminator performed worse. The problem is that discrete tokens block flow, and alternatives (like REINFORCE) introduced too much variance. Maximum likelihood training for the generator was simpler and better.

Generator size: A generator the same size as the discriminator hurt performance. Intuitively, a too-good generator creates replacements so realistic that the discriminator cannot learn from them. The sweet spot was 1/4 to 1/3 of the discriminator's hidden dimension.

ELECTRA 15% (loss on masked positions only): To isolate why ELECTRA works better, the authors tried computing the discriminator loss only on the 15% of positions that were masked — same positions as BERT. Performance dropped from 85.0 to 82.4 on GLUE, confirming that the benefit comes specifically from computing loss on all tokens.

Replacing [MASK] with generated tokens: The authors tried replacing [MASK] tokens in BERT with generator samples but keeping the MLM objective. This helped slightly but didn't match ELECTRA's performance, confirming that the discriminative (binary classification) task itself is the key innovation.

Open in Lab
Explore the design decisions: see how each alternative affects GLUE performance.
The demo wakes as you arrive…

Results: more with less

ELECTRA's results tell a compelling efficiency story across three model sizes:

ELECTRA-Small (14M params): Trained on a single GPU in 4 days, it achieved 79.9 on GLUE dev — outperforming GPT (117M params) which scored 78.8. A model 8× smaller, trained with a fraction of the compute, beat a much larger model. This demonstrated that the pre-training task matters more than raw scale.

ELECTRA-Base (110M params): Matched BERT-Large (340M params, GLUE 84.0) at 85.1 on GLUE dev, using only 1/4 of BERT-Large's compute. Same architecture size as BERT-Base, but the replaced token detection task pushed it beyond a model 3× its size.

ELECTRA-Large (330M params): With the same compute as RoBERTa, achieved 88.5 on GLUE dev (vs. RoBERTa's 88.5) and set a new state-of-the-art of 88.7 EM / 90.6 F1 on SQuAD 2.0 — surpassing both RoBERTa and XLNet.

The consistent pattern: at every scale, ELECTRA matches or beats models that use significantly more compute or parameters. The training task is the multiplier.

Open in Lab
GLUE performance vs. compute (FLOPs). ELECTRA consistently achieves more with less.
The demo wakes as you arrive…

From BERT to ELECTRA: the efficiency frontier

  1. 2018

    BERT — the MLM paradigm

    Proved that deep bidirectional pre-training works. But MLM only learns from 15% of tokens, requiring massive compute to converge.

  2. 2019

    RoBERTa — more data, more compute

    Showed BERT was undertrained. Fixed it by training longer on more data. Same task, same inefficiency — just more brute force.

  3. 2020

    ELECTRA — efficient pre-training via RTD

    Replaced the generative MLM task with discriminative replaced token detection. Every token contributes to learning. Matched RoBERTa with 1/4 the compute.

  4. 2021

    DeBERTa — disentangled attention builds on ELECTRA

    Combined disentangled attention with ELECTRA-style training. Showed that the replaced token detection task composes well with architectural improvements.

ELECTRA's lasting contribution is philosophical as much as technical. It challenged the assumption that pre-training must be generative — that models must learn to produce language to understand it. Instead, ELECTRA showed that learning to judge language can be even more efficient. The detective doesn't need to write novels to spot a forgery.

This insight extends beyond NLP. In vision, discriminative pre-training objectives have shown similar efficiency gains. The principle is general: when you can design a task that extracts signal from every input element (not just a masked subset), you get more learning per compute dollar.

CitationClark, Luong, Le, Manning. ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators. ICLR, 2020.

Terms in this paper