Language Models2018intermediate10 min read

Deep Contextualized Word Representations

التمثيلات السياقية العميقة للكلمات

Peters, M. E. · Neumann, M. · Iyyer, M. · Gardner, M. · Clark, C. · Lee, K. · Zettlemoyer, L. — NAACL

The problem

Before ELMo, word embeddings were static: Word2Vec and GloVe assigned each word a single fixed learned from co-occurrence statistics. The word "play" got the same vector whether it meant a theatrical performance, a sports action, or a children's activity. These frozen representations forced downstream models to do all the disambiguation work themselves, with limited context and limited capacity. in NLP was shallow — you could initialize with pre-trained embeddings, but those embeddings carried no sentence-level understanding.

The contribution

ELMo (Embeddings from Language Models): deep contextualized word representations derived from the internal states of a (biLM). A builds context-free embeddings, then a 2- biLSTM reads the full sentence in both directions. Each layer captures different linguistic properties — syntax in lower layers, semantics in upper layers. The final ELMo vector is a task-specific learned linear combination of all layers, allowing downstream models to mix the right level of abstraction. Simply adding ELMo to existing models produced new state-of-the-art results on six NLP benchmarks with relative error reductions of 6–25%.

The impact

ELMo proved that pre-trained internals are a rich, transferable source of linguistic knowledge — a paradigm shift from static embeddings. It directly inspired BERT (which replaced the biLSTM with a and used masking instead of language modeling) and ULMFiT (which explored strategies). ELMo established that different layers encode different linguistic levels, an insight that persists in modern Transformer analysis. It was the bridge between the Word2Vec era and the pre-train–fine-tune revolution.

Think of GloVe as a passport photo: it captures your face, but it looks the same whether you're laughing at a party or concentrating at work. The photo is you, but it misses everything about the moment.

ELMo is a live video feed: it sees your face in context — your expression shifts based on who you're talking to, what you just heard, and what you're about to say. Two frames of "you" in different conversations look genuinely different.

And ELMo doesn't just film from one angle. It runs two cameras simultaneously — one filming from the beginning of the sentence forward, one from the end backward — and blends both feeds so every word gets the full picture.

The problem: one vector per word is not enough

GloVe and Word2Vec were breakthroughs: they turned words into vectors where similar words land nearby. But every word gets exactly one vector, no matter where it appears. Consider:

  • "He sat by the river bank" — a muddy slope
  • "She went to the bank to deposit money" — a financial institution
  • "You can bank on me" — to rely on someone

GloVe mashes all three meanings into a single point in vector space. The downstream model is left to figure out which meaning applies, starting from this blurred average. For rare senses, the signal is almost entirely lost.

Open in Lab
Click a word with multiple meanings and see how GloVe gives it one fixed position while ELMo moves it based on context.
The demo wakes as you arrive…

The architecture: a bidirectional language model with deep layers

ELMo builds word representations in three stages, each adding richer information:

Stage 1: Character-level CNN. Instead of looking up a word in a fixed vocabulary, ELMo reads the word's characters through a with 2048 filters, followed by two highway layers and a projection down to 512 dimensions. This means ELMo can handle any word — even misspellings and rare terms — because it builds representations from characters up, much like how you can sound out an unfamiliar word letter by letter.

Stage 2: Two-layer biLSTM. The character-based embeddings flow into a two-layer bidirectional , each layer with 4096 units projected down to 512 dimensions, connected by a . The forward LSTM reads left-to-right, predicting each next word. The backward LSTM reads right-to-left, predicting each previous word. Together, they produce representations that absorb the full sentence context from both directions.

Stage 3: Layer mixing. Each layer captures something different. The bottom (character CNN) provides morphological features. Layer 1 captures syntactic patterns. Layer 2 captures semantic meaning. Rather than using just the top layer, ELMo takes a learned weighted average of all three — and the weights are learned separately for each downstream task.

Open in Lab
Click any stage to see how a word's representation gets built from characters to context-aware vectors.
The demo wakes as you arrive…

The training objective: predict the next word, both ways

A forward language model reads a sentence left to right and, at each position, tries to predict the next word given everything before it. A backward language model does the same in reverse — predicting each previous word from everything after it. ELMo trains both directions jointly, sharing the character CNN and the softmax layer while keeping the LSTM parameters separate for each direction. The objective is to maximize the combined log-:

∑k=1N(log⁡p(tk∣t1,…,tk−1; Θfwd)  +  log⁡p(tk∣tk+1,…,tN; Θbwd))\sum_{k=1}^{N}\bigl(\log p(t_k \mid t_1, \dots, t_{k-1};\, \Theta_{\text{fwd}}) \;+\; \log p(t_k \mid t_{k+1}, \dots, t_N;\, \Theta_{\text{bwd}})\bigr)
Bidirectional language model objective — The forward LM predicts each token from its left context; the backward LM predicts each token from its right context. Both share the embedding and softmax parameters.

Think of it this way: the forward LSTM has read "The cat sat on the" and is trying to predict "mat." The backward LSTM has read "mat the on sat" (reversed) and is trying to predict "cat." Each direction builds a different kind of anticipation, and together they give every word a view of the complete sentence.

The key innovation: mixing layers for each task

Previous work typically used only the top layer of a pre-trained network. ELMo's insight is that every layer matters, and different tasks need different layers. Syntax-heavy tasks like POS tagging benefit from lower layers; semantic tasks like benefit from higher layers. Rather than choosing one layer, ELMo lets each task learn its own recipe:

ELMoktask=γtask∑j=0Lsjtask hk,j\mathbf{ELMo}_k^{\text{task}} = \gamma^{\text{task}} \sum_{j=0}^{L} s_j^{\text{task}}\, \mathbf{h}_{k,j}
ELMo representation — the learned layer cocktail — h_{k,j} = representation of word k at layer j (j=0 is the char CNN, j=1,2 are biLSTM layers) · s_j = softmax-normalized task-specific weight for layer j · γ = task-specific scaling factor · The task learns which layers matter most for its needs.
Open in Lab
Drag the layer weights and see how the ELMo representation changes. Notice how syntax tasks prefer lower layers and semantic tasks prefer higher layers.
The demo wakes as you arrive…

Using ELMo: plug it into any model

One of ELMo's most practical virtues is how easy it is to integrate. You don't need to redesign your model. You just:

  1. Run your sentence through the pre-trained biLM (weights frozen).
  2. Compute the weighted layer combination to get an ELMo vector for each token.
  3. Concatenate the ELMo vector with your existing word : [xk; ELMok][\mathbf{x}_k;\, \mathbf{ELMo}_k].
  4. Feed this enriched into your model as usual.

For some tasks (like SQuAD and SNLI), the authors found that also adding ELMo vectors at the output of the task model's RNN — not just the input — gave further improvements. This works especially well when the task model uses an mechanism, because the attention can directly query the biLM's internal representations.

Open in Lab
Watch how ELMo vectors are concatenated with existing embeddings and fed into a downstream task model.
The demo wakes as you arrive…

What does each layer learn?

The paper includes a striking analysis. They tested what each biLSTM layer knows independently:

  • Layer 1 (lower) is better at syntax. When used alone for POS tagging, the first layer outperforms the second. It captures grammatical structure — noun vs. verb, subject vs. object.
  • Layer 2 (upper) is better at semantics. When used alone for word sense disambiguation (WSD), the second layer achieves F1 of 69.0, competitive with supervised WSD systems that use hand-crafted features. It captures meaning in context — which sense of "bank" is active.

This separation is intuitive: you need to parse the grammar before you can understand meaning, so lower layers learn structure and upper layers build on that structure to capture semantics. It's like how a human reader first recognizes that "bank" is a noun in this sentence (syntax) and then decides it means a riverbank (semantics).

Open in Lab
Toggle between Layer 1 and Layer 2 to see how the same words cluster differently — syntactically or semantically.
The demo wakes as you arrive…

The same idea in code

Computing ELMo representations — simplifiedpython

Simplified to show the idea — not the real implementation.

import numpy as np

def elmo_representation(char_cnn_out, bilstm_layer1, bilstm_layer2,
                        task_weights, task_gamma):
    """
    Compute ELMo vector for one token.

    char_cnn_out:    (d,) — context-free character-level embedding
    bilstm_layer1:   (d,) — biLSTM layer 1 output (syntax-rich)
    bilstm_layer2:   (d,) — biLSTM layer 2 output (semantics-rich)
    task_weights:    (3,) — softmax-normalized, learned per task
    task_gamma:      scalar — learned per task, scales final vector
    """
    # Stack all layers: [char CNN, biLSTM L1, biLSTM L2]
    layers = np.stack([char_cnn_out, bilstm_layer1, bilstm_layer2])

    # Weighted sum — each task learns its own recipe
    elmo_vec = task_gamma * np.sum(task_weights[:, None] * layers, axis=0)
    return elmo_vec  # shape: (d,)

def add_elmo_to_model(word_embedding, elmo_vec):
    """Just concatenate ELMo with the existing embedding."""
    return np.concatenate([word_embedding, elmo_vec])
    # shape: (d_word + d_elmo,) — feed into your task model

Results: state-of-the-art on everything

The most striking aspect of ELMo's results is their breadth. Simply adding ELMo vectors to existing models — without changing architectures — set new state-of-the-art on six different NLP benchmarks simultaneously:

  • SQuAD (question answering): F1 improved from 81.1 to 85.8 (+4.7, 24.9% error reduction)
  • SNLI (textual entailment): accuracy from 88.0 to 88.7
  • SRL (semantic role labeling): F1 from 81.4 to 84.6 (+3.2)
  • Coref (coreference resolution): F1 from 67.2 to 70.4 (+3.2)
  • NER (): F1 from 90.15 to 92.22 (+2.06)
  • SST-5 (): accuracy from 51.4 to 54.7 (+3.3)

The improvements were especially dramatic for SQuAD — the 4.7% gain was more than double what CoVe (an earlier contextualized approach using machine translation) achieved. Even more revealing: with only 1% of the SRL training data, ELMo matched the baseline that used 10% of the data, showing that better representations dramatically reduce the need for labeled examples.

Open in Lab
Compare ELMo's improvements across all six benchmarks.
The demo wakes as you arrive…

Why it mattered: the bridge to BERT

ELMo's limitation was also clear: it concatenated two separate directions rather than truly integrating them. The forward and backward LSTMs never interact during encoding — each builds its representation in isolation, then they're pasted together. This is a shallow form of bidirectionality. BERT would solve this six months later by using a Transformer that lets every token attend to every other token simultaneously — deep bidirectionality. But BERT's core insight — that pre-trained language model representations transfer powerfully — was ELMo's insight first.

  1. 2013

    Word2Vec

    Static word vectors from co-occurrence. "King - Man + Woman ≈ Queen" demonstrated that linear structure exists in word space, but every word is frozen to one vector.

  2. 2014

    GloVe

    Global co-occurrence statistics baked into vectors. Better training, same limitation: one static vector per word.

  3. 2018

    ELMo — contextualized embeddings

    First widely adopted contextualized representations. BiLSTM language model internals, layer mixing, state-of-the-art on 6 tasks. Proved pre-trained representations transfer.

  4. 2018

    ULMFiT

    Howard & Ruder showed that fine-tuning (not just feature extraction) a pre-trained LM with careful learning rate schedules works across tasks. Parallel innovation to ELMo.

  5. 2018

    BERT — deep bidirectionality

    Replaced the biLSTM with a Transformer encoder and the LM objective with masked language modeling. Every token sees every other token during encoding — solving ELMo's shallow concatenation. The pre-train → fine-tune paradigm became the standard.

  6. 2019

    GPT-2 — scaling up

    OpenAI scales the Transformer decoder to 1.5B parameters. Showed that more data + more parameters = emergent abilities. The ELMo insight (pre-trained LMs transfer) taken to its extreme.

ELMo's deepest contribution is not its architecture — biLSTMs were soon replaced by Transformers. Its contribution is the idea: that a language model trained on raw text builds internal representations rich enough to boost any NLP task, and that the right strategy is to expose all layers and let the task choose. Every pre-trained model since — BERT, GPT, T5, Claude — builds on that foundation.

CitationPeters, Neumann, Iyyer, Gardner, Clark, Lee, Zettlemoyer. Deep Contextualized Word Representations. NAACL, 2018.

Terms in this paper