Language Models2018intermediate10 min read
Deep Contextualized Word Representations
التمثيلات السياقية العميقة للكلمات
Peters, M. E. · Neumann, M. · Iyyer, M. · Gardner, M. · Clark, C. · Lee, K. · Zettlemoyer, L. — NAACL
The problem
Before ELMo, word embeddings were static: Word2Vec and GloVe assigned each word a single fixed learned from co-occurrence statistics. The word "play" got the same vector whether it meant a theatrical performance, a sports action, or a children's activity. These frozen representations forced downstream models to do all the disambiguation work themselves, with limited context and limited capacity. in NLP was shallow — you could initialize with pre-trained embeddings, but those embeddings carried no sentence-level understanding.
The contribution
ELMo (Embeddings from Language Models): deep contextualized word representations derived from the internal states of a (biLM). A builds context-free embeddings, then a 2- biLSTM reads the full sentence in both directions. Each layer captures different linguistic properties — syntax in lower layers, semantics in upper layers. The final ELMo vector is a task-specific learned linear combination of all layers, allowing downstream models to mix the right level of abstraction. Simply adding ELMo to existing models produced new state-of-the-art results on six NLP benchmarks with relative error reductions of 6–25%.
The impact
ELMo proved that pre-trained internals are a rich, transferable source of linguistic knowledge — a paradigm shift from static embeddings. It directly inspired BERT (which replaced the biLSTM with a and used masking instead of language modeling) and ULMFiT (which explored strategies). ELMo established that different layers encode different linguistic levels, an insight that persists in modern Transformer analysis. It was the bridge between the Word2Vec era and the pre-train–fine-tune revolution.
Think of GloVe as a passport photo: it captures your face, but it looks the same whether you're laughing at a party or concentrating at work. The photo is you, but it misses everything about the moment.
ELMo is a live video feed: it sees your face in context — your expression shifts based on who you're talking to, what you just heard, and what you're about to say. Two frames of "you" in different conversations look genuinely different.
And ELMo doesn't just film from one angle. It runs two cameras simultaneously — one filming from the beginning of the sentence forward, one from the end backward — and blends both feeds so every word gets the full picture.
The problem: one vector per word is not enough
GloVe and Word2Vec were breakthroughs: they turned words into vectors where similar words land nearby. But every word gets exactly one vector, no matter where it appears. Consider:
- "He sat by the river bank" — a muddy slope
- "She went to the bank to deposit money" — a financial institution
- "You can bank on me" — to rely on someone
GloVe mashes all three meanings into a single point in vector space. The downstream model is left to figure out which meaning applies, starting from this blurred average. For rare senses, the signal is almost entirely lost.
The architecture: a bidirectional language model with deep layers
ELMo builds word representations in three stages, each adding richer information:
Stage 1: Character-level CNN. Instead of looking up a word in a fixed vocabulary, ELMo reads the word's characters through a with 2048 filters, followed by two highway layers and a projection down to 512 dimensions. This means ELMo can handle any word — even misspellings and rare terms — because it builds representations from characters up, much like how you can sound out an unfamiliar word letter by letter.
Stage 2: Two-layer biLSTM. The character-based embeddings flow into a two-layer bidirectional , each layer with 4096 units projected down to 512 dimensions, connected by a . The forward LSTM reads left-to-right, predicting each next word. The backward LSTM reads right-to-left, predicting each previous word. Together, they produce representations that absorb the full sentence context from both directions.
Stage 3: Layer mixing. Each layer captures something different. The bottom (character CNN) provides morphological features. Layer 1 captures syntactic patterns. Layer 2 captures semantic meaning. Rather than using just the top layer, ELMo takes a learned weighted average of all three — and the weights are learned separately for each downstream task.
The training objective: predict the next word, both ways
A forward language model reads a sentence left to right and, at each position, tries to predict the next word given everything before it. A backward language model does the same in reverse — predicting each previous word from everything after it. ELMo trains both directions jointly, sharing the character CNN and the softmax layer while keeping the LSTM parameters separate for each direction. The objective is to maximize the combined log-:
Think of it this way: the forward LSTM has read "The cat sat on the" and is trying to predict "mat." The backward LSTM has read "mat the on sat" (reversed) and is trying to predict "cat." Each direction builds a different kind of anticipation, and together they give every word a view of the complete sentence.
The key innovation: mixing layers for each task
Previous work typically used only the top layer of a pre-trained network. ELMo's insight is that every layer matters, and different tasks need different layers. Syntax-heavy tasks like POS tagging benefit from lower layers; semantic tasks like benefit from higher layers. Rather than choosing one layer, ELMo lets each task learn its own recipe:
Using ELMo: plug it into any model
One of ELMo's most practical virtues is how easy it is to integrate. You don't need to redesign your model. You just:
- Run your sentence through the pre-trained biLM (weights frozen).
- Compute the weighted layer combination to get an ELMo vector for each token.
- Concatenate the ELMo vector with your existing word : .
- Feed this enriched into your model as usual.
For some tasks (like SQuAD and SNLI), the authors found that also adding ELMo vectors at the output of the task model's RNN — not just the input — gave further improvements. This works especially well when the task model uses an mechanism, because the attention can directly query the biLM's internal representations.
What does each layer learn?
The paper includes a striking analysis. They tested what each biLSTM layer knows independently:
- Layer 1 (lower) is better at syntax. When used alone for POS tagging, the first layer outperforms the second. It captures grammatical structure — noun vs. verb, subject vs. object.
- Layer 2 (upper) is better at semantics. When used alone for word sense disambiguation (WSD), the second layer achieves F1 of 69.0, competitive with supervised WSD systems that use hand-crafted features. It captures meaning in context — which sense of "bank" is active.
This separation is intuitive: you need to parse the grammar before you can understand meaning, so lower layers learn structure and upper layers build on that structure to capture semantics. It's like how a human reader first recognizes that "bank" is a noun in this sentence (syntax) and then decides it means a riverbank (semantics).
The same idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
def elmo_representation(char_cnn_out, bilstm_layer1, bilstm_layer2,
task_weights, task_gamma):
"""
Compute ELMo vector for one token.
char_cnn_out: (d,) — context-free character-level embedding
bilstm_layer1: (d,) — biLSTM layer 1 output (syntax-rich)
bilstm_layer2: (d,) — biLSTM layer 2 output (semantics-rich)
task_weights: (3,) — softmax-normalized, learned per task
task_gamma: scalar — learned per task, scales final vector
"""
# Stack all layers: [char CNN, biLSTM L1, biLSTM L2]
layers = np.stack([char_cnn_out, bilstm_layer1, bilstm_layer2])
# Weighted sum — each task learns its own recipe
elmo_vec = task_gamma * np.sum(task_weights[:, None] * layers, axis=0)
return elmo_vec # shape: (d,)
def add_elmo_to_model(word_embedding, elmo_vec):
"""Just concatenate ELMo with the existing embedding."""
return np.concatenate([word_embedding, elmo_vec])
# shape: (d_word + d_elmo,) — feed into your task modelResults: state-of-the-art on everything
The most striking aspect of ELMo's results is their breadth. Simply adding ELMo vectors to existing models — without changing architectures — set new state-of-the-art on six different NLP benchmarks simultaneously:
- SQuAD (question answering): F1 improved from 81.1 to 85.8 (+4.7, 24.9% error reduction)
- SNLI (textual entailment): accuracy from 88.0 to 88.7
- SRL (semantic role labeling): F1 from 81.4 to 84.6 (+3.2)
- Coref (coreference resolution): F1 from 67.2 to 70.4 (+3.2)
- NER (): F1 from 90.15 to 92.22 (+2.06)
- SST-5 (): accuracy from 51.4 to 54.7 (+3.3)
The improvements were especially dramatic for SQuAD — the 4.7% gain was more than double what CoVe (an earlier contextualized approach using machine translation) achieved. Even more revealing: with only 1% of the SRL training data, ELMo matched the baseline that used 10% of the data, showing that better representations dramatically reduce the need for labeled examples.
Why it mattered: the bridge to BERT
ELMo's limitation was also clear: it concatenated two separate directions rather than truly integrating them. The forward and backward LSTMs never interact during encoding — each builds its representation in isolation, then they're pasted together. This is a shallow form of bidirectionality. BERT would solve this six months later by using a Transformer that lets every token attend to every other token simultaneously — deep bidirectionality. But BERT's core insight — that pre-trained language model representations transfer powerfully — was ELMo's insight first.
2013
Word2Vec
Static word vectors from co-occurrence. "King - Man + Woman ≈ Queen" demonstrated that linear structure exists in word space, but every word is frozen to one vector.
2014
GloVe
Global co-occurrence statistics baked into vectors. Better training, same limitation: one static vector per word.
2018
ELMo — contextualized embeddings
First widely adopted contextualized representations. BiLSTM language model internals, layer mixing, state-of-the-art on 6 tasks. Proved pre-trained representations transfer.
2018
ULMFiT
Howard & Ruder showed that fine-tuning (not just feature extraction) a pre-trained LM with careful learning rate schedules works across tasks. Parallel innovation to ELMo.
2018
BERT — deep bidirectionality
Replaced the biLSTM with a Transformer encoder and the LM objective with masked language modeling. Every token sees every other token during encoding — solving ELMo's shallow concatenation. The pre-train → fine-tune paradigm became the standard.
2019
GPT-2 — scaling up
OpenAI scales the Transformer decoder to 1.5B parameters. Showed that more data + more parameters = emergent abilities. The ELMo insight (pre-trained LMs transfer) taken to its extreme.
ELMo's deepest contribution is not its architecture — biLSTMs were soon replaced by Transformers. Its contribution is the idea: that a language model trained on raw text builds internal representations rich enough to boost any NLP task, and that the right strategy is to expose all layers and let the task choose. Every pre-trained model since — BERT, GPT, T5, Claude — builds on that foundation.
CitationPeters, Neumann, Iyyer, Gardner, Clark, Lee, Zettlemoyer. Deep Contextualized Word Representations. NAACL, 2018.
Terms in this paper
- Contextual Embeddingالتضمين السياقي
- Bidirectional Language Modelالنموذج اللغوي ثنائي الاتجاه
- LSTMشبكة الذاكرة الطويلة قصيرة المدى
- Character-Level CNNشبكة التفافية على مستوى الأحرف
- Pretrainingالتدريب المسبق
- Fine-Tuningالضبط الدقيق
- Word Sense Disambiguationإزالة غموض معاني الكلمات
- Transfer Learningنقل التعلم
- Feature Extractionاستخلاص السمات
- Residual Connectionالوصلة التجاوزية
- Highway Networkالشبكة الطُرُقية