Language Models2014intermediate12 min read

GloVe: Global Vectors for Word Representation

GloVe: متجهات شاملة لتمثيل الكلمات

Pennington, J. · Socher, R. · Manning, C. D. — EMNLP

The problem

By 2014, two families of methods existed but neither told the full story. Count-based methods like LSA built a global and factorized it — capturing corpus-wide statistics but poor at word analogies. Prediction-based methods like Word2Vec's trained on local context windows — excellent at analogies but blind to global co-occurrence patterns. No unified both views, and nobody had shown why the analogy property (king − man + woman ≈ queen) actually emerged.

The contribution

GloVe: a log-bilinear regression model that trains on the nonzero entries of a global word-word co-occurrence matrix using a weighted least-squares objective. The key insight is that ratios of co-occurrence probabilities — not raw probabilities — encode meaning. The paper derives an objective where each word gets two vectors (word and context), and their plus terms should approximate the log of their co-occurrence count. A weighting function down-weights rare pairs and caps frequent ones. The result: vectors whose differences encode semantic relations, achieving 75% on the task and state-of-the-art on word similarity and named entity recognition benchmarks.

The impact

GloVe became one of the two standard pre-trained word embeddings (alongside Word2Vec) used across all of NLP for half a decade. Its pre-trained vectors (6B, 42B, 840B tokens) are still downloaded millions of times. More fundamentally, GloVe showed that count-based and prediction-based methods are two sides of the same coin — skip-gram implicitly factorizes a matrix. This unification shaped research on embeddings that followed: FastText, ELMo, and even the layers inside Transformers owe conceptual debts to GloVe's analysis.

Imagine a library card catalogue with a card for every word and a number on each card for every other word it has ever been shelved near.

LSA photocopies the whole catalogue and squashes it with a steamroller (SVD) — you get a compact summary but the fine neighbourhood gossip is lost.

Word2Vec never looks at the catalogue. Instead it plays a guessing game: cover a word, predict its shelf-neighbours from a small window. It learns patterns street by street but never sees the city map.

GloVe reads the full catalogue and plays the guessing game. It notices that the ratio of how often "ice" vs "steam" appears near "solid" tells you more than raw counts. It compresses the catalogue into short number-lists — vectors — where arithmetic on vectors mirrors arithmetic on meaning.

The two worlds: counting vs predicting

Before GloVe, word embedding methods split into two camps.

Count-based methods (LSA, HAL) build a giant matrix where row ii, column jj holds how often word ii appears near word jj across the entire corpus. They then apply — typically SVD — to compress this matrix into dense vectors. These methods see the whole corpus at once, so they capture global statistics well. But the resulting vectors performed poorly on the word analogy task (king − man + woman ≈ queen) that had become the gold standard.

Prediction-based methods (Word2Vec's skip-gram and ) never build an explicit matrix. Instead, they slide a small window across the text and train a shallow to predict context words from a (or vice versa). These models excelled at analogies and captured fine-grained semantic regularities, but they only ever saw local windows — they had no way to leverage global co-occurrence patterns directly.

The question GloVe asked was: can we get the best of both worlds?

Open in Lab
Left: count-based methods build a full co-occurrence matrix then compress it. Right: prediction-based methods learn from local windows. GloVe bridges both.
The demo wakes as you arrive…

The key insight: ratios, not raw counts

The starting point of GloVe is a simple but powerful observation. Consider two target words, "ice" and "steam", and various context words. The raw co-occurrence probabilities P(solid∣ice)P(\text{solid}|\text{ice}) and P(solid∣steam)P(\text{solid}|\text{steam}) are noisy — "water" co-occurs with both, "fashion" co-occurs with neither, and neither of these uninformative contexts tells you anything about the difference between ice and steam.

But the ratio P(solid∣ice)/P(solid∣steam)P(\text{solid}|\text{ice}) / P(\text{solid}|\text{steam}) is large (≫ 1), because "solid" is specific to ice. Conversely, P(gas∣ice)/P(gas∣steam)P(\text{gas}|\text{ice}) / P(\text{gas}|\text{steam}) is small (≪ 1), because "gas" is specific to steam. For non-discriminative words like "water" and "fashion", the ratio is close to 1.

This ratio is what encodes meaning — and GloVe's objective is designed so that differences capture exactly these ratios.

Open in Lab
Explore how the ratio of co-occurrence probabilities separates discriminative context words from non-discriminative ones.
The demo wakes as you arrive…

From ratios to vectors: deriving the objective

GloVe's derivation starts from a requirement: we want a function FF of word vectors that encodes the co-occurrence ratio. Specifically, for words ii, jj, and context word kk:

F(wi−wj,w~k)=PikPjkF(w_i - w_j, \tilde{w}_k) = \frac{P_{ik}}{P_{jk}}

The difference wi−wjw_i - w_j captures the contrast between two words, and FF maps that contrast (combined with the w~k\tilde{w}_k) to the ratio of their co-occurrence probabilities.

Through a series of mathematical constraints — requiring FF to be a homomorphism between addition in and multiplication in probability space — the authors show that the only solution is:

wiTw~k=log⁡(Xik)−bi−b~kw_i^T \tilde{w}_k = \log(X_{ik}) - b_i - \tilde{b}_k

In words: the dot product of a word vector and a context vector should equal the log of their co-occurrence count, minus bias terms that absorb frequency effects.

J=∑i,j=1Vf(Xij)(wiTw~j+bi+b~j−log⁡Xij)2J = \sum_{i,j=1}^{V} f(X_{ij}) \left( w_i^T \tilde{w}_j + b_i + \tilde{b}_j - \log X_{ij} \right)^2
GloVe's weighted least-squares objective — the full cost function — Sum over all observed word pairs: the squared error between the predicted log co-occurrence (dot product + biases) and the actual log co-occurrence, weighted by f(Xij)f(X_{ij}) to down-weight rare pairs and cap frequent ones. VV is vocabulary size, XijX_{ij} is the co-occurrence count, wiw_i and w~j\tilde{w}_j are the word and context vectors, and bib_i, b~j\tilde{b}_j are bias scalars.

The weighting function: taming rare and frequent pairs

Not all co-occurrence counts are equally informative. A word pair that co-occurs once might be noise; a pair that co-occurs a million times (like "the" + "of") dominates the if treated equally. GloVe addresses this with a weighting function f(x)f(x) that has three properties:

  • f(0)=0f(0) = 0: zero co-occurrence contributes nothing.
  • f(x)f(x) is non-decreasing: more frequent pairs matter more, up to a point.
  • f(x)f(x) saturates at a cap xmax⁡x_{\max}: beyond this threshold (typically 100), the weight stays at 1 — preventing hyper-frequent pairs from dominating.

The chosen form is a piecewise function: below xmax⁡x_{\max} the weight grows as (x/xmax⁡)α(x / x_{\max})^\alpha with α=0.75\alpha = 0.75, and above xmax⁡x_{\max} the weight is clamped to 1. The sub-linear exponent α=0.75\alpha = 0.75 means that doubling the count increases the weight by only about 68%, not 100%.

f(x)={(x/xmax⁡)αif x<xmax⁡1otherwisef(x) = \begin{cases} (x / x_{\max})^\alpha & \text{if } x < x_{\max} \\ 1 & \text{otherwise} \end{cases}
Weighting function — sub-linear ramp with a hard cap — Below xmax⁡x_{\max} (typically 100), the weight ramps up sub-linearly (α=0.75\alpha = 0.75). Above xmax⁡x_{\max}, it saturates at 1. This prevents hyper-frequent pairs like ("the", "of") from dominating the objective.
Open in Lab
Drag the α slider to see how the weighting curve changes. At α = 0.75, common pairs are gently down-weighted relative to a linear ramp.
The demo wakes as you arrive…

Building the co-occurrence matrix

Before can begin, GloVe needs the global co-occurrence matrix XX. This is built by sliding a window of size LL (typically 10 words on each side) across the entire corpus. For each center word, every context word within the window contributes to XijX_{ij}, but with a twist: distant words within the window contribute less. Specifically, a context word at distance dd adds 1/d1/d to the count, so the immediate neighbor adds 1, a word two positions away adds 0.5, and so on.

This distance-weighted counting is elegant: it makes the matrix capture not just whether two words co-occur, but how closely they tend to appear. The resulting matrix is symmetric (Xij=XjiX_{ij} = X_{ji}) and sparse — most word pairs never co-occur, so only nonzero entries need storage and computation. Training iterates only over observed (nonzero) cells, which gives GloVe a significant speed advantage over methods that must process every cell or every .

Open in Lab
Watch how the co-occurrence matrix fills up as the window slides across a sample sentence. Closer words get higher counts.
The demo wakes as you arrive…

Training: stochastic gradient descent on the matrix

With the co-occurrence matrix built, training is straightforward. The parameters are: a word vector wiw_i and bias bib_i for each word, plus a separate context vector w~j\tilde{w}_j and bias b~j\tilde{b}_j. The model randomly samples nonzero entries from XX and updates the parameters by gradient descent on the weighted least-squares loss.

An interesting detail: GloVe maintains two sets of vectors — word and context. After training, both carry useful information. The final embedding is their sum: wi+w~iw_i + \tilde{w}_i. This sum consistently outperforms either set alone, because the word and context vectors encode slightly different perspectives of the same word — averaging smooths out noise and captures a richer .

Training complexity scales as O(∣C∣0.8)O(|C|^{0.8}) where ∣C∣|C| is the corpus size — better than the worst case and competitive with window-based methods.

Why does king − man + woman = queen?

The most celebrated property of word vectors is the analogy: wking−wman+wwoman≈wqueenw_{king} - w_{man} + w_{woman} \approx w_{queen}. GloVe explains why this happens.

Since wiTw~k≈log⁡P(k∣i)w_i^T \tilde{w}_k \approx \log P(k|i), the difference wkingTw~k−wmanTw~k≈log⁡P(k∣king)P(k∣man)w_{king}^T \tilde{w}_k - w_{man}^T \tilde{w}_k \approx \log\frac{P(k|\text{king})}{P(k|\text{man})}. This ratio isolates the "royalty" component. Adding wwomanw_{woman} re-introduces the "female" component. The resulting vector is closest to "queen" because queen = royalty + female.

In other words, vector arithmetic works because GloVe's objective is log-probability ratios, and log turns division into subtraction. The analogy property is not a happy accident — it is a direct consequence of the .

Open in Lab
Try different analogy queries. The vectors encode relationships as directions in space: the "gender" direction, the "capital" direction, etc.
The demo wakes as you arrive…

The unification: skip-gram is implicit matrix factorization

Perhaps GloVe's deepest contribution is showing that count-based and prediction-based methods are not fundamentally different. The paper demonstrates that the skip-gram model with (SGNS) is implicitly factorizing a word-context matrix whose entries are shifted pointwise mutual information (PMI):

wiTw~j≈PMI(i,j)−log⁡kw_i^T \tilde{w}_j \approx \text{PMI}(i,j) - \log k

where kk is the number of negative samples. GloVe makes this factorization explicit and replaces the noisy stochastic training of skip-gram with a clean weighted least-squares regression on the full co-occurrence matrix.

This unification was influential: it meant researchers could analyze both families with the same mathematical tools. Later work by Levy & Goldberg (2014) formalized this connection further, showing that many apparent differences between methods boil down to choices rather than fundamental architecture.

Results: analogies, similarity, and NER

GloVe was evaluated on three standard benchmarks:

  • Word analogy (Google dataset, 19,544 questions): GloVe achieved 75% accuracy with 300-dimensional vectors trained on 42B tokens — the best result at the time. It outperformed skip-gram, CBOW, and SVD-based methods.
  • Word similarity (WordSim-353, MEN, etc.): GloVe matched or surpassed existing methods, confirming that its vectors capture semantic similarity.
  • Named Entity Recognition (CoNLL-2003): Using GloVe vectors as input features for a CRF-based NER system improved F1 scores, demonstrating that the vectors carry useful information for downstream tasks beyond intrinsic benchmarks.

The authors also studied the effect of hyperparameters: larger corpora, larger vectors (up to 300 dimensions), and symmetric context windows all helped. The training data matters: Wikipedia + Gigaword outperformed either alone.

Open in Lab
Compare GloVe's performance against skip-gram and SVD across three benchmark categories.
The demo wakes as you arrive…

Using GloVe vectors in practice

Loading pre-trained GloVe and testing analogiespython

Simplified to show the idea — not the real implementation.

import numpy as np
# Load pre-trained GloVe vectors (50-dimensional for demo) def load_glove(path):
    """Read glove.6B.50d.txt into a {word: vector} dict."""
    vectors = {}
    with open(path, encoding='utf-8') as f:
        for line in f:
            parts = line.split()
            word = parts[0]
            vec = np.array(parts[1:], dtype=np.float32)
            vectors[word] = vec
    return vectors

glove = load_glove('glove.6B.50d.txt')
# Analogy: king - man + woman ≈ ? def analogy(a, b, c, vecs, top_k=5):
    """Solve: a is to b as c is to ?"""
    query = vecs[b] - vecs[a] + vecs[c]
    # Cosine similarity against every word
    scores = {w: np.dot(query, v) / (np.linalg.norm(query) * np.linalg.norm(v))
              for w, v in vecs.items() if w not in {a, b, c}}
    return sorted(scores, key=scores.get, reverse=True)[:top_k]

print(analogy('man', 'king', 'woman', glove)) # Expected: ['queen', 'princess', 'monarch', ...]

Legacy and impact

GloVe demonstrated two lasting principles. First, explicit matrix factorization and implicit neural prediction are mathematically equivalent — a unification that clarified the entire field of distributional semantics. Second, simple models with the right objective can compete with or surpass complex neural architectures, a reminder that inductive bias matters more than model complexity.

The pre-trained GloVe vectors served as the default initialization for downstream NLP models from 2014 to 2018, before contextualized embeddings (ELMo and BERT) replaced static word vectors. Even then, GloVe's intellectual DNA persists: the embedding layers in Transformers learn a similar factorization, and modern retrieval systems still use GloVe-inspired dot-product similarity.

  1. 2013

    Word2Vec

    Mikolov et al. introduced skip-gram and CBOW — prediction-based word embeddings that popularized the analogy task and vector arithmetic.

  2. 2014

    GloVe published

    Pennington, Socher & Manning unified count-based and prediction-based methods. Pre-trained vectors on 6B, 42B, and 840B tokens became an NLP standard.

  3. 2014

    Levy & Goldberg analysis

    Formalized the connection between skip-gram and PMI matrix factorization, confirming GloVe's unification thesis with rigorous proofs.

  4. 2016

    FastText

    Extended Word2Vec by representing words as bags of character n-grams. This handled morphology and out-of-vocabulary words, but still used static vectors like GloVe.

  5. 2018

    ELMo — contextualized embeddings

    Peters et al. showed that the same word should get different vectors in different contexts. ELMo used deep BiLSTMs over characters, making static vectors (GloVe, Word2Vec) a pre-ELMo era artifact.

  6. 2025

    GloVe 2024 vectors released

    Stanford released updated GloVe vectors trained on the Dolma corpus (220B tokens), proving that the model remains relevant a decade later for lightweight, interpretable word representations.

GloVe's pre-trained vectors were the invisible backbone of NLP for half a decade. Every sentiment classifier, every question-answering system, every model between 2014 and 2018 likely started with GloVe or Word2Vec embeddings. The era of static word vectors may have passed, but the mathematical insight — that co-occurrence ratios encode meaning, and that counting and predicting are the same — remains foundational.

CitationPennington, J., Socher, R., & Manning, C. D.. GloVe: Global Vectors for Word Representation. EMNLP, 2014.

Terms in this paper