Language Models2013foundational10 min read
Distributed Representations of Words and Phrases and Their Compositionality
التمثيلات الموزَّعة للكلمات والعبارات وخاصية التركيب
Mikolov, T. · Sutskever, I. · Chen, K. · Corrado, G. · Dean, J. — NeurIPS
The problem
Before , representing words for machine learning was crude: either a one-hot of size (sparse, no similarity signal) or hand-crafted features. Full neural language models like Bengio's 2003 model could learn good representations, but the over the entire vocabulary at each step made prohibitively expensive for large corpora. There was no scalable method to learn dense word vectors that captured rich semantic and syntactic patterns.
The contribution
Three practical advances that made the model trainable at billion-word scale: (1) — replace the expensive softmax over the full vocabulary with a binary classification task that contrasts true context words against a few randomly drawn "" words. (2) of frequent words — randomly discard common words like "the" and "a" during training, yielding speedup and better rare-word vectors. (3) — a statistical scoring method identifies multi-word expressions ("New York", "ice cream") and treats them as single tokens. Together these produced word vectors with remarkable compositional arithmetic: vec("king") − vec("man") + vec("woman") ≈ vec("queen").
The impact
Word2Vec democratized word embeddings. Within two years GloVe and fastText followed, and dense word vectors became the default input to virtually every NLP pipeline. The idea of learning representations from context — rather than hand-engineering features — foreshadowed the pre-training revolution that led to BERT, GPT, and every modern large . DeepWalk extended the skip-gram idea to graphs, and Contrastive Predictive Coding generalized it to speech and images.
Imagine a dictionary where every word's definition is just a list of numbers — not characters, but coordinates in a vast conceptual space. Words used in similar sentences drift toward each other like magnets on a whiteboard: "happy" and "joyful" cluster together, while "happy" and "concrete" repel to distant corners.
The astonishing part: directions in this space encode meaning. Walk from "man" to "king" and the same step takes you from "woman" to "queen". The model never learned royal semantics — it learned geometry of context.
The problem: words as meaningless IDs
The simplest way to feed a word to a is : a vector of zeros with a single 1 at the word's index. If the vocabulary has 100,000 words, every word becomes a 100,000-dimensional vector — enormous, sparse, and completely blind to similarity. "Cat" and "kitten" are as far apart as "cat" and "parliament".
Full neural language models (Bengio et al., 2003) solved this by learning dense embeddings, but their computed a softmax over the entire vocabulary at every training step. For a vocabulary of V words, that means V dot products plus a normalization — an operation that scales linearly with V and becomes a bottleneck when V reaches hundreds of thousands.
The Skip-gram model: predict context from a center word
The Skip-gram model flips the typical language model around. Instead of predicting the next word from its history, it takes a center word and tries to predict the surrounding context words within a window of size . Think of it as a spotlight: the center word asks, "Who usually stands near me?"
Given a sequence of training words , the objective is to maximize the average log probability:
Each word has two roles — and therefore two vectors. As a center word it uses its input vector ; as a context word it uses its output vector . The probability of seeing context word given center word is defined via softmax over the full vocabulary.
Negative Sampling: the key speedup
The expensive part of Skip-gram is the softmax denominator — summing over every word in the vocabulary. Negative Sampling (NEG) replaces this with a far cheaper surrogate objective.
The idea is elegant: instead of asking "what is the probability of this context word among all V words?", ask a binary question — "did this word-context pair come from real data, or is it noise?" For each real pair the model should output high probability. Then draw random "negative" words from a and the model should output low probability for each of them.
Imagine a customs officer. Instead of checking every passenger on the plane (full softmax), she checks the ticket of the arriving passenger (the positive pair) and then spot-checks random people from the airport crowd (negative samples). Much faster, and still catches counterfeits.
Subsampling frequent words: less is more
Words like "the", "a", and "in" appear millions of times in any large . They provide much less information than rare words — seeing "the" next to "France" tells you almost nothing, while "Paris" next to "France" is highly informative. Yet without correction, the model spends most of its compute on these uninformative pairs.
The paper introduces a simple subsampling formula: each word in the training data is discarded with a probability that grows with its frequency.
Learning phrases: "New York" is not "New" + "York"
Some word combinations carry meaning that their parts alone do not convey. "New York" is a city, not something new and something called York. "Ice cream" is a dessert, not frozen cream. Treating these as separate tokens loses critical meaning.
The paper uses a data-driven scoring function based on to identify such phrases:
When the score exceeds a threshold, the bigram is merged into a single . Running this process multiple times discovers longer phrases like "New_York_Times". The approach is simple but effective — no linguistic rules, purely statistical co-occurrence.
Vector arithmetic: the magic of compositionality
The most celebrated result of Word2Vec is that its vector space encodes semantic relationships as directions. The classic demonstration:
This is not a trick or cherry-picked example — it works across many relationship types. The direction from "man" to "woman" captures a gender axis; the direction from "Paris" to "France" captures a capital-country axis. These directions are consistent: the same offset that maps "Paris" → "France" also maps "Berlin" → "Germany".
This arises because the Skip-gram objective implicitly factorizes a word-context co-occurrence matrix. Words that share similar contexts end up at similar positions, and systematic relationships in language create systematic geometric patterns in the space.
Hierarchical Softmax: the tree-based alternative
Before Negative Sampling, the paper also discusses — an earlier approach to avoiding the full-vocabulary sum. The idea: arrange all words as leaves of a binary tree. Instead of computing one V-way softmax, the model makes a series of binary decisions as it walks down the tree from root to the target word's leaf.
Each internal node has a learned vector, and at each branch the model computes a to decide left or right. The total path length is — so the cost drops from to . A Huffman tree assigns shorter paths to frequent words, further speeding up common predictions.
The paper found that Negative Sampling outperformed Hierarchical Softmax on analogy tasks, particularly for frequent words, while being simpler to implement. This is why Negative Sampling became the default training method for Word2Vec.
Full training pipeline
The complete Word2Vec training pipeline combines all three contributions into a streamlined process:
Step 1 — Phrase detection. Scan the corpus and merge high-scoring bigrams into single tokens. Run multiple passes for longer phrases.
Step 2 — Build vocabulary. Count all (merged) tokens. Discard any below a minimum frequency threshold.
Step 3 — Subsample frequent words. For each training sentence, randomly drop words with probability tied to their frequency. This both accelerates training and improves quality.
Step 4 — Train with Negative Sampling. For each surviving center-context pair: push their vectors closer together; sample noise words and push their vectors apart from the center word. Update via .
The result: a matrix of dense vectors — one per word — where geometric relationships encode semantic meaning. Training on a billion-word corpus takes hours, not weeks.
Why addition works: the log-linear connection
It may seem magical that vector addition captures semantic analogies. The paper offers an intuitive explanation rooted in the training objective.
Skip-gram's probability is defined via exponentiated dot products. Taking logs converts multiplication to addition. If words that share a relationship consistently appear in similar contexts, the offset vector between them points in a consistent direction. Adding and subtracting these offsets navigates the space along meaningful axes.
This is not a guaranteed property of any embedding — it emerges because the Skip-gram objective implicitly captures log-probability ratios of co-occurrence statistics. Later work (Levy & Goldberg, 2014) showed this connection rigorously: Skip-gram with negative sampling implicitly factorizes a shifted PMI (pointwise mutual information) matrix.
What Word2Vec unlocked
2013
Word2Vec (this paper)
Skip-gram + Negative Sampling. Showed that simple log-linear models trained on vast text produce vectors with remarkable compositionality.
2014
GloVe
Pennington et al. combined global co-occurrence statistics with local context windows. Showed Word2Vec implicitly factorizes a PMI matrix and proposed an explicit factorization that matched or beat it.
2014
DeepWalk
Applied Skip-gram to random walks on graphs — nodes become "words", walks become "sentences". Extended embeddings beyond language to social networks and knowledge graphs.
2017
fastText
Bojanowski et al. enriched Word2Vec with character n-grams, allowing the model to construct embeddings for unseen words by combining sub-word pieces. Essential for morphologically rich languages.
2018
ELMo → contextualized embeddings
Peters et al. showed that static vectors (one vector per word) miss polysemy. Contextual embeddings from deep LSTMs gave "bank" different vectors in "river bank" vs "bank account". The idea Word2Vec planted evolved into context-dependent representations.
2019
CPC — Contrastive Predictive Coding
Van den Oord et al. generalized the "predict context from target" idea to speech, images, and video. The contrastive objective descends directly from Negative Sampling.
Word2Vec's most enduring contribution is not any particular set of vectors — it is the idea that unsupervised co-occurrence prediction creates structured representations. Every modern pre-trained model, from BERT to GPT, is a descendant of this insight: learn from context, and meaning emerges in the geometry of the learned space.
CitationMikolov, Sutskever, Chen, Corrado, Dean. Distributed Representations of Words and Phrases and Their Compositionality. NeurIPS, 2013.
Terms in this paper
- Skip-gramنموذج التخطي (Skip-gram)
- Negative Samplingالتعيين السلبي
- Word2Vecخوارزمية تحويل الكلمات إلى متجهات
- subsamplingالتقليص المكاني
- Phrase Detectionكشف العبارات
- Compositionalityالتركيبية الدلالية
- Hierarchical Softmaxسوفت ماكس الهرمي
- Noise Distributionتوزيع الضوضاء
- Pointwise Mutual Information (PMI)المعلومات المتبادلة النقطية (PMI)