Language Models2017beginner10 min read
Enriching Word Vectors with Subword Information
إثراء متجهات الكلمات بمعلومات الوحدات الفرعية
Bojanowski, P. · Grave, E. · Joulin, A. · Mikolov, T. — TACL
The problem
Word2Vec assigns one per word — if a word never appeared in , it gets no at all. This is devastating for morphologically rich languages like Arabic, Finnish, or Turkish, where a single root can generate hundreds of inflected forms. Even in English, rare words like "unbreakableness" get poor vectors because they appear too few times. The internal structure of words — prefixes, suffixes, roots — is completely ignored.
The contribution
FastText extends the by representing each word as a bag of character n-grams. The word "where" with n=3 becomes: <wh, whe, her, ere, re>, plus the whole token <where>. Each n-gram gets its own vector; the word's vector is the sum of its n-gram vectors. This means unseen words still get meaningful representations from their character pieces, and morphologically related words (run, runs, running) naturally share most of their n-grams. On word benchmarks for morphologically rich languages, FastText outperforms Word2Vec by a large margin.
The impact
FastText became the practical standard for word embeddings in production systems, especially for multilingual NLP. Its pre-trained vectors covering 157 languages remain among the most downloaded resources in NLP. The idea directly influenced tokenizers used in every modern LLM — GPT, BERT, and Claude all tokenize at the subword level, a principle FastText helped popularize. It bridged the gap between word-level models and the character-aware architectures that followed.
Word2Vec treats every word like an ID badge: "running" and "runner" are two completely unrelated badges, even though they clearly share a root.
FastText treats words like LEGO bricks: the word "running" is built from pieces — "run", "unn", "nni", "nin", "ing" — and those pieces are reusable. A word you've never seen before? No problem — if its bricks exist in the collection, you can still build a meaning for it.
It's like learning a language by understanding word roots and patterns, rather than memorizing every word in the dictionary.
The problem: one vector per word is not enough
Word2Vec was a breakthrough: it showed that you can learn meaningful word vectors from raw text by predicting context words. But it treats every word as an atomic unit — an opaque symbol with no internal structure. This creates three concrete problems:
-
words. Any word not seen during training gets no vector at all. In a production system, this means a typo, a new product name, or a rare inflection produces a crash or a fallback to a generic "unknown" token.
-
Morphological blindness. "teach", "teacher", "teaching", and "teaches" are four unrelated entries. The obvious shared meaning carried by the root "teach" is invisible to the model. For Arabic, where a single root like «ك-ت-ب» generates dozens of forms (كتب، كاتب، مكتوب، كتابة، مكتبة), this blindness is catastrophic.
-
Data hunger for rare words. Words that appear only a few times get poor vectors because there isn't enough context to learn from. Morphological information could bootstrap these rare words, but Word2Vec has no mechanism to use it.
The idea: break words into character n-grams
The core insight is simple: represent each word not as one vector, but as the sum of vectors of its character n-grams. Here's how it works step by step:
Step 1 — Add boundary markers. Wrap the word in special boundary symbols to distinguish prefixes from suffixes. "where" becomes <where>.
Step 2 — Extract character n-grams. For n-gram sizes 3 to 6, slide a window across the word. For "where" with n=3: <wh, whe, her, ere, re>.
Step 3 — Include the whole word. Add a special entry for the full word <where> to preserve word-level information that character pieces alone might miss.
Step 4 — Sum the vectors. The final word vector is the sum of all its n-gram vectors plus the whole-word vector. That's it — no architecture change, no mechanism. Just a richer input representation.
The scoring function: from words to subword sums
In the original skip-gram, the score between a center word and a context word is a simple of their vectors: . FastText replaces the center word's vector with a sum over its n-gram set :
This is the only change to the skip-gram model. Everything else — the sliding context window, , — stays identical. Yet this one change has profound consequences: now the model shares parameters across words that share character sequences, and it can compose vectors for words never seen during training.
Why this works: morphology for free
Consider the English words "teach", "teacher", "teaching", "teaches", and "taught". In Word2Vec, these are five separate entries that must each accumulate enough training examples independently. In FastText, they all share n-grams like "tea", "eac", "ach" — so training data for one helps all the others.
The effect is even more dramatic in morphologically rich languages. German compounds like "Donaudampfschifffahrtsgesellschaftskapitän" share n-grams with simpler words like "Donau" (Danube) and "Schiff" (ship). Arabic's root system means that «كِتاب» (book), «كاتِب» (writer), «مَكتوب» (written), and «مَكتَبة» (library) all share character sequences from the root «ك-ت-ب». FastText captures this automatically, without any explicit morphological analyzer.
The OOV solution: composing vectors for unseen words
This is perhaps the most practically important feature. When FastText encounters a word not in its vocabulary, it:
- Wraps it in boundary markers 2. Extracts all character n-grams of sizes 3–6 3. Looks up the vectors for any n-grams it has seen before 4. Sums those vectors to produce a representation
The result won't be as good as a well-trained word, but it's enormously better than nothing. A misspelling like "languge" shares most n-grams with "language", so its composed vector will be close. A new technical term like "transformerized" inherits meaning from its familiar pieces.
This matters in production: real users misspell, new words emerge daily, and languages have productive that generates valid words faster than any can cover.
Training: negative sampling with subwords
FastText uses the same training objective as skip-gram with negative sampling. For each word–context pair observed in the corpus, the model also samples negative (random) context words. The objective maximizes the log- of the score for the true pair and minimizes it for the negatives:
The key difference from Word2Vec: when we compute gradients and update vectors, the flows back to every n-gram vector in . So each training example updates not just one word's vector, but all the character n-grams that compose it — and those n-grams are shared with other words. A training example for "running" also improves the representations of "run", "runner", and every other word containing the same pieces.
Practical trick: hashing n-grams to save memory
A natural worry: if every word produces dozens of n-grams, and there are millions of words in the vocabulary, the total number of distinct n-grams could be enormous. FastText uses a hashing trick: each n-gram string is mapped through a hash function to one of buckets (typically ). N-grams that hash to the same bucket share a vector.
This bounds the memory at exactly parameters for the n-gram table (where is the dimension), regardless of how many distinct n-grams exist. The collision rate is low enough that it barely affects quality — a remarkably elegant engineering solution.
Results: where FastText shines
The authors evaluated on two families of tasks across many languages:
Word similarity — how well do vector distances match human similarity judgments? On morphologically rich languages (German, Czech, Arabic), FastText dramatically outperformed skip-gram, with gains of 5–10 points in Spearman correlation. The improvement was smaller for English, which has simpler morphology.
— can the model solve "king:queen :: man:?" type tasks? FastText excelled on syntactic analogies (which test morphological patterns like pluralization and tense), while performing comparably on semantic analogies (which test meaning relationships like capitals and currencies). This confirms that the subword mechanism primarily captures morphological structure.
The full picture
Here is the complete FastText pipeline — from raw text to a trained model that can handle any word, seen or unseen:
The idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
def get_ngrams(word, min_n=3, max_n=6):
"""Break a word into character n-grams with boundary markers."""
word = f"<{word}>" # add boundary markers
ngrams = []
for n in range(min_n, max_n + 1):
for i in range(len(word) - n + 1):
ngrams.append(word[i:i+n])
ngrams.append(word) # include the whole word too
return ngrams
def fasttext_score(word, context_vec, ngram_vectors):
"""Score a word–context pair using subword decomposition."""
ngrams = get_ngrams(word)
# Word vector = sum of all its n-gram vectors
word_vec = sum(ngram_vectors[ng] for ng in ngrams if ng in ngram_vectors)
# Score = dot product of (summed n-gram vector) and context vector
return float(np.dot(word_vec, context_vec))
# Example: "where" → ['<wh', 'whe', 'her', 'ere', 're>', '<whe', 'wher', ...]
# "runs" shares n-grams with "run", "running", "runner"
# An unseen word like "kitchenette" inherits from "kitchen" n-gramsLegacy: from subwords to modern tokenizers
2013
Word2Vec
Proved that simple neural networks trained on word co-occurrence learn rich semantic vectors. One vector per word, no internal structure.
2014
GloVe
Combined count-based and predictive methods for word vectors. Still one vector per word, no morphological awareness.
2016
BPE for Neural MT
Sennrich et al. applied Byte Pair Encoding to create subword vocabularies for translation. Solved the OOV problem for neural MT.
2017
FastText (this paper)
Showed that character n-grams enrich word vectors — handling OOV words, capturing morphology, and improving multilingual embeddings.
2018
ELMo — Contextualized Embeddings
Took subword awareness further: each word gets a different vector depending on its sentence context, using deep bidirectional LSTMs.
2018
BERT — WordPiece Tokenization
Combined subword tokenization (WordPiece) with deep bidirectional Transformers. The subword principle became standard for all LLMs.
FastText occupies a pivotal position: it showed that you don't need to choose between word-level and character-level models. By operating at the subword level, it captured morphological structure without sacrificing the efficiency and simplicity of Word2Vec's skip-gram. This principle — subword decomposition — became the foundation for every in modern NLP.
CitationBojanowski, Grave, Joulin, Mikolov. Enriching Word Vectors with Subword Information. TACL, 2017.
Terms in this paper
- Subwordجزء الكلمة
- Word Embeddingتضمين الكلمة
- Character N-gramسلسلة حرفية فرعية
- Skip-gramنموذج التخطي (Skip-gram)
- Morphologyالصرف
- Out-of-Vocabulary (OOV)خارج قاموس المفردات
- Negative Samplingالتعيين السلبي
- Byte Pair Encoding (BPE)ترميز زوج البايت