Language Models2017beginner10 min read

Enriching Word Vectors with Subword Information

إثراء متجهات الكلمات بمعلومات الوحدات الفرعية

Bojanowski, P. · Grave, E. · Joulin, A. · Mikolov, T. — TACL

The problem

Word2Vec assigns one per word — if a word never appeared in , it gets no at all. This is devastating for morphologically rich languages like Arabic, Finnish, or Turkish, where a single root can generate hundreds of inflected forms. Even in English, rare words like "unbreakableness" get poor vectors because they appear too few times. The internal structure of words — prefixes, suffixes, roots — is completely ignored.

The contribution

FastText extends the by representing each word as a bag of character n-grams. The word "where" with n=3 becomes: <wh, whe, her, ere, re>, plus the whole token <where>. Each n-gram gets its own vector; the word's vector is the sum of its n-gram vectors. This means unseen words still get meaningful representations from their character pieces, and morphologically related words (run, runs, running) naturally share most of their n-grams. On word benchmarks for morphologically rich languages, FastText outperforms Word2Vec by a large margin.

The impact

FastText became the practical standard for word embeddings in production systems, especially for multilingual NLP. Its pre-trained vectors covering 157 languages remain among the most downloaded resources in NLP. The idea directly influenced tokenizers used in every modern LLM — GPT, BERT, and Claude all tokenize at the subword level, a principle FastText helped popularize. It bridged the gap between word-level models and the character-aware architectures that followed.

Word2Vec treats every word like an ID badge: "running" and "runner" are two completely unrelated badges, even though they clearly share a root.

FastText treats words like LEGO bricks: the word "running" is built from pieces — "run", "unn", "nni", "nin", "ing" — and those pieces are reusable. A word you've never seen before? No problem — if its bricks exist in the collection, you can still build a meaning for it.

It's like learning a language by understanding word roots and patterns, rather than memorizing every word in the dictionary.

The problem: one vector per word is not enough

Word2Vec was a breakthrough: it showed that you can learn meaningful word vectors from raw text by predicting context words. But it treats every word as an atomic unit — an opaque symbol with no internal structure. This creates three concrete problems:

  • words. Any word not seen during training gets no vector at all. In a production system, this means a typo, a new product name, or a rare inflection produces a crash or a fallback to a generic "unknown" token.

  • Morphological blindness. "teach", "teacher", "teaching", and "teaches" are four unrelated entries. The obvious shared meaning carried by the root "teach" is invisible to the model. For Arabic, where a single root like «ك-ت-ب» generates dozens of forms (كتب، كاتب، مكتوب، كتابة، مكتبة), this blindness is catastrophic.

  • Data hunger for rare words. Words that appear only a few times get poor vectors because there isn't enough context to learn from. Morphological information could bootstrap these rare words, but Word2Vec has no mechanism to use it.

Open in Lab
Type a word — Word2Vec can only handle words it memorized. FastText builds a vector from character pieces, even for words it never saw.
The demo wakes as you arrive…

The idea: break words into character n-grams

The core insight is simple: represent each word not as one vector, but as the sum of vectors of its character n-grams. Here's how it works step by step:

Step 1 — Add boundary markers. Wrap the word in special boundary symbols to distinguish prefixes from suffixes. "where" becomes <where>.

Step 2 — Extract character n-grams. For n-gram sizes 3 to 6, slide a window across the word. For "where" with n=3: <wh, whe, her, ere, re>.

Step 3 — Include the whole word. Add a special entry for the full word <where> to preserve word-level information that character pieces alone might miss.

Step 4 — Sum the vectors. The final word vector is the sum of all its n-gram vectors plus the whole-word vector. That's it — no architecture change, no mechanism. Just a richer input representation.

Open in Lab
Type any word and see how FastText breaks it into character n-grams with boundary markers.
The demo wakes as you arrive…

The scoring function: from words to subword sums

In the original skip-gram, the score between a center word ww and a context word cc is a simple of their vectors: s(w,c)=uw⊤vcs(w, c) = \mathbf{u}_w^\top \mathbf{v}_c. FastText replaces the center word's vector with a sum over its n-gram set Gw\mathcal{G}_w:

s(w,c)=∑g∈Gwzg⊤vcs(w, c) = \sum_{g \in \mathcal{G}_w} \mathbf{z}_g^\top \mathbf{v}_c
FastText scoring function — the word's score is the sum of its n-gram scores — Each n-gram g has a learned vector z_g · The context word c has a vector v_c · The overall score sums every n-gram's dot product with the context — if most character pieces point toward the context word, the score is high

This is the only change to the skip-gram model. Everything else — the sliding context window, , — stays identical. Yet this one change has profound consequences: now the model shares parameters across words that share character sequences, and it can compose vectors for words never seen during training.

Open in Lab
See how each n-gram contributes its dot product to the total word–context score.
The demo wakes as you arrive…

Why this works: morphology for free

Consider the English words "teach", "teacher", "teaching", "teaches", and "taught". In Word2Vec, these are five separate entries that must each accumulate enough training examples independently. In FastText, they all share n-grams like "tea", "eac", "ach" — so training data for one helps all the others.

The effect is even more dramatic in morphologically rich languages. German compounds like "Donaudampfschifffahrtsgesellschaftskapitän" share n-grams with simpler words like "Donau" (Danube) and "Schiff" (ship). Arabic's root system means that «كِتاب» (book), «كاتِب» (writer), «مَكتوب» (written), and «مَكتَبة» (library) all share character sequences from the root «ك-ت-ب». FastText captures this automatically, without any explicit morphological analyzer.

Open in Lab
Compare how Word2Vec and FastText handle morphologically related words. Notice how shared n-grams create natural similarity.
The demo wakes as you arrive…

The OOV solution: composing vectors for unseen words

This is perhaps the most practically important feature. When FastText encounters a word not in its vocabulary, it:

  1. Wraps it in boundary markers 2. Extracts all character n-grams of sizes 3–6 3. Looks up the vectors for any n-grams it has seen before 4. Sums those vectors to produce a representation

The result won't be as good as a well-trained word, but it's enormously better than nothing. A misspelling like "languge" shares most n-grams with "language", so its composed vector will be close. A new technical term like "transformerized" inherits meaning from its familiar pieces.

This matters in production: real users misspell, new words emerge daily, and languages have productive that generates valid words faster than any can cover.

Open in Lab
Enter an invented or misspelled word and watch FastText compose its vector from known n-grams.
The demo wakes as you arrive…

Training: negative sampling with subwords

FastText uses the same training objective as skip-gram with negative sampling. For each word–context pair (w,c)(w, c) observed in the corpus, the model also samples kk negative (random) context words. The objective maximizes the log- of the score for the true pair and minimizes it for the negatives:

L=log⁡σ ⁣(s(w,c))+∑i=1kEni ⁣[log⁡σ ⁣(−s(w,ni))]\mathcal{L} = \log\sigma\!\left(s(w,c)\right) + \sum_{i=1}^{k}\mathbb{E}_{n_i}\!\left[\log\sigma\!\left(-s(w,n_i)\right)\right]
Negative sampling loss — push real pairs up, random pairs down — s(w,c) is the subword-sum score from the previous formula · σ is the sigmoid function · The first term rewards high scores for real word–context pairs · The sum penalizes high scores for random (negative) pairs · k is typically 5–20

The key difference from Word2Vec: when we compute gradients and update vectors, the flows back to every n-gram vector in Gw\mathcal{G}_w. So each training example updates not just one word's vector, but all the character n-grams that compose it — and those n-grams are shared with other words. A training example for "running" also improves the representations of "run", "runner", and every other word containing the same pieces.

Practical trick: hashing n-grams to save memory

A natural worry: if every word produces dozens of n-grams, and there are millions of words in the vocabulary, the total number of distinct n-grams could be enormous. FastText uses a hashing trick: each n-gram string is mapped through a hash function to one of BB buckets (typically B=2,000,000B = 2{,}000{,}000). N-grams that hash to the same bucket share a vector.

This bounds the memory at exactly B×dB \times d parameters for the n-gram table (where dd is the dimension), regardless of how many distinct n-grams exist. The collision rate is low enough that it barely affects quality — a remarkably elegant engineering solution.

Open in Lab
Watch how different n-grams map to hash buckets. Collisions are rare but possible.
The demo wakes as you arrive…

Results: where FastText shines

The authors evaluated on two families of tasks across many languages:

Word similarity — how well do vector distances match human similarity judgments? On morphologically rich languages (German, Czech, Arabic), FastText dramatically outperformed skip-gram, with gains of 5–10 points in Spearman correlation. The improvement was smaller for English, which has simpler morphology.

— can the model solve "king:queen :: man:?" type tasks? FastText excelled on syntactic analogies (which test morphological patterns like pluralization and tense), while performing comparably on semantic analogies (which test meaning relationships like capitals and currencies). This confirms that the subword mechanism primarily captures morphological structure.

Open in Lab
Compare word similarity scores: FastText vs. skip-gram across languages with different morphological richness.
The demo wakes as you arrive…

The full picture

Here is the complete FastText pipeline — from raw text to a trained model that can handle any word, seen or unseen:

Open in Lab
Click each stage to see what happens inside the FastText training pipeline.
The demo wakes as you arrive…

The idea in code

FastText subword scoring — the complete mechanismpython

Simplified to show the idea — not the real implementation.

import numpy as np

def get_ngrams(word, min_n=3, max_n=6):
    """Break a word into character n-grams with boundary markers."""
    word = f"<{word}>"              # add boundary markers
    ngrams = []
    for n in range(min_n, max_n + 1):
        for i in range(len(word) - n + 1):
            ngrams.append(word[i:i+n])
    ngrams.append(word)             # include the whole word too
    return ngrams

def fasttext_score(word, context_vec, ngram_vectors):
    """Score a word–context pair using subword decomposition."""
    ngrams = get_ngrams(word)
    # Word vector = sum of all its n-gram vectors
    word_vec = sum(ngram_vectors[ng] for ng in ngrams if ng in ngram_vectors)
    # Score = dot product of (summed n-gram vector) and context vector
    return float(np.dot(word_vec, context_vec))

# Example: "where" → ['<wh', 'whe', 'her', 'ere', 're>', '<whe', 'wher', ...]
# "runs" shares n-grams with "run", "running", "runner"
# An unseen word like "kitchenette" inherits from "kitchen" n-grams

Legacy: from subwords to modern tokenizers

  1. 2013

    Word2Vec

    Proved that simple neural networks trained on word co-occurrence learn rich semantic vectors. One vector per word, no internal structure.

  2. 2014

    GloVe

    Combined count-based and predictive methods for word vectors. Still one vector per word, no morphological awareness.

  3. 2016

    BPE for Neural MT

    Sennrich et al. applied Byte Pair Encoding to create subword vocabularies for translation. Solved the OOV problem for neural MT.

  4. 2017

    FastText (this paper)

    Showed that character n-grams enrich word vectors — handling OOV words, capturing morphology, and improving multilingual embeddings.

  5. 2018

    ELMo — Contextualized Embeddings

    Took subword awareness further: each word gets a different vector depending on its sentence context, using deep bidirectional LSTMs.

  6. 2018

    BERT — WordPiece Tokenization

    Combined subword tokenization (WordPiece) with deep bidirectional Transformers. The subword principle became standard for all LLMs.

FastText occupies a pivotal position: it showed that you don't need to choose between word-level and character-level models. By operating at the subword level, it captured morphological structure without sacrificing the efficiency and simplicity of Word2Vec's skip-gram. This principle — subword decomposition — became the foundation for every in modern NLP.

CitationBojanowski, Grave, Joulin, Mikolov. Enriching Word Vectors with Subword Information. TACL, 2017.

Terms in this paper