Language Models2003foundational10 min read

A Neural Probabilistic Language Model

نموذج لغوي عصبي احتمالي

Bengio, Y. · Ducharme, R. · Vincent, P. · Jauvin, C. — JMLR

The problem

Traditional n-gram language models treat every word as an opaque symbol. A of 100,000 words produces 10¹⁵ possible trigrams — most never appear in any corpus. Even if the has seen "The cat is walking in the bedroom", it learns nothing about "A dog was running in a room". This is the : the number of possible word sequences grows exponentially, but training data is finite.

The contribution

Replace the discrete word table with a learned continuous vector space. Each word is mapped to a dense via a shared lookup table C. The previous n−1 embeddings are concatenated and fed through a feedforward ( with tanh, then over the full vocabulary) to predict the next word. The network and the embeddings are trained jointly by maximizing log-likelihood with descent. Because similar words get similar vectors, the model generalizes to word sequences never seen during training.

The impact

This paper introduced word embeddings — the idea that launched a revolution. Word2Vec, GloVe, FastText, and ultimately every embedding layer descend from this insight. It proved that neural networks can beat n-grams at language modeling, planted the seed of representation learning in NLP, and directly inspired the sequence-to-sequence models and pre-training paradigms that define modern AI.

An n-gram model is like a phone book: it knows that "John Smith" is at 555-0123, but if you ask for "Jonathan Smith" — a person who lives next door — it draws a blank. Every name is an isolated entry.

Bengio's neural replaces the phone book with a map of the neighborhood: people who live close together share similar addresses. If you know where John lives, you can guess where Jonathan lives — because they're neighbors. The map is the embedding space, and closeness on the map is closeness in meaning.

The problem: the curse of dimensionality in language

Before 2003, language models were built on counting. A trigram model estimates P(wt∣wt−1,wt−2)P(w_t \mid w_{t-1}, w_{t-2}) by counting how often that triplet appeared in the training data. The problem is combinatorial explosion: with a vocabulary of size ∣V∣|V|, there are ∣V∣3|V|^3 possible trigrams. For ∣V∣=100,000|V| = 100{,}000, that's 101510^{15} entries — far more than any corpus could ever cover. Most legitimate word sequences will never be seen, and the model assigns them zero .

Smoothing techniques (Kneser-Ney, Good-Turing) redistribute a small amount of probability mass to unseen n-grams, but this is a patch on a fundamentally discrete representation. The model has no notion of similarity: it cannot use the fact that "dog" and "cat" are both animals to transfer knowledge from one context to another. Each word is an opaque symbol — slot 4721 in a table.

Open in Lab
Watch the number of possible n-grams explode as vocabulary or context length grows. Most cells remain empty — the model can't score them.
The demo wakes as you arrive…

The idea: give every word a position on a map

Bengio's breakthrough has two parts:

1. Distributed word representations. Associate each word ww in the vocabulary with a real-valued C(w)∈RmC(w) \in \mathbb{R}^m. This is a lookup table — a matrix of size ∣V∣×m|V| \times m — whose rows are learned by backpropagation. Instead of treating "dog" as symbol #4721, the model knows dog as [0.32,−0.15,0.87,…][0.32, -0.15, 0.87, \ldots]. Similar words end up with similar vectors because the function rewards the model for assigning similar probabilities to similar contexts.

2. A neural probability function. Instead of counting, use a neural network to compute P(wt∣wt−1,…,wt−n+1)P(w_t \mid w_{t-1}, \ldots, w_{t-n+1}). The embeddings of the context words are concatenated into a single long vector and passed through a hidden layer to produce a probability distribution over the full vocabulary.

The key insight: if "A dog was running" and "A cat was running" have similar embeddings for "dog" and "cat", the probability of "in the room" following both is automatically similar — without ever seeing both sequences in training. happens through the continuous structure of the embedding space, not through brute-force memorization.

Open in Lab
Drag words around to see how proximity in embedding space enables generalization. Similar words cluster together.
The demo wakes as you arrive…

The architecture: lookup → concat → hidden → softmax

The model has four stages. Given the context words wt−n+1,…,wt−1w_{t-n+1}, \ldots, w_{t-1}:

Stage 1 — Embedding lookup. Look up each context word in the shared matrix CC to get its feature vector C(wt−i)∈RmC(w_{t-i}) \in \mathbb{R}^m. This is the same table for every position, so the model learns one representation per word, shared across all contexts.

Stage 2 — Concatenation. Concatenate the n−1n{-}1 feature vectors into a single vector x∈R(n−1)⋅mx \in \mathbb{R}^{(n-1) \cdot m}. This is how the model sees its context — as one long corridor of numbers.

Stage 3 — Hidden layer. Transform xx through a hidden layer: h=tanh⁡(Hx+d)h = \tanh(Hx + d). The tanh activation squashes values to [−1,1][-1, 1] and introduces the nonlinearity the model needs to learn complex patterns.

Stage 4 — Output layer. A matrix UU maps hh to a vector of ∣V∣|V| scores, one per word. The softmax function normalizes these scores into a probability distribution. Optionally, a direct connection from xx to the output adds a skip path: y=b+Wx+Utanh⁡(Hx+d)y = b + Wx + U\tanh(Hx+d).

Open in Lab
Step through the four stages of the architecture. Watch how each context word's embedding flows through the network to produce a probability distribution.
The demo wakes as you arrive…
P(wt∣wt−1,…,wt−n+1)=eywt∑i=1∣V∣eyiP(w_t \mid w_{t-1}, \ldots, w_{t-n+1}) = \frac{e^{y_{w_t}}}{\sum_{i=1}^{|V|} e^{y_i}}
Softmax probability for the next word — The output yy is a vector of raw scores (logits), one per vocabulary word. Softmax exponentiates each score, then divides by the sum, giving a valid probability distribution. The word with the highest probability is the model's best guess for the next word.
y=b+Wx+Utanh⁡(Hx+d)y = b + Wx + U\tanh(Hx + d)
The full network output (with optional direct connection) — The model first transforms the input embeddings through a hidden layer to learn useful patterns and relationships. These hidden representations are then converted into scores for all possible output words. An additional direct connection allows information from the original embeddings to influence the output without passing through the hidden layer, improving information flow and making learning easier.

Training: everything learns together

The training objective is to maximize the log-likelihood of the training corpus. For each position tt in the text, the model computes log⁡P(wt∣wt−1,…,wt−n+1)\log P(w_t \mid w_{t-1}, \ldots, w_{t-n+1}) and the total loss is the average plus a regularizer:

The critical detail: the gradients flow back through the output layer, through the hidden layer, and all the way into the embedding matrix CC. Every training example doesn't just improve the network weights — it also adjusts the feature vectors of the context words. Words that appear in similar contexts gradually move closer together in the embedding space. This is joint learning: the representations and the prediction function evolve together.

In the 2003 experiments, training used on the Brown corpus (1.18M words, vocabulary 16,383) and AP News (14M words, vocabulary 17,964). A context window of n=5n=5 with embedding dimension m=30–60m=30\text{–}60 and hidden size h=50–100h=50\text{–}100 was typical. Training took days on the hardware of the era — the softmax bottleneck (computing scores for every vocabulary word at every step) consumed 99.7% of the computation.

L=−1T∑t=1Tlog⁡P(wt∣wt−1,…,wt−n+1)+λ∥θ∥2L = -\frac{1}{T}\sum_{t=1}^{T}\log P(w_t \mid w_{t-1}, \ldots, w_{t-n+1}) + \lambda \|\theta\|^2
Training objective — regularized negative log-likelihood — The objective combines two goals. First, it encourages the model to assign high probability to the correct next word in the training data. Second, it penalizes excessively large parameter values through regularization. This helps the model generalize better to unseen text and reduces the risk of overfitting.

Why it works: the geometry of generalization

The magic of the model is in the count. An n-gram table for a trigram model with ∣V∣=100,000|V| = 100{,}000 needs up to 101510^{15} entries. Bengio's model needs roughly ∣V∣×m+(n−1)⋅m⋅h+h⋅∣V∣|V| \times m + (n{-}1) \cdot m \cdot h + h \cdot |V| parameters. In the AP News experiment (∣V∣=17,964|V| = 17{,}964, m=60m = 60, h=100h = 100, n=6n = 6) that's about 2.9 million — orders of magnitude fewer.

But it's not just about fewer parameters. The continuous embedding space creates a geometric structure where generalization is built in. When the model sees "The cat sat on the mat", the gradient adjusts the embeddings of "cat" so that contexts containing "cat" become more likely. If "dog" already has a similar embedding, this same probability mass bleeds over to contexts with "dog" — without any explicit rule. This is the transfer of knowledge through the geometry of the embedding space.

The paper reported improvements of 10–20% over state-of-the-art smoothed trigram models on both Brown corpus and AP News. More importantly, interpolating the neural model with a trigram gave even better results, suggesting the two models capture complementary patterns.

Open in Lab
Explore how the output layer dominates computation. Adjust vocabulary size and hidden dimension to see the parameter explosion.
The demo wakes as you arrive…

The model in code

Bengio's Neural Language Model in PyTorchpython

Simplified to show the idea — not the real implementation.

import torch import torch.nn as nn
class NeuralLM(nn.Module):
    def __init__(self, vocab_size, embed_dim, context_len, hidden_dim):
        super().__init__()
        # Stage 1: shared embedding lookup table C
        self.C = nn.Embedding(vocab_size, embed_dim)
        # Stage 3: hidden layer H
        self.hidden = nn.Linear(context_len * embed_dim, hidden_dim)
        # Stage 4: output layer U
        self.output = nn.Linear(hidden_dim, vocab_size)
        # Optional direct connection W
        self.direct = nn.Linear(context_len * embed_dim, vocab_size, bias=False)

    def forward(self, context_ids):
        # Stage 1: lookup embeddings
        x = self.C(context_ids)              # (batch, n-1, m)
        # Stage 2: concatenate
        x = x.view(x.size(0), -1)            # (batch, (n-1)*m)
        # Stage 3: hidden layer with tanh
        h = torch.tanh(self.hidden(x))        # (batch, h)
        # Stage 4: output scores + direct path
        logits = self.output(h) + self.direct(x)  # (batch, |V|)
        return logits  # apply softmax or cross-entropy loss

The legacy: from embeddings to everything

The paper's influence radiates outward in concentric circles. It established three ideas that pervade modern AI:

Learned representations over engineered features. Before Bengio, NLP features were hand-crafted (POS tags, parse trees, lexicons). After Bengio, the model learns its own features — and they turn out to be better. This principle now drives every system.

Joint training of representations and task. The embedding table and the prediction network share a loss function. This idea of end-to-end learning, where the representation is shaped by what it needs to predict, is the foundation of GPT and every modern language model.

Continuous spaces enable combinatorial generalization. A discrete symbol can only be itself. A point in a continuous space inherits properties from its neighbors. This geometric insight is why Transformers generalize — every head operates in a continuous space descended from Bengio's embedding layer.

  1. 2003

    Bengio — Neural Probabilistic Language Model

    Introduced learned word embeddings and proved neural networks beat n-grams at language modeling. Planted the seed of representation learning in NLP.

  2. 2008

    Collobert & Weston — Unified NLP Architecture

    Showed that a single neural architecture with pre-trained embeddings could handle multiple NLP tasks (POS, NER, chunking, SRL) — extending Bengio's embeddings to multi-task learning.

  3. 2013

    Word2Vec — Embeddings at scale

    Mikolov et al. stripped Bengio's architecture to its essence (no hidden layer, or a shallow one) and trained on billions of words. Made word embeddings a practical tool for every NLP pipeline.

  4. 2014

    Seq2Seq — Embeddings power translation

    Sutskever, Vinyals & Le used learned embeddings inside encoder-decoder RNNs for machine translation — the direct descendant of Bengio's joint-training idea.

  5. 2018

    ULMFiT — Embeddings become pre-training

    Howard & Ruder showed that a language model pre-trained on general text could be fine-tuned for classification with very little labeled data — the pre-train/fine-tune paradigm born from Bengio's seed.

  6. 2017

    Transformer — Embeddings meet attention

    Every Transformer starts with an embedding layer mapping tokens to dense vectors — the same idea Bengio introduced in 2003, now feeding into self-attention instead of a feedforward hidden layer.

Every modern language model — GPT, Claude, Gemini, LLaMA — begins with a learned embedding table. That table is Bengio's matrix CC, scaled up. The 2003 paper didn't just propose a model. It proposed a way of thinking: that the right representation, learned from data, is more powerful than any amount of manual engineering.

CitationBengio, Ducharme, Vincent, Jauvin. A Neural Probabilistic Language Model. JMLR, 2003.

Terms in this paper