Language Models2017intermediate6 min read

Attention Is All You Need

الانتباه هو كل ما تحتاجه

Vaswani, A. · Shazeer, N. · Parmar, N. · Uszkoreit, J. · Jones, L. · Gomez, A. · Kaiser, Ł. · Polosukhin, I. — NeurIPS

The problem

Recurrent models (LSTMs) process tokens strictly in order — impossible to parallelize on GPUs — and connect distant words only through a long chain of steps, through which information and gradients decay. Translation quality plateaued and training took weeks.

The contribution

The : an architecture with no recurrence at all. connects every pair of positions in a single step — each token builds a , compares it against every token's , and gathers a weighted mix of their values. runs several such 'searches' in parallel; positional encodings restore word order. Every position is computed simultaneously, so training parallelizes perfectly.

The impact

The single most consequential architecture of modern AI. GPT, Claude, Gemini, BERT, Vision Transformers, AlphaFold, Whisper — all are Transformers. Its is what made trillion-token training runs, and therefore large language models, physically possible.

An reading a sentence is a game of telephone: each word whispers a compressed summary to the next, and by word 100 the first word's contribution has decayed to a rumor.

The Transformer replaces the telephone line with a meeting room: every word sits at the same table and can address any other word directly.

The word "it" doesn't hope the summary still contains its referent — it asks the whole table, "who am I about?", and "the animal" answers loudest.

The problem: sequential chains are slow and forgetful

By 2017, LSTM-based translation had two walls it couldn't break:

  • No parallelism. Step tt needs step t−1t-1's output. GPUs — machines built to do thousands of things at once — sat mostly idle. Training big models took weeks.
  • Long paths between related words. In "The animal didn't cross the street because it was too tired", connecting it to animal takes many recurrent steps; each step is a chance to forget (you saw this decay in the LSTM chapter).
Open in Lab
Press "Race!" to watch the Transformer finish while LSTM is still on token 1.
The demo wakes as you arrive…

The idea: every word queries every word

Give each word three learned vectors, playing three roles:

  • a Query — what am I looking for?
  • a Key — what can I be found by?
  • a — what do I contribute if chosen?

A word scores its query against every key, softmaxes the scores into weights that sum to 1, and takes the weighted average of the values. That average — a blend of the words it chose to attend to — becomes its new .

Attention(Q,K,V)=softmax(QKT/√dk)VAttention(Q, K, V) = softmax(QKᵀ / √d_k) V
Scaled dot-product attention — the whole engine — QKᵀ = all query·key match scores at once · √d_k = keeps scores in softmax's sensitive range · the result: each row is a word's new, context-aware representation

Read it as a library visit: your query is the request slip, each book's key is its index card, matching gives you relevance scores, and you walk out with a weighted armful of the books' values.

Open in Lab
Click a token, then step through the 5 stages of self-attention.
The demo wakes as you arrive…

Why divide by √d_k?

Dot products of high-dimensional vectors grow large, pushing into a regime where one weight ≈ 1 and the rest ≈ 0 — and gradients there are nearly zero (a vanishing-gradient echo of the LSTM chapter). Scaling by √d_k keeps the scores in softmax's responsive range.

Open in Lab
Drag d_k to 256+ and watch the left (unscaled) bars collapse to one-hot.
The demo wakes as you arrive…

See attention in action

Below is a simulated heatmap for the sentence "The animal didn't cross the street because it was tired." Try switching heads — each one learns different relationship patterns. Notice how "it" strongly attends to "animal": the Transformer resolves coreference in a single step, with no chain for the signal to decay through.

Open in Lab
The demo wakes as you arrive…

The supporting cast: multi-head attention & positional encoding

Two supporting ideas complete the architecture:

Multi-head attention — run 8 attention 'searches' in parallel with different learned projections: one head may track grammar, another coreference, another adjacency. Concatenate the 8 answers and project back to the original dimension. It's like asking 8 different librarians the same question — each brings back different books, and together they give you a richer answer than any single search.

Positional encodings — attention itself is order-blind (a set operation), so "dog bites man" and "man bites dog" look the same to it. A sinusoidal position signature is added to each to restore word order. The waves at different frequencies give each position a unique fingerprint that the model can learn to decode.

Open in Lab
Each row is a word position. Notice how the wave patterns make each row unique.
The demo wakes as you arrive…

Putting it all together

The Transformer stacks these pieces into an block that repeats 6 times. Each block has the same pattern: multi-head self-attention → add & normalize → feed-forward network → add & normalize. The "add" is a (it adds the input back), which keeps gradients flowing through the deep stack.

Click any layer below to see what it does:

Open in Lab
The demo wakes as you arrive…

The same idea in code

Scaled dot-product attention, completepython

Simplified to show the idea — not the real implementation.

import numpy as np

def softmax(x):
    e = np.exp(x - x.max(axis=-1, keepdims=True))   # stable softmax
    return e / e.sum(axis=-1, keepdims=True)

def attention(Q, K, V):
    """Q, K, V: (n_words, d) matrices — one row per word."""
    d_k = K.shape[-1]
    scores = Q @ K.T / np.sqrt(d_k)   # every word scores every word: (n, n)
    weights = softmax(scores)         # each row sums to 1: "where do I look?"
    return weights @ V                # each word = weighted mix of all values

def multi_head_attention(X, W_q, W_k, W_v, W_o, n_heads=8):
    """Run n_heads parallel attention searches, concatenate results."""
    d = X.shape[-1]
    head_dim = d // n_heads
    heads = []
    for h in range(n_heads):
        Q = X @ W_q[h]   # (n, head_dim)
        K = X @ W_k[h]
        V = X @ W_v[h]
        heads.append(attention(Q, K, V))
    return np.concatenate(heads, axis=-1) @ W_o   # (n, d)

# In self-attention, Q, K, V are three linear projections of the SAME
# word embeddings: X @ W_q, X @ W_k, X @ W_v. The W's are learned.
# That's it. GPT, Claude, and Gemini run this loop billions of times.

Why it mattered

  1. 2018

    GPT-1

    The Transformer's decoder stack, pre-trained to predict the next word then fine-tuned on downstream tasks. Proved that unsupervised pre-training transfers.

  2. 2018

    BERT

    The encoder stack, trained to fill in masked blanks bidirectionally. Dominated every NLP benchmark and established the pre-train → fine-tune paradigm.

  3. 2020

    GPT-3 — 175B parameters

    Scaling unlocked emergent abilities — translation, arithmetic, coding — that nobody programmed. Demonstrated that size alone is a research variable.

  4. 2020

    Vision Transformer (ViT)

    Proved Transformers conquer images too — patches of pixels treated as tokens, self-attention replacing convolution. Challenged two decades of CNN dominance.

  5. 2021

    AlphaFold 2

    Transformers solved protein folding — a 50-year grand challenge in biology — with accuracy matching experimental methods. Science's first Transformer-driven breakthrough.

  6. 2022

    ChatGPT

    Transformers went mainstream. Instruction fine-tuning and RLHF turned a language model into a conversational product used by 100 million people in two months.

  7. 2023

    Claude

    Constitutional AI built on Transformers — training with a written set of principles rather than purely human ratings, aiming for safer and more steerable models.

  8. 2024

    DeepSeek

    Transformers combined with Mixture of Experts — only a fraction of parameters activate per token, dramatically cutting compute while scaling capacity.

  9. 2026

    Qwen

    A model that excels at heavy software engineering and multi-step reasoning tasks, pushing the frontier of what Transformer-based agents can autonomously accomplish.

GPT is the Transformer's stack; BERT its encoder stack. You are reading this because countless researchers added their lines to a story larger than themselves, extending the frontier one idea at a time.

CitationVaswani, Shazeer, Parmar, Uszkoreit, Jones, Gomez, Kaiser, Polosukhin. Attention Is All You Need. NeurIPS, 2017.

Terms in this paper