Language Models2017intermediate6 min read
Attention Is All You Need
الانتباه هو كل ما تحتاجه
Vaswani, A. · Shazeer, N. · Parmar, N. · Uszkoreit, J. · Jones, L. · Gomez, A. · Kaiser, Ł. · Polosukhin, I. — NeurIPS
The problem
Recurrent models (LSTMs) process tokens strictly in order — impossible to parallelize on GPUs — and connect distant words only through a long chain of steps, through which information and gradients decay. Translation quality plateaued and training took weeks.
The contribution
The : an architecture with no recurrence at all. connects every pair of positions in a single step — each token builds a , compares it against every token's , and gathers a weighted mix of their values. runs several such 'searches' in parallel; positional encodings restore word order. Every position is computed simultaneously, so training parallelizes perfectly.
The impact
The single most consequential architecture of modern AI. GPT, Claude, Gemini, BERT, Vision Transformers, AlphaFold, Whisper — all are Transformers. Its is what made trillion-token training runs, and therefore large language models, physically possible.
An reading a sentence is a game of telephone: each word whispers a compressed summary to the next, and by word 100 the first word's contribution has decayed to a rumor.
The Transformer replaces the telephone line with a meeting room: every word sits at the same table and can address any other word directly.
The word "it" doesn't hope the summary still contains its referent — it asks the whole table, "who am I about?", and "the animal" answers loudest.
The problem: sequential chains are slow and forgetful
By 2017, LSTM-based translation had two walls it couldn't break:
- No parallelism. Step needs step 's output. GPUs — machines built to do thousands of things at once — sat mostly idle. Training big models took weeks.
- Long paths between related words. In "The animal didn't cross the street because it was too tired", connecting it to animal takes many recurrent steps; each step is a chance to forget (you saw this decay in the LSTM chapter).
The idea: every word queries every word
Give each word three learned vectors, playing three roles:
- a Query — what am I looking for?
- a Key — what can I be found by?
- a — what do I contribute if chosen?
A word scores its query against every key, softmaxes the scores into weights that sum to 1, and takes the weighted average of the values. That average — a blend of the words it chose to attend to — becomes its new .
Read it as a library visit: your query is the request slip, each book's key is its index card, matching gives you relevance scores, and you walk out with a weighted armful of the books' values.
Why divide by √d_k?
Dot products of high-dimensional vectors grow large, pushing into a regime where one weight ≈ 1 and the rest ≈ 0 — and gradients there are nearly zero (a vanishing-gradient echo of the LSTM chapter). Scaling by √d_k keeps the scores in softmax's responsive range.
See attention in action
Below is a simulated heatmap for the sentence "The animal didn't cross the street because it was tired." Try switching heads — each one learns different relationship patterns. Notice how "it" strongly attends to "animal": the Transformer resolves coreference in a single step, with no chain for the signal to decay through.
The supporting cast: multi-head attention & positional encoding
Two supporting ideas complete the architecture:
Multi-head attention — run 8 attention 'searches' in parallel with different learned projections: one head may track grammar, another coreference, another adjacency. Concatenate the 8 answers and project back to the original dimension. It's like asking 8 different librarians the same question — each brings back different books, and together they give you a richer answer than any single search.
Positional encodings — attention itself is order-blind (a set operation), so "dog bites man" and "man bites dog" look the same to it. A sinusoidal position signature is added to each to restore word order. The waves at different frequencies give each position a unique fingerprint that the model can learn to decode.
Putting it all together
The Transformer stacks these pieces into an block that repeats 6 times. Each block has the same pattern: multi-head self-attention → add & normalize → feed-forward network → add & normalize. The "add" is a (it adds the input back), which keeps gradients flowing through the deep stack.
Click any layer below to see what it does:
The same idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
def softmax(x):
e = np.exp(x - x.max(axis=-1, keepdims=True)) # stable softmax
return e / e.sum(axis=-1, keepdims=True)
def attention(Q, K, V):
"""Q, K, V: (n_words, d) matrices — one row per word."""
d_k = K.shape[-1]
scores = Q @ K.T / np.sqrt(d_k) # every word scores every word: (n, n)
weights = softmax(scores) # each row sums to 1: "where do I look?"
return weights @ V # each word = weighted mix of all values
def multi_head_attention(X, W_q, W_k, W_v, W_o, n_heads=8):
"""Run n_heads parallel attention searches, concatenate results."""
d = X.shape[-1]
head_dim = d // n_heads
heads = []
for h in range(n_heads):
Q = X @ W_q[h] # (n, head_dim)
K = X @ W_k[h]
V = X @ W_v[h]
heads.append(attention(Q, K, V))
return np.concatenate(heads, axis=-1) @ W_o # (n, d)
# In self-attention, Q, K, V are three linear projections of the SAME
# word embeddings: X @ W_q, X @ W_k, X @ W_v. The W's are learned.
# That's it. GPT, Claude, and Gemini run this loop billions of times.Why it mattered
2018
GPT-1
The Transformer's decoder stack, pre-trained to predict the next word then fine-tuned on downstream tasks. Proved that unsupervised pre-training transfers.
2018
BERT
The encoder stack, trained to fill in masked blanks bidirectionally. Dominated every NLP benchmark and established the pre-train → fine-tune paradigm.
2020
GPT-3 — 175B parameters
Scaling unlocked emergent abilities — translation, arithmetic, coding — that nobody programmed. Demonstrated that size alone is a research variable.
2020
Vision Transformer (ViT)
Proved Transformers conquer images too — patches of pixels treated as tokens, self-attention replacing convolution. Challenged two decades of CNN dominance.
2021
AlphaFold 2
Transformers solved protein folding — a 50-year grand challenge in biology — with accuracy matching experimental methods. Science's first Transformer-driven breakthrough.
2022
ChatGPT
Transformers went mainstream. Instruction fine-tuning and RLHF turned a language model into a conversational product used by 100 million people in two months.
2023
Claude
Constitutional AI built on Transformers — training with a written set of principles rather than purely human ratings, aiming for safer and more steerable models.
2024
DeepSeek
Transformers combined with Mixture of Experts — only a fraction of parameters activate per token, dramatically cutting compute while scaling capacity.
2026
Qwen
A model that excels at heavy software engineering and multi-step reasoning tasks, pushing the frontier of what Transformer-based agents can autonomously accomplish.
GPT is the Transformer's stack; BERT its encoder stack. You are reading this because countless researchers added their lines to a story larger than themselves, extending the frontier one idea at a time.
CitationVaswani, Shazeer, Parmar, Uszkoreit, Jones, Gomez, Kaiser, Polosukhin. Attention Is All You Need. NeurIPS, 2017.
Terms in this paper
- Transformerالمحوِّل
- Self-Attentionالانتباه الذاتي
- Multi-Head Attentionالانتباه المتعدد المسارات
- Queryمصفوفة الطلب (الاستعلام)
- Keyمصفوفة المفتاح (الاستدلال)
- Valueمصفوفة القيمة (المحتوى)
- Positional Encodingالترميز الموضعي