Computational Biology2021advanced13 min read
Highly Accurate Protein Structure Prediction with AlphaFold
التنبؤ الدقيق ببنية البروتينات باستخدام ألفافولد
Jumper, J. · Evans, R. · Pritzel, A. · Green, T. · Figurnov, M. · Ronneberger, O. · Tunyasuvunakool, K. · Bates, R. · Žídek, A. · Potapenko, A. · Bridgland, A. · Meyer, C. · Kohl, S. A. A. · Ballard, A. J. · Cowie, A. · Romera-Paredes, B. · Nikolov, S. · Jain, R. · Adler, J. · Back, T. · Petersen, S. · Reiman, D. · Clancy, E. · Zielinski, M. · Steinegger, M. · Pacholska, M. · Berghammer, T. · Bodenstein, S. · Silver, D. · Vinyals, O. · Senior, A. W. · Kavukcuoglu, K. · Kohli, P. · Hassabis, D. — Nature
The problem
Predicting the 3D structure a protein will adopt from its sequence alone — the problem — had been a grand challenge in biology for over 50 years. Existing computational methods fell far short of atomic accuracy, especially for proteins without known homologous structures. Experimental methods like X-ray crystallography and cryo-EM are accurate but take months to years per protein, while the human body contains over 20,000 different proteins.
The contribution
AlphaFold introduces an end-to-end neural network that predicts protein structure with atomic accuracy. Its core is the — a novel -based architecture that jointly reasons over evolutionary sequences (MSA) and pairwise relationships using triangular updates. A then converts these abstract representations into 3D coordinates using (IPA), which is equivariant to rotations and translations. The system uses to iteratively refine its predictions, trained with a novel Frame Aligned Point Error () .
The impact
AlphaFold solved a 50-year grand challenge in biology, achieving accuracy competitive with experimental methods at CASP14. DeepMind subsequently released predicted structures for over 200 million proteins — nearly every known protein — accelerating drug discovery, enzyme engineering, and our understanding of life at the molecular level. The work earned Demis Hassabis and John Jumper the 2024 Nobel Prize in Chemistry.
A protein's amino acid sequence is like a long strip of origami paper with fold instructions written at each crease. The paper could fold into millions of possible shapes, but only one is the "right" one — the one that nature actually makes.
For 50 years, scientists tried to fold the paper by simulating every physical force. AlphaFold takes a different approach: it looks at thousands of similar origami strips from related species (an MSA), notices which creases always fold together, then uses two "workbenches" — one tracking what each crease knows about every other crease, and another tracking which crease pairs tend to end up close — iterating until the 3D shape emerges.
The problem: how does a chain of amino acids become a 3D machine?
Proteins are the molecular machines of life. A protein starts as a linear chain of amino acids — there are 20 types, each with different chemical properties. Within milliseconds of being synthesized, this chain folds into a precise 3D shape that determines its function: hemoglobin carries oxygen because of its shape, antibodies recognize viruses because of theirs.
The sequence-to-structure mapping is called the protein folding problem. Experimental methods (X-ray crystallography, cryo-EM, NMR) can determine structure, but each takes months and costs tens of thousands of dollars. By 2020, only ~170,000 structures had been experimentally determined — a tiny fraction of the billions of proteins in nature.
(Critical Assessment of protein Structure Prediction) is a biennial blind competition where teams predict structures of proteins whose shapes have been determined but not yet published. In CASP14 (2020), AlphaFold achieved a median GDT score of 92.4 — for the first time matching experimental accuracy. The second-best method scored 67.
The big picture: from sequence to structure in three stages
AlphaFold's architecture has three main stages, and understanding this pipeline is key to understanding the paper:
Stage 1 — Input . The amino acid sequence is searched against genetic databases to build a (MSA) — a collection of related sequences from other organisms. These evolutionary relatives reveal which positions co-evolve (change together), hinting at 3D proximity. Two initial representations are created: an (capturing per-sequence, per-residue features) and a (capturing relationships between every pair of residues).
Stage 2 — Evoformer. 48 blocks of a novel attention architecture jointly refine both representations. The MSA track uses row-wise and column-wise attention; the pair track uses triangular updates and triangular attention. Information flows continuously between the two tracks. This is where the network reasons about which residues are close in 3D space.
Stage 3 — Structure module. Converts the refined representations into actual 3D coordinates. Each residue gets a local coordinate frame (rotation + translation), refined by Invariant Point Attention (IPA). The entire pipeline is recycled 3 times, with each pass refining the previous prediction.
Inputs: learning from evolution
The key insight driving AlphaFold is . When two amino acid positions are in contact in the 3D structure, a mutation at one position creates evolutionary pressure for a compensating mutation at the other — just like two puzzle pieces that must fit together. If one piece changes shape, the other must change too.
To capture this, AlphaFold builds a Multiple Sequence Alignment: it searches databases like UniRef90 and BFD for sequences similar to the target protein, aligning them column by column. Each column represents the same structural position across species. Columns that show correlated mutations are likely in 3D contact.
From this, two representations are initialized:
-
The MSA representation has shape (sequences × residues × features). Think of it as a spreadsheet where each row is a related species and each column is a position in the protein. Each cell holds a vector describing that residue in that sequence.
-
The pair representation has shape (residues × residues × features). Think of it as a relationship map: cell holds everything the network currently believes about the spatial relationship between residue and residue .
The Evoformer: where evolution meets geometry
The Evoformer is the heart of AlphaFold — 48 stacked blocks that iteratively refine both representations. Each block contains two interleaved tracks, one for the MSA representation and one for the pair representation, with information flowing between them. Think of it as two teams of analysts working side by side: one team studies the evolutionary evidence (MSA), the other studies spatial relationships (pairs), and they constantly share notes.
MSA track — Row-wise attention with pair bias: For each sequence in the MSA, attention is computed across residue positions (the columns). This lets the network learn which positions in the sequence are related. Crucially, the pair representation is injected as a bias into this attention — so what we already know about residue proximity directly influences how the MSA is read.
MSA track — Column-wise attention: Attention is computed down each column of the MSA — across different species at the same position. This lets the network compare how different organisms handle the same position, learning conservation patterns.
— MSA to pair communication: The MSA representation is summarized into the pair representation via an outer product operation. For each pair of positions , the network takes the mean outer product of their MSA column features across all sequences. This converts evolutionary covariance into spatial information.
Triangular updates: enforcing geometric consistency
The pair representation encodes pairwise distances — but distances in 3D space must satisfy the : if residue A is close to B, and B is close to C, then A must be reasonably close to C. Standard attention has no mechanism to enforce this.
AlphaFold introduces triangular multiplicative updates — a novel operation designed specifically for this constraint. The intuition: for every edge in the pair representation, look at all intermediate residues that form triangles . The edges and provide information about what edge should look like.
There are two variants: outgoing edges update using and — both starting from the same node; incoming edges update using and — both ending at the same node. These are followed by triangular , where the attention bias for updating row of the pair representation comes from column information of the same pair matrix.
The structure module: from abstract features to 3D atoms
After 48 Evoformer blocks, the representations are rich with structural information — but they're still abstract feature vectors, not 3D coordinates. The structure module bridges this gap.
Each residue is represented by a — a rotation matrix and translation vector that define the position and orientation of its backbone in 3D space. Initially, all frames are at the origin. The structure module's 8 blocks (with shared weights) iteratively refine these frames.
The core mechanism is Invariant Point Attention (IPA) — a novel attention mechanism designed for 3D structure. Standard attention computes similarity between abstract feature vectors. IPA additionally computes similarity between 3D points — each query and key generate points in their local coordinate frames, and the attention scores include the distances between these points. The crucial property: IPA is invariant to global rotations and translations. Moving or rotating the entire protein doesn't change the attention scores, because the distances are computed in local frames.
After IPA updates the single representation, a small network predicts a rotation and translation update for each residue's frame. The side-chain positions are then computed from predicted torsion angles.
Recycling: iterative refinement
A single pass through the network produces a reasonable structure, but not the best one. AlphaFold uses recycling — running the entire pipeline 3 times, feeding the previous output back as additional input. Think of it like a sculptor: the first pass roughs out the shape, the second refines proportions, the third polishes details.
Specifically, three things are recycled: (1) the predicted backbone coordinates from the structure module, (2) the output pair representation from the Evoformer, and (3) the first row of the MSA representation. This recycling is remarkably efficient — it lets the network act as if it were 3× deeper without 3× the parameters, because the weights are shared across cycles.
Training: the FAPE loss
How do you teach a network to predict 3D structures? You need a loss function that compares predicted and true atom positions — but there's a subtlety. If the predicted structure is simply rotated or translated relative to the ground truth, a naive loss would penalize this even though the structure is correct.
AlphaFold introduces Frame Aligned Point Error (FAPE): for each residue's local frame, transform all atoms into that frame and compute the error. By averaging over all frames as reference points, FAPE is invariant to global rotations and translations — it measures structural accuracy regardless of orientation. This is more informative than traditional RMSD because it captures local structural quality (each residue judges the structure from its own viewpoint), not just global alignment.
Knowing what you know: confidence estimation
A prediction is only useful if you know how much to trust it. AlphaFold predicts a per-residue confidence score called (predicted Local Distance Difference Test), ranging from 0 to 100. Regions with pLDDT > 90 are typically accurate to within 1 Å of the experimental structure; regions below 50 are likely disordered (no fixed structure in nature).
AlphaFold also provides a Predicted Aligned Error (PAE) matrix — a residue × residue map showing the expected position error of residue when the prediction is aligned on residue . This reveals domain boundaries: within a rigid domain, PAE is low everywhere; between independently-moving domains, PAE is high.
The key ideas in code
Simplified to show the idea — not the real implementation.
import numpy as np
def softmax(x, axis=-1):
e = np.exp(x - x.max(axis=axis, keepdims=True))
return e / e.sum(axis=axis, keepdims=True)
def row_attention_with_pair_bias(msa, pair_rep, W_q, W_k, W_v):
"""
Row-wise attention on MSA, biased by pair representation.
msa: (n_seq, n_res, d_msa) — MSA features
pair_rep: (n_res, n_res, d_pair) — pair features
"""
n_seq, n_res, d = msa.shape
for s in range(n_seq): # each sequence independently
Q = msa[s] @ W_q # (n_res, d_head)
K = msa[s] @ W_k
V = msa[s] @ W_v
scores = Q @ K.T / np.sqrt(d) # (n_res, n_res)
# Inject pair representation as attention bias
pair_bias = pair_rep.mean(axis=-1) # (n_res, n_res)
scores = scores + pair_bias # spatial info guides attention
weights = softmax(scores)
msa[s] = weights @ V
return msa
def outer_product_mean(msa, W_a, W_b):
"""
Convert MSA covariance into pair representation updates.
For each pair (i, j): average the outer product of their MSA columns.
"""
n_seq, n_res, _ = msa.shape
a = msa @ W_a # (n_seq, n_res, c)
b = msa @ W_b # (n_seq, n_res, c)
# outer product for each pair, averaged over sequences
pair_update = np.einsum('ski,skj->ij', a, b) / n_seq
return pair_update # (n_res, n_res) — co-evolution → spatial info
# This is the core loop: MSA reads spatial clues from pairs,
# then MSA covariance updates pairs. 48 blocks of this back-and-forth
# is what makes AlphaFold reason about protein structure.Why it worked — the design principles
The ripple effect
2018
AlphaFold 1 — CASP13
First version uses distance prediction + gradient descent to fold proteins. Wins CASP13 but accuracy still far from experimental.
2020
AlphaFold 2 — CASP14
A completely redesigned architecture with the Evoformer and structure module achieves experimental-level accuracy. Median GDT = 92.4, solving the protein folding problem.
2021
Open-source release + 350K structures
DeepMind publishes the code and releases predictions for the entire human proteome and 20 other organisms via the AlphaFold Protein Structure Database.
2022
200 million structures
The database expands to cover nearly every known protein in nature — over 200 million structures — revolutionizing structural biology overnight.
2024
AlphaFold 3 + Nobel Prize
AlphaFold 3 extends to predict interactions between proteins, DNA, RNA, and small molecules. Demis Hassabis and John Jumper awarded the Nobel Prize in Chemistry.
AlphaFold didn't just solve a computational problem — it gave biology an entirely new tool. Drug designers use it to find binding sites. Enzyme engineers use it to design proteins that don't exist in nature. Evolutionary biologists use it to study proteins too ancient for experimental determination. The architecture that powers modern language models also powers the prediction of life's molecular machinery.
CitationJumper, Evans, Pritzel, Green, Figurnov, Ronneberger, et al.. Highly Accurate Protein Structure Prediction with AlphaFold. Nature, 2021.
Terms in this paper
- Protein Foldingطيّ البروتين
- Multiple Sequence Alignmentمحاذاة التسلسلات المتعددة
- Co-evolutionالتطور المشترك
- Pair Representationالتمثيل الثنائي
- MSA Representationتمثيل محاذاة التسلسلات المتعددة
- Evoformerالإيفوفورمر
- Triangular Multiplicative Updateالتحديث المثلثي الضربيّ
- Invariant Point Attentionانتباه النقاط الثابتة
- FAPEخطأ النقاط المحاذية للإطار
- pLDDTدرجة الثقة المحلية المتنبأة
- Rigid Body Frameإطار الجسم الجاسئ
- Recyclingإعادة التدوير
- Outer Product Meanمتوسط الجداء الخارجي
- Structure Moduleوحدة البنية
- Torsion Angleزاوية الالتواء