Computational Biology2021advanced13 min read

Highly Accurate Protein Structure Prediction with AlphaFold

التنبؤ الدقيق ببنية البروتينات باستخدام ألفافولد

Jumper, J. · Evans, R. · Pritzel, A. · Green, T. · Figurnov, M. · Ronneberger, O. · Tunyasuvunakool, K. · Bates, R. · Žídek, A. · Potapenko, A. · Bridgland, A. · Meyer, C. · Kohl, S. A. A. · Ballard, A. J. · Cowie, A. · Romera-Paredes, B. · Nikolov, S. · Jain, R. · Adler, J. · Back, T. · Petersen, S. · Reiman, D. · Clancy, E. · Zielinski, M. · Steinegger, M. · Pacholska, M. · Berghammer, T. · Bodenstein, S. · Silver, D. · Vinyals, O. · Senior, A. W. · Kavukcuoglu, K. · Kohli, P. · Hassabis, D. — Nature

The problem

Predicting the 3D structure a protein will adopt from its sequence alone — the problem — had been a grand challenge in biology for over 50 years. Existing computational methods fell far short of atomic accuracy, especially for proteins without known homologous structures. Experimental methods like X-ray crystallography and cryo-EM are accurate but take months to years per protein, while the human body contains over 20,000 different proteins.

The contribution

AlphaFold introduces an end-to-end neural network that predicts protein structure with atomic accuracy. Its core is the — a novel -based architecture that jointly reasons over evolutionary sequences (MSA) and pairwise relationships using triangular updates. A then converts these abstract representations into 3D coordinates using (IPA), which is equivariant to rotations and translations. The system uses to iteratively refine its predictions, trained with a novel Frame Aligned Point Error () .

The impact

AlphaFold solved a 50-year grand challenge in biology, achieving accuracy competitive with experimental methods at CASP14. DeepMind subsequently released predicted structures for over 200 million proteins — nearly every known protein — accelerating drug discovery, enzyme engineering, and our understanding of life at the molecular level. The work earned Demis Hassabis and John Jumper the 2024 Nobel Prize in Chemistry.

A protein's amino acid sequence is like a long strip of origami paper with fold instructions written at each crease. The paper could fold into millions of possible shapes, but only one is the "right" one — the one that nature actually makes.

For 50 years, scientists tried to fold the paper by simulating every physical force. AlphaFold takes a different approach: it looks at thousands of similar origami strips from related species (an MSA), notices which creases always fold together, then uses two "workbenches" — one tracking what each crease knows about every other crease, and another tracking which crease pairs tend to end up close — iterating until the 3D shape emerges.

The problem: how does a chain of amino acids become a 3D machine?

Proteins are the molecular machines of life. A protein starts as a linear chain of amino acids — there are 20 types, each with different chemical properties. Within milliseconds of being synthesized, this chain folds into a precise 3D shape that determines its function: hemoglobin carries oxygen because of its shape, antibodies recognize viruses because of theirs.

The sequence-to-structure mapping is called the protein folding problem. Experimental methods (X-ray crystallography, cryo-EM, NMR) can determine structure, but each takes months and costs tens of thousands of dollars. By 2020, only ~170,000 structures had been experimentally determined — a tiny fraction of the billions of proteins in nature.

(Critical Assessment of protein Structure Prediction) is a biennial blind competition where teams predict structures of proteins whose shapes have been determined but not yet published. In CASP14 (2020), AlphaFold achieved a median GDT score of 92.4 — for the first time matching experimental accuracy. The second-best method scored 67.

Open in Lab
Drag amino acids to see how sequence determines 3D shape. Notice how distant residues can end up close in space.
The demo wakes as you arrive…

The big picture: from sequence to structure in three stages

AlphaFold's architecture has three main stages, and understanding this pipeline is key to understanding the paper:

Stage 1 — Input . The amino acid sequence is searched against genetic databases to build a (MSA) — a collection of related sequences from other organisms. These evolutionary relatives reveal which positions co-evolve (change together), hinting at 3D proximity. Two initial representations are created: an (capturing per-sequence, per-residue features) and a (capturing relationships between every pair of residues).

Stage 2 — Evoformer. 48 blocks of a novel attention architecture jointly refine both representations. The MSA track uses row-wise and column-wise attention; the pair track uses triangular updates and triangular attention. Information flows continuously between the two tracks. This is where the network reasons about which residues are close in 3D space.

Stage 3 — Structure module. Converts the refined representations into actual 3D coordinates. Each residue gets a local coordinate frame (rotation + translation), refined by Invariant Point Attention (IPA). The entire pipeline is recycled 3 times, with each pass refining the previous prediction.

Open in Lab
Click each stage to explore the three main modules of AlphaFold.
The demo wakes as you arrive…

Inputs: learning from evolution

The key insight driving AlphaFold is . When two amino acid positions are in contact in the 3D structure, a mutation at one position creates evolutionary pressure for a compensating mutation at the other — just like two puzzle pieces that must fit together. If one piece changes shape, the other must change too.

To capture this, AlphaFold builds a Multiple Sequence Alignment: it searches databases like UniRef90 and BFD for sequences similar to the target protein, aligning them column by column. Each column represents the same structural position across species. Columns that show correlated mutations are likely in 3D contact.

From this, two representations are initialized:

  • The MSA representation msim_{si} has shape (sequences × residues × features). Think of it as a spreadsheet where each row is a related species and each column is a position in the protein. Each cell holds a vector describing that residue in that sequence.

  • The pair representation zijz_{ij} has shape (residues × residues × features). Think of it as a relationship map: cell (i,j)(i, j) holds everything the network currently believes about the spatial relationship between residue ii and residue jj.

Open in Lab
See how co-evolving columns in the MSA (left) inform the pair representation (right). Correlated mutations reveal 3D contacts.
The demo wakes as you arrive…

The Evoformer: where evolution meets geometry

The Evoformer is the heart of AlphaFold — 48 stacked blocks that iteratively refine both representations. Each block contains two interleaved tracks, one for the MSA representation and one for the pair representation, with information flowing between them. Think of it as two teams of analysts working side by side: one team studies the evolutionary evidence (MSA), the other studies spatial relationships (pairs), and they constantly share notes.

MSA track — Row-wise attention with pair bias: For each sequence in the MSA, attention is computed across residue positions (the columns). This lets the network learn which positions in the sequence are related. Crucially, the pair representation is injected as a bias into this attention — so what we already know about residue proximity directly influences how the MSA is read.

MSA track — Column-wise attention: Attention is computed down each column of the MSA — across different species at the same position. This lets the network compare how different organisms handle the same position, learning conservation patterns.

— MSA to pair communication: The MSA representation is summarized into the pair representation via an outer product operation. For each pair of positions (i,j)(i, j), the network takes the mean outer product of their MSA column features across all sequences. This converts evolutionary covariance into spatial information.

Open in Lab
Click each sub-layer to see how MSA and pair representations talk to each other.
The demo wakes as you arrive…

Triangular updates: enforcing geometric consistency

The pair representation encodes pairwise distances — but distances in 3D space must satisfy the : if residue A is close to B, and B is close to C, then A must be reasonably close to C. Standard attention has no mechanism to enforce this.

AlphaFold introduces triangular multiplicative updates — a novel operation designed specifically for this constraint. The intuition: for every edge (i,j)(i, j) in the pair representation, look at all intermediate residues kk that form triangles (i,k,j)(i, k, j). The edges (i,k)(i, k) and (k,j)(k, j) provide information about what edge (i,j)(i, j) should look like.

There are two variants: outgoing edges update (i,j)(i, j) using (i,k)(i, k) and (j,k)(j, k) — both starting from the same node; incoming edges update (i,j)(i, j) using (k,i)(k, i) and (k,j)(k, j) — both ending at the same node. These are followed by triangular , where the attention bias for updating row ii of the pair representation comes from column information of the same pair matrix.

Open in Lab
Hover over edge (i, j) to see which triangles constrain it. Notice how outgoing and incoming edges provide complementary information.
The demo wakes as you arrive…
zij←zij+∑kaik⊙bjkz_{ij} \leftarrow z_{ij} + \sum_k a_{ik} \odot b_{jk}
Triangular multiplicative update (outgoing edges) — For each pair (i, j), aggregate information from all intermediate nodes k via element-wise products of projected pair features. a and b are gated linear projections of the pair representation.

The structure module: from abstract features to 3D atoms

After 48 Evoformer blocks, the representations are rich with structural information — but they're still abstract feature vectors, not 3D coordinates. The structure module bridges this gap.

Each residue is represented by a — a rotation matrix and translation vector that define the position and orientation of its backbone in 3D space. Initially, all frames are at the origin. The structure module's 8 blocks (with shared weights) iteratively refine these frames.

The core mechanism is Invariant Point Attention (IPA) — a novel attention mechanism designed for 3D structure. Standard attention computes similarity between abstract feature vectors. IPA additionally computes similarity between 3D points — each query and key generate points in their local coordinate frames, and the attention scores include the distances between these points. The crucial property: IPA is invariant to global rotations and translations. Moving or rotating the entire protein doesn't change the attention scores, because the distances are computed in local frames.

After IPA updates the single representation, a small network predicts a rotation and translation update for each residue's frame. The side-chain positions are then computed from predicted torsion angles.

Open in Lab
See how IPA computes attention using both feature similarity and 3D point distances. Rotate the structure to verify invariance.
The demo wakes as you arrive…
aij=softmaxj(wLqiTkj⏟features+wP∑p∥Ti∘q^ip−Tj∘k^jp∥2⏟3D points+bij)a_{ij} = \text{softmax}_j\Big( w_L \underbrace{q_i^T k_j}_{\text{features}} + w_P \underbrace{\sum_p \|T_i \circ \hat{q}_i^p - T_j \circ \hat{k}_j^p\|^2}_{\text{3D points}} + b_{ij}\Big)
Invariant Point Attention — attention weights — Standard feature dot products plus L2 distances between 3D query/key points transformed by local frames T. The L2 norm is invariant to global rigid motion, so the attention scores don't change when you rotate or translate the entire protein.

Recycling: iterative refinement

A single pass through the network produces a reasonable structure, but not the best one. AlphaFold uses recycling — running the entire pipeline 3 times, feeding the previous output back as additional input. Think of it like a sculptor: the first pass roughs out the shape, the second refines proportions, the third polishes details.

Specifically, three things are recycled: (1) the predicted backbone coordinates from the structure module, (2) the output pair representation from the Evoformer, and (3) the first row of the MSA representation. This recycling is remarkably efficient — it lets the network act as if it were 3× deeper without 3× the parameters, because the weights are shared across cycles.

Open in Lab
Watch the structure refine across 3 recycling iterations. Each pass sharpens the prediction.
The demo wakes as you arrive…

Training: the FAPE loss

How do you teach a network to predict 3D structures? You need a loss function that compares predicted and true atom positions — but there's a subtlety. If the predicted structure is simply rotated or translated relative to the ground truth, a naive loss would penalize this even though the structure is correct.

AlphaFold introduces Frame Aligned Point Error (FAPE): for each residue's local frame, transform all atoms into that frame and compute the error. By averaging over all frames as reference points, FAPE is invariant to global rotations and translations — it measures structural accuracy regardless of orientation. This is more informative than traditional RMSD because it captures local structural quality (each residue judges the structure from its own viewpoint), not just global alignment.

LFAPE=1N2∑i∑j∥Ti−1∘x^j−Titrue−1∘xjtrue∥\mathcal{L}_{\text{FAPE}} = \frac{1}{N^2} \sum_i \sum_j \| T_i^{-1} \circ \hat{x}_j - T_i^{\text{true}^{-1}} \circ x_j^{\text{true}} \|
Frame Aligned Point Error — the main training loss — For each residue i's frame, transform every atom j's position into that frame, and compare predicted vs. true. Averaging over all reference frames i makes the loss invariant to global rigid transformations.

Knowing what you know: confidence estimation

A prediction is only useful if you know how much to trust it. AlphaFold predicts a per-residue confidence score called (predicted Local Distance Difference Test), ranging from 0 to 100. Regions with pLDDT > 90 are typically accurate to within 1 Å of the experimental structure; regions below 50 are likely disordered (no fixed structure in nature).

AlphaFold also provides a Predicted Aligned Error (PAE) matrix — a residue × residue map showing the expected position error of residue jj when the prediction is aligned on residue ii. This reveals domain boundaries: within a rigid domain, PAE is low everywhere; between independently-moving domains, PAE is high.

The key ideas in code

Simplified Evoformer block — MSA attention with pair bias + outer product meanpython

Simplified to show the idea — not the real implementation.

import numpy as np

def softmax(x, axis=-1):
    e = np.exp(x - x.max(axis=axis, keepdims=True))
    return e / e.sum(axis=axis, keepdims=True)

def row_attention_with_pair_bias(msa, pair_rep, W_q, W_k, W_v):
    """
    Row-wise attention on MSA, biased by pair representation.
    msa:      (n_seq, n_res, d_msa)   — MSA features
    pair_rep: (n_res, n_res, d_pair)   — pair features
    """
    n_seq, n_res, d = msa.shape
    for s in range(n_seq):                   # each sequence independently
        Q = msa[s] @ W_q                     # (n_res, d_head)
        K = msa[s] @ W_k
        V = msa[s] @ W_v
        scores = Q @ K.T / np.sqrt(d)        # (n_res, n_res)
        # Inject pair representation as attention bias
        pair_bias = pair_rep.mean(axis=-1)    # (n_res, n_res)
        scores = scores + pair_bias           # spatial info guides attention
        weights = softmax(scores)
        msa[s] = weights @ V
    return msa

def outer_product_mean(msa, W_a, W_b):
    """
    Convert MSA covariance into pair representation updates.
    For each pair (i, j): average the outer product of their MSA columns.
    """
    n_seq, n_res, _ = msa.shape
    a = msa @ W_a    # (n_seq, n_res, c)
    b = msa @ W_b    # (n_seq, n_res, c)
    # outer product for each pair, averaged over sequences
    pair_update = np.einsum('ski,skj->ij', a, b) / n_seq
    return pair_update   # (n_res, n_res) — co-evolution → spatial info

# This is the core loop: MSA reads spatial clues from pairs,
# then MSA covariance updates pairs. 48 blocks of this back-and-forth
# is what makes AlphaFold reason about protein structure.

Why it worked — the design principles

The ripple effect

  1. 2018

    AlphaFold 1 — CASP13

    First version uses distance prediction + gradient descent to fold proteins. Wins CASP13 but accuracy still far from experimental.

  2. 2020

    AlphaFold 2 — CASP14

    A completely redesigned architecture with the Evoformer and structure module achieves experimental-level accuracy. Median GDT = 92.4, solving the protein folding problem.

  3. 2021

    Open-source release + 350K structures

    DeepMind publishes the code and releases predictions for the entire human proteome and 20 other organisms via the AlphaFold Protein Structure Database.

  4. 2022

    200 million structures

    The database expands to cover nearly every known protein in nature — over 200 million structures — revolutionizing structural biology overnight.

  5. 2024

    AlphaFold 3 + Nobel Prize

    AlphaFold 3 extends to predict interactions between proteins, DNA, RNA, and small molecules. Demis Hassabis and John Jumper awarded the Nobel Prize in Chemistry.

AlphaFold didn't just solve a computational problem — it gave biology an entirely new tool. Drug designers use it to find binding sites. Enzyme engineers use it to design proteins that don't exist in nature. Evolutionary biologists use it to study proteins too ancient for experimental determination. The architecture that powers modern language models also powers the prediction of life's molecular machinery.

CitationJumper, Evans, Pritzel, Green, Figurnov, Ronneberger, et al.. Highly Accurate Protein Structure Prediction with AlphaFold. Nature, 2021.

Terms in this paper