Learning Theory2018advanced11 min read

Neural Tangent Kernel: Convergence and Generalization in Neural Networks

نواة المماس العصبية: التقارب والتعميم في الشبكات العصبية

Jacot, A. · Gabriel, F. · Hongler, C. — NeurIPS

The problem

Neural networks work astonishingly well in practice — they consistently converge to good solutions even when massively over-parameterized — but in 2018 there was no rigorous theory explaining why. The landscape is non-convex with countless local minima, yet different random initializations reliably reach similar performance. Why don't these networks get stuck? And why do they generalize to new data instead of memorizing the training set?

The contribution

The NTK paper proved that a neural network trained by can be described by a — the — defined by the inner product of the network's gradients. In the infinite-width limit, this kernel converges to a deterministic, fixed object that stays constant during training. This transforms the intractable non-convex optimization into well-understood , providing the first rigorous framework to study (via positive-definiteness of the kernel) and of deep networks.

The impact

The NTK paper launched the modern theory of over-parameterized neural networks. It showed that infinitely wide networks trained by gradient descent are equivalent to kernel machines, opening the door to proving convergence guarantees. It directly influenced the study of double descent, the development of maximal update parameterization (µP) for hyperparameter transfer, and the broader understanding of when and why neural networks escape the "" regime to learn useful features.

Imagine you're trying to understand a massive factory with millions of moving parts. Studying each gear and piston seems hopeless. But the NTK paper discovered something remarkable: if you make the factory infinitely large, all those moving parts synchronize into a single, predictable rhythm — like a metronome.

The factory's output changes smoothly and predictably, governed by one fixed object (the kernel), and you can prove mathematically that the production line will always converge to the target. The "metronome" is the Neural Tangent Kernel — and it turns chaotic optimization into a song you can follow note by note.

The puzzle: why does training work at all?

A neural network with more parameters than training examples can memorize anything — there are infinitely many solutions that fit the data perfectly. Classical statistics says such a model should overfit catastrophically and fail on new data. Yet in practice, large networks generalize beautifully.

Worse, the is non-convex: full of saddle points and local minima. Gradient descent has no guarantee of finding a . Yet different random initializations consistently reach similarly low loss. Something deep is happening that the classical theory cannot explain.

The NTK paper asked: what if we could simplify the dynamics of training? The key insight was to study what happens when the network becomes infinitely wide.

Open in Lab
Drag the width slider to see how increasing width smooths the loss landscape. At infinite width, all initializations converge to the same global minimum.
The demo wakes as you arrive…

What is the Neural Tangent Kernel?

To understand the NTK, start with a simple question: when we update the network's parameters on one training example, how does the prediction on every other example change?

Consider a network f(x;θ)f(\mathbf{x}; \theta) with parameters θ\theta. During gradient descent with an infinitesimally small , the output evolves as:

df(x)dt=−η∑i∇θf(x)⊤∇θf(xi)⋅∇fℓi\frac{df(\mathbf{x})}{dt} = -\eta \sum_{i} \nabla_\theta f(\mathbf{x})^\top \nabla_\theta f(\mathbf{x}_i) \cdot \nabla_f \ell_i

The term ∇θf(x)⊤∇θf(xi)\nabla_\theta f(\mathbf{x})^\top \nabla_\theta f(\mathbf{x}_i) is the Neural Tangent Kernel K(x,xi)K(\mathbf{x}, \mathbf{x}_i). It measures how "aligned" the gradients at two inputs are — if updating on xi\mathbf{x}_i also pushes f(x)f(\mathbf{x}) toward the correct answer, these inputs are kernel-similar.

Think of it like ripples in a pond: dropping a stone at one point (training on one example) creates waves that reach other points (affect other predictions). The NTK describes exactly how these ripples propagate.

K(x,x′)=∇θf(x;θ)⊤∇θf(x′;θ)=∑p=1P∂f(x)∂θp∂f(x′)∂θpK(\mathbf{x}, \mathbf{x}') = \nabla_\theta f(\mathbf{x};\theta)^\top \nabla_\theta f(\mathbf{x}';\theta) = \sum_{p=1}^{P} \frac{\partial f(\mathbf{x})}{\partial \theta_p} \frac{\partial f(\mathbf{x}')}{\partial \theta_p}
The Neural Tangent Kernel — the inner product of parameter gradients — Each parameter contributes one term to this sum. The kernel value tells you how correlated the learning signal is between two inputs. High K means training on x also helps x′.
Open in Lab
Click a training point to "train" on it. Watch how the prediction surface changes everywhere — the NTK governs the shape of this influence.
The demo wakes as you arrive…

The infinite-width miracle: a frozen kernel

For a finite network, the NTK changes during training — as the parameters move, the gradients change, and so does the kernel. This makes analysis intractable.

The breakthrough: Jacot et al. proved that when every width n1,n2,…,nL→∞n_1, n_2, \ldots, n_L \to \infty, two remarkable things happen:

  • Deterministic at initialization. The NTK converges to a fixed kernel that depends only on the architecture (depth, , variance) — not on the particular random initialization. By the law of large numbers, each entry in the kernel matrix becomes a sum of so many independent terms that randomness washes out.

  • Constant during training. Parameters move so little relative to their total norm that the kernel barely changes throughout training. The training dynamics become linear — equivalent to kernel regression with this fixed kernel.

This is the infinite-width limit: the chaotic non-convex optimization becomes a well-understood linear system, solvable in closed form for the squared loss.

Open in Lab
Watch the residual decompose along the NTK's eigenvectors. Large eigenvalues learn fast; small eigenvalues learn slowly. Convergence is guaranteed when all eigenvalues are positive.
The demo wakes as you arrive…

The Gaussian process connection

The NTK result builds on an earlier discovery: at initialization, an infinitely wide network's outputs follow a . Neal (1994) first showed this for single-layer networks, and Lee & Bahri et al. (2018) extended it to deep networks.

The idea is elegant: each neuron's pre-activation is a weighted sum of many independent random terms. By the , this sum converges to a Gaussian. The covariance between outputs at two inputs is defined recursively, layer by layer, giving the Neural Network Gaussian Process (NNGP) kernel.

The NTK goes further. It shows that not just the initial outputs but the training dynamics are governed by a kernel. The NTK at layer L+1L+1 builds recursively on the NTK at layer LL:

K(L+1)=Σ(L+1)+K(L)⋅Σ˙(L+1)K^{(L+1)} = \Sigma^{(L+1)} + K^{(L)} \cdot \dot{\Sigma}^{(L+1)}

where Σ(L+1)\Sigma^{(L+1)} is the NNGP covariance and Σ˙(L+1)\dot{\Sigma}^{(L+1)} involves the derivative of the activation function. Each layer adds its own contribution to the kernel, building a hierarchy of similarity measures.

Σ(l+1)(x,x′)=Ef∼N(0,Λ(l))[σ(f(x))⋅σ(f(x′))]+β2\Sigma^{(l+1)}(\mathbf{x}, \mathbf{x}') = \mathbb{E}_{f \sim \mathcal{N}(0, \Lambda^{(l)})}[\sigma(f(\mathbf{x})) \cdot \sigma(f(\mathbf{x}'))] + \beta^2
Recursive NNGP covariance — each layer transforms the similarity — Starting from the input inner product Σ⁽¹⁾ = x·x′/n₀ + β², each subsequent layer passes the covariance through the activation function's second moment. The nonlinearity creates a new, richer notion of similarity at every depth.
Open in Lab
Step through layers to see how the NTK kernel matrix builds up recursively. Each layer transforms and enriches the notion of similarity between inputs.
The demo wakes as you arrive…

Lazy training: when parameters barely move

The reason the NTK stays constant is intimately connected to a phenomenon called lazy training (Chizat et al., 2019). In the infinite-width regime, the network has so many parameters that each individual parameter needs to change only infinitesimally to collectively fit the training data.

Formally, the linearized model — a first-order of the network around its initialization — becomes an exact description of training:

flin(θ(t))=f(θ0)+∇θf(θ0)⋅(θ(t)−θ0)f^{\text{lin}}(\theta(t)) = f(\theta_0) + \nabla_\theta f(\theta_0) \cdot (\theta(t) - \theta_0)

Since θ(t)−θ0\theta(t) - \theta_0 stays small, the ∇θf\nabla_\theta f barely changes, so the NTK K=J⊤JK = J^\top J stays approximately constant. The network effectively behaves like a linear model in a very high-dimensional feature space defined by the gradient features ∇θf(x;θ0)\nabla_\theta f(\mathbf{x}; \theta_0).

This is both the power and the limitation of the NTK regime. Power: it makes everything analytically tractable. Limitation: a frozen kernel means no — the network never discovers new representations beyond what the random initialization already encodes. Real networks of practical width do learn features, which is why the NTK theory describes the starting point, not the full story.

Open in Lab
Compare: in the lazy regime (left), gradient features stay fixed; in the feature-learning regime (right), the network discovers new representations during training.
The demo wakes as you arrive…

The same idea in code

Computing the empirical NTK for a two-layer networkpython

Simplified to show the idea — not the real implementation.

import numpy as np

def relu(x):
    return np.maximum(0, x)

def relu_deriv(x):
    return (x > 0).astype(float)

def empirical_ntk(X, W1, b1, W2):
    """Compute the empirical NTK for a 2-layer ReLU network.
    X: (n_samples, d_in), W1: (d_in, width), b1: (width,), W2: (width, 1)
    Returns: (n_samples, n_samples) kernel matrix."""
    h = relu(X @ W1 + b1)                    # hidden activations: (n, width)
    # Jacobian w.r.t. W2: dF/dW2 = h
    J_W2 = h                                  # (n, width)
    # Jacobian w.r.t. W1: chain rule through ReLU
    mask = relu_deriv(X @ W1 + b1)            # (n, width)
    # For each sample i, dF/dW1[i] is outer(x_i, mask_i * W2.T)
    # Flattened: (n, d_in * width)
    J_W1 = (X[:, :, None] * (mask * W2.T)[None, :, :]).reshape(X.shape[0], -1)
    # Full Jacobian and NTK
    J = np.concatenate([J_W1, J_W2], axis=1)  # (n, total_params)
    K = J @ J.T                               # NTK: (n, n)
    return K

# Initialize a wide network
np.random.seed(42)
width = 2048
X = np.random.randn(5, 3)                     # 5 samples, 3 features
W1 = np.random.randn(3, width) / np.sqrt(3)   # NTK parameterization
b1 = np.zeros(width)
W2 = np.random.randn(width, 1) / np.sqrt(width)

K = empirical_ntk(X, W1, b1, W2)
print("NTK matrix (5×5):\n", np.round(K, 2))
print("All eigenvalues positive?", np.all(np.linalg.eigvalsh(K) > 0))

Convergence and generalization

The NTK theory provides two key guarantees:

Convergence. When the limiting NTK K∞K_\infty is positive-definite on the training set, gradient descent drives the training loss to zero exponentially fast. Each component of the residual — decomposed along the eigenvectors of K∞K_\infty — decays at a rate proportional to the corresponding . Components aligned with the top eigenvectors (the "easy" directions in function space) are learned first; components aligned with the bottom eigenvectors (the "hard" directions) take longest.

Generalization. The NTK framework shows that infinitely wide networks generalize like kernel machines: predictions on new inputs are weighted combinations of training labels, with weights determined by the kernel. The smoothness properties of the kernel — encoded in the decay rate of its eigenvalues — determine the effective complexity of the learned function, providing a natural form of implicit regularization even without explicit or .

f(θ(t))=Y+(f(θ0)−Y) e−ηK∞tf(\theta(t)) = \mathcal{Y} + (f(\theta_0) - \mathcal{Y}) \, e^{-\eta K_\infty t}
Closed-form solution for MSE loss in the NTK regime — Training predictions converge exponentially toward targets. The rate is controlled by the NTK eigenvalues — large eigenvalues learn fast, small eigenvalues learn slowly but all eventually converge.

How architecture shapes the NTK

A key insight from the NTK framework is that the architecture determines the kernel, and the kernel determines what the network can learn. Different design choices create different NTK kernels:

  • Depth. Deeper networks produce NTKs with richer structure, but two extreme regimes emerge. In the ordered (freeze) regime, the NTK becomes nearly constant across all inputs — making different inputs look identical, which hurts generalization. In the chaotic regime, the NTK approaches a Kronecker delta — making everything look different, which speeds training but may hurt generalization too.

  • Activation function. The choice of nonlinearity (ReLU, tanh, etc.) directly determines the "second moment" in the recursive formula, changing how similarity between inputs transforms across layers.

  • Initialization variance. The ratio σw2/nl\sigma_w^2 / n_l controls whether the network is in the ordered or chaotic regime, with the — the boundary between the two — often giving the best performance.

Open in Lab
Adjust depth and initialization variance to explore the ordered → edge of chaos → chaotic transition. The NTK heatmap changes dramatically between regimes.
The demo wakes as you arrive…

Why it changed everything

  1. 2018

    NTK paper published

    Jacot, Gabriel, and Hongler introduce the Neural Tangent Kernel at NeurIPS. Prove convergence to a deterministic, constant kernel in the infinite-width limit.

  2. 2019

    Lazy training formalized

    Chizat et al. coin "lazy training" to describe the regime where the NTK stays approximately constant — and show it as a consequence of over-parameterization.

  3. 2019

    Wide networks evolve as linear models

    Lee, Xiao et al. show that wide networks of any depth evolve exactly as linearized models under gradient descent, completing the theoretical picture.

  4. 2020

    Double descent discovered

    The NTK framework helped explain why test error can decrease again past the interpolation threshold — the "double descent" curve that contradicts classical bias-variance tradeoff intuitions.

  5. 2021

    µP for hyperparameter transfer

    Yang and Hu develop Maximal Update Parameterization (µP), directly building on NTK theory to enable hyperparameter transfer across model widths.

  6. 2022

    Finite-width corrections

    Researchers develop perturbative corrections to the NTK theory, describing how finite-width effects produce feature learning — bridging the gap between theory and practice.

The NTK didn't just give us a theorem — it gave us a language for discussing neural network training dynamics. Concepts like "lazy vs. feature-learning regime", "kernel alignment", and "" all trace back to this foundational work. Every theoretical advance in understanding why large models work — from double descent to scaling laws — stands on the bridge that the NTK built between kernel methods and .

CitationJacot, Gabriel, Hongler. Neural Tangent Kernel: Convergence and Generalization in Neural Networks. NeurIPS, 2018.

Terms in this paper