Learning Theory2018advanced11 min read
Neural Tangent Kernel: Convergence and Generalization in Neural Networks
نواة المماس العصبية: التقارب والتعميم في الشبكات العصبية
Jacot, A. · Gabriel, F. · Hongler, C. — NeurIPS
The problem
Neural networks work astonishingly well in practice — they consistently converge to good solutions even when massively over-parameterized — but in 2018 there was no rigorous theory explaining why. The landscape is non-convex with countless local minima, yet different random initializations reliably reach similar performance. Why don't these networks get stuck? And why do they generalize to new data instead of memorizing the training set?
The contribution
The NTK paper proved that a neural network trained by can be described by a — the — defined by the inner product of the network's gradients. In the infinite-width limit, this kernel converges to a deterministic, fixed object that stays constant during training. This transforms the intractable non-convex optimization into well-understood , providing the first rigorous framework to study (via positive-definiteness of the kernel) and of deep networks.
The impact
The NTK paper launched the modern theory of over-parameterized neural networks. It showed that infinitely wide networks trained by gradient descent are equivalent to kernel machines, opening the door to proving convergence guarantees. It directly influenced the study of double descent, the development of maximal update parameterization (µP) for hyperparameter transfer, and the broader understanding of when and why neural networks escape the "" regime to learn useful features.
Imagine you're trying to understand a massive factory with millions of moving parts. Studying each gear and piston seems hopeless. But the NTK paper discovered something remarkable: if you make the factory infinitely large, all those moving parts synchronize into a single, predictable rhythm — like a metronome.
The factory's output changes smoothly and predictably, governed by one fixed object (the kernel), and you can prove mathematically that the production line will always converge to the target. The "metronome" is the Neural Tangent Kernel — and it turns chaotic optimization into a song you can follow note by note.
The puzzle: why does training work at all?
A neural network with more parameters than training examples can memorize anything — there are infinitely many solutions that fit the data perfectly. Classical statistics says such a model should overfit catastrophically and fail on new data. Yet in practice, large networks generalize beautifully.
Worse, the is non-convex: full of saddle points and local minima. Gradient descent has no guarantee of finding a . Yet different random initializations consistently reach similarly low loss. Something deep is happening that the classical theory cannot explain.
The NTK paper asked: what if we could simplify the dynamics of training? The key insight was to study what happens when the network becomes infinitely wide.
What is the Neural Tangent Kernel?
To understand the NTK, start with a simple question: when we update the network's parameters on one training example, how does the prediction on every other example change?
Consider a network with parameters . During gradient descent with an infinitesimally small , the output evolves as:
The term is the Neural Tangent Kernel . It measures how "aligned" the gradients at two inputs are — if updating on also pushes toward the correct answer, these inputs are kernel-similar.
Think of it like ripples in a pond: dropping a stone at one point (training on one example) creates waves that reach other points (affect other predictions). The NTK describes exactly how these ripples propagate.
The infinite-width miracle: a frozen kernel
For a finite network, the NTK changes during training — as the parameters move, the gradients change, and so does the kernel. This makes analysis intractable.
The breakthrough: Jacot et al. proved that when every width , two remarkable things happen:
-
Deterministic at initialization. The NTK converges to a fixed kernel that depends only on the architecture (depth, , variance) — not on the particular random initialization. By the law of large numbers, each entry in the kernel matrix becomes a sum of so many independent terms that randomness washes out.
-
Constant during training. Parameters move so little relative to their total norm that the kernel barely changes throughout training. The training dynamics become linear — equivalent to kernel regression with this fixed kernel.
This is the infinite-width limit: the chaotic non-convex optimization becomes a well-understood linear system, solvable in closed form for the squared loss.
The Gaussian process connection
The NTK result builds on an earlier discovery: at initialization, an infinitely wide network's outputs follow a . Neal (1994) first showed this for single-layer networks, and Lee & Bahri et al. (2018) extended it to deep networks.
The idea is elegant: each neuron's pre-activation is a weighted sum of many independent random terms. By the , this sum converges to a Gaussian. The covariance between outputs at two inputs is defined recursively, layer by layer, giving the Neural Network Gaussian Process (NNGP) kernel.
The NTK goes further. It shows that not just the initial outputs but the training dynamics are governed by a kernel. The NTK at layer builds recursively on the NTK at layer :
where is the NNGP covariance and involves the derivative of the activation function. Each layer adds its own contribution to the kernel, building a hierarchy of similarity measures.
Lazy training: when parameters barely move
The reason the NTK stays constant is intimately connected to a phenomenon called lazy training (Chizat et al., 2019). In the infinite-width regime, the network has so many parameters that each individual parameter needs to change only infinitesimally to collectively fit the training data.
Formally, the linearized model — a first-order of the network around its initialization — becomes an exact description of training:
Since stays small, the barely changes, so the NTK stays approximately constant. The network effectively behaves like a linear model in a very high-dimensional feature space defined by the gradient features .
This is both the power and the limitation of the NTK regime. Power: it makes everything analytically tractable. Limitation: a frozen kernel means no — the network never discovers new representations beyond what the random initialization already encodes. Real networks of practical width do learn features, which is why the NTK theory describes the starting point, not the full story.
The same idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
def relu(x):
return np.maximum(0, x)
def relu_deriv(x):
return (x > 0).astype(float)
def empirical_ntk(X, W1, b1, W2):
"""Compute the empirical NTK for a 2-layer ReLU network.
X: (n_samples, d_in), W1: (d_in, width), b1: (width,), W2: (width, 1)
Returns: (n_samples, n_samples) kernel matrix."""
h = relu(X @ W1 + b1) # hidden activations: (n, width)
# Jacobian w.r.t. W2: dF/dW2 = h
J_W2 = h # (n, width)
# Jacobian w.r.t. W1: chain rule through ReLU
mask = relu_deriv(X @ W1 + b1) # (n, width)
# For each sample i, dF/dW1[i] is outer(x_i, mask_i * W2.T)
# Flattened: (n, d_in * width)
J_W1 = (X[:, :, None] * (mask * W2.T)[None, :, :]).reshape(X.shape[0], -1)
# Full Jacobian and NTK
J = np.concatenate([J_W1, J_W2], axis=1) # (n, total_params)
K = J @ J.T # NTK: (n, n)
return K
# Initialize a wide network
np.random.seed(42)
width = 2048
X = np.random.randn(5, 3) # 5 samples, 3 features
W1 = np.random.randn(3, width) / np.sqrt(3) # NTK parameterization
b1 = np.zeros(width)
W2 = np.random.randn(width, 1) / np.sqrt(width)
K = empirical_ntk(X, W1, b1, W2)
print("NTK matrix (5×5):\n", np.round(K, 2))
print("All eigenvalues positive?", np.all(np.linalg.eigvalsh(K) > 0))Convergence and generalization
The NTK theory provides two key guarantees:
Convergence. When the limiting NTK is positive-definite on the training set, gradient descent drives the training loss to zero exponentially fast. Each component of the residual — decomposed along the eigenvectors of — decays at a rate proportional to the corresponding . Components aligned with the top eigenvectors (the "easy" directions in function space) are learned first; components aligned with the bottom eigenvectors (the "hard" directions) take longest.
Generalization. The NTK framework shows that infinitely wide networks generalize like kernel machines: predictions on new inputs are weighted combinations of training labels, with weights determined by the kernel. The smoothness properties of the kernel — encoded in the decay rate of its eigenvalues — determine the effective complexity of the learned function, providing a natural form of implicit regularization even without explicit or .
How architecture shapes the NTK
A key insight from the NTK framework is that the architecture determines the kernel, and the kernel determines what the network can learn. Different design choices create different NTK kernels:
-
Depth. Deeper networks produce NTKs with richer structure, but two extreme regimes emerge. In the ordered (freeze) regime, the NTK becomes nearly constant across all inputs — making different inputs look identical, which hurts generalization. In the chaotic regime, the NTK approaches a Kronecker delta — making everything look different, which speeds training but may hurt generalization too.
-
Activation function. The choice of nonlinearity (ReLU, tanh, etc.) directly determines the "second moment" in the recursive formula, changing how similarity between inputs transforms across layers.
-
Initialization variance. The ratio controls whether the network is in the ordered or chaotic regime, with the — the boundary between the two — often giving the best performance.
Why it changed everything
2018
NTK paper published
Jacot, Gabriel, and Hongler introduce the Neural Tangent Kernel at NeurIPS. Prove convergence to a deterministic, constant kernel in the infinite-width limit.
2019
Lazy training formalized
Chizat et al. coin "lazy training" to describe the regime where the NTK stays approximately constant — and show it as a consequence of over-parameterization.
2019
Wide networks evolve as linear models
Lee, Xiao et al. show that wide networks of any depth evolve exactly as linearized models under gradient descent, completing the theoretical picture.
2020
Double descent discovered
The NTK framework helped explain why test error can decrease again past the interpolation threshold — the "double descent" curve that contradicts classical bias-variance tradeoff intuitions.
2021
µP for hyperparameter transfer
Yang and Hu develop Maximal Update Parameterization (µP), directly building on NTK theory to enable hyperparameter transfer across model widths.
2022
Finite-width corrections
Researchers develop perturbative corrections to the NTK theory, describing how finite-width effects produce feature learning — bridging the gap between theory and practice.
The NTK didn't just give us a theorem — it gave us a language for discussing neural network training dynamics. Concepts like "lazy vs. feature-learning regime", "kernel alignment", and "" all trace back to this foundational work. Every theoretical advance in understanding why large models work — from double descent to scaling laws — stands on the bridge that the NTK built between kernel methods and .
CitationJacot, Gabriel, Hongler. Neural Tangent Kernel: Convergence and Generalization in Neural Networks. NeurIPS, 2018.
Terms in this paper
- Kernelالنواة الحسابية
- Gaussian Processالعملية الغاوسية
- Gradient Descentالانحدار التدريجي
- Convergenceالتقارب الحسابي
- Generalizationالتعميم
- Jacobianمصفوفة جاكوبي
- Taylor Expansionتفكيك تايلور
- Dot Productالضرب النقطي
- Feature Extractionاستخلاص السمات
- Overfittingفرط التخصيص
- Loss functionدالة الخسارة
- Backpropagationالتحديث التراجعي
- Parameterالمعلمة البنيوية
- Weightالوزن البنيوي
- Biasالانحياز الحسابي