Learning Theory2019intermediate13 min read

Reconciling Modern Machine-Learning Practice and the Classical Bias–Variance Trade-Off

كيف نوفّق بين نجاح النماذج الكبيرة والفهم الكلاسيكي لمقايضة الانحياز والتباين

Belkin, M. · Hsu, D. · Ma, S. · Mandal, S. — PNAS

The problem

By 2018, a glaring contradiction sat at the heart of machine learning. The classical – trade-off — a cornerstone taught in every textbook — says that models should balance complexity: too simple means , too complex means . The "sweet spot" in the middle gives the best test performance. Yet practitioners were routinely enormous neural networks to perfectly fit every training point (zero training error) and getting excellent test accuracy. Classically, this should be a disaster. The gap between theory and practice had become untenable.

The contribution

The paper introduces the "double descent" risk curve — a unified framework that extends the classical U-shaped bias–variance curve beyond the threshold. The key insight: when is pushed far enough past the point where it perfectly fits (interpolates) training data, test risk decreases again, often falling below the classical "sweet spot." The authors demonstrate this phenomenon across neural networks, random Fourier features, random forests, and boosted trees on multiple datasets (MNIST, CIFAR-10, SVHN, TIMIT, 20-Newsgroups). They identify the mechanism: larger function classes contain smoother, lower- interpolating solutions — a form of Occam's razor where the simplest perfect fit improves as the search space grows.

The impact

This paper named and formalized a phenomenon that reshaped how the field thinks about model selection and . It inspired Nakkiran et al. (2020) to discover "deep double descent" (epoch-wise and dataset-size-wise variants), influenced the theoretical study of overparameterized models through the lens, and contributed to the understanding of grokking. The paper gave practitioners theoretical permission to use very large models without guilt — if the model is big enough to interpolate, making it bigger helps rather than hurts. This insight is now foundational to the scaling-laws era of modern AI.

Imagine hiring chefs for a cooking competition. A chef with too few skills undercooks every dish (underfitting). A chef with just enough skill to replicate every recipe in the cookbook — but no room to improvise — panics under pressure and produces bizarre dishes (overfitting at the interpolation threshold). But a master chef who knows thousands of recipes can reproduce the cookbook perfectly and improvise gracefully, because their vast repertoire lets them default to the simplest, most natural interpretation of each dish.

This paper discovered that machine learning models behave the same way. Push a model past the danger zone of "just barely enough capacity," and it finds smoother, simpler solutions — not despite its size, but because of it.

The classical story: the U-shaped curve

Every machine learning textbook tells the same story. You have training data, and you want to find a predictor hh from some function class H\mathcal{H} that performs well on new, unseen data. The standard approach is (ERM): pick the function that minimizes the training loss.

The classical wisdom says: if H\mathcal{H} is too small, every function in it underfits — none can capture the true pattern. If H\mathcal{H} is too large, the ERM solution overfits — it memorizes noise and spurious patterns. The textbook conclusion: find the sweet spot in between. This produces the famous U-shaped test risk curve, where performance first improves then degrades as model capacity grows.

This thinking guided decades of practice. , , cross-validated model selection — all are tools for finding that sweet spot. The unspoken rule was: a model with zero training error is overfit and will generalize poorly.

Open in Lab
The classical U-shaped curve: drag the capacity slider to see underfitting on the left, overfitting on the right, and the sweet spot in between.
The demo wakes as you arrive…

The modern contradiction: interpolation works

Then came deep learning. Practitioners discovered that the recipe for success was not finding the sweet spot — it was building networks so large that they could easily memorize every training example with zero loss. This is called interpolation: the model passes through every training point exactly, like a curve drawn through every dot on a scatter plot.

According to classical theory, this should be catastrophic overfitting. The model has memorized the training data including all its noise — surely it will fail on new data. But empirically, these interpolating models achieved state-of-the-art test accuracy. Zhang et al. (2017) demonstrated this dramatically: neural networks that could memorize completely random labels still generalized well when trained on real data.

A best practice emerged in deep learning: the network should be large enough to achieve zero training loss effortlessly. This directly contradicted the textbook rule. Something fundamental was missing from the classical picture.

The resolution: the double descent curve

Belkin et al. proposed a unified framework that resolves the contradiction. The classical U-shaped curve is not wrong — it is incomplete. It only shows what happens in the under-parameterized regime, where the model has fewer parameters than training examples. If you keep increasing model capacity past the interpolation threshold, a second phenomenon emerges.

The full picture is the double descent curve. It has three distinct regions. On the left, the classical regime: underfitting gives way to a sweet spot, then overfitting increases as capacity approaches nn (the number of training examples). At the interpolation threshold (N≈nN \approx n, where NN is the number of parameters), the model just barely has enough capacity to fit all training points — and test risk peaks sharply. This is the worst possible model: all its capacity is consumed by the constraint of fitting the data exactly, with no room for any useful .

But past this threshold, in the over-parameterized regime (N≫nN \gg n), something remarkable happens: test risk decreases again. The model has so many parameters that it can fit the training data in many different ways — and among all interpolating solutions, the learning algorithm finds progressively smoother, simpler ones. More capacity means more choices, and more choices means better solutions.

Open in Lab
The double descent risk curve: drag the capacity slider through the classical regime, the interpolation peak, and into the modern regime. Watch test risk rise, peak, then fall again.
The demo wakes as you arrive…

The mechanism: why bigger is smoother

All interpolating models (those to the right of the threshold) have zero training error. So what distinguishes a good interpolator from a bad one? The answer lies in the inductive bias — the implicit preference of the learning algorithm for certain solutions over others.

For Random Fourier Features (RFF), the learning algorithm finds the minimum-norm solution: among all coefficient vectors (a1,…,aN)(a_1, \ldots, a_N) that achieve zero training loss, it picks the one with the smallest ℓ2\ell_2 norm. This is the mathematical formulation of Occam's razor — choose the simplest explanation compatible with the data.

The key mathematical fact: as the number of features NN grows beyond nn, the minimum-norm interpolating function approximates the minimum-norm function in the full Reproducing Hilbert Space (RKHS). This RKHS-norm solution is the smoothest possible interpolator. More features means a closer approximation to this ideal smooth function.

h^n,N=arg⁡min⁡h∈HN∥h∥s.t.h(xi)=yi    ∀ i∈{1,…,n}\hat{h}_{n,N} = \arg\min_{h \in \mathcal{H}_N} \|h\| \quad \text{s.t.} \quad h(x_i) = y_i \;\; \forall\, i \in \{1,\ldots,n\}
Minimum-norm interpolation — Occam's razor for interpolating models — Among all functions in HN\mathcal{H}_N that perfectly fit the training data, choose the one with the smallest norm. As NN increases, this solution becomes smoother and its norm decreases. In the limit N→∞N \to \infty, the solution converges to the RKHS minimum-norm interpolator h^n,∞\hat{h}_{n,\infty}, which has the best generalization performance.
Open in Lab
Watch the norm shrink: as we add more features past the interpolation threshold, the minimum-norm solution becomes smoother. The fitted function visually flattens out.
The demo wakes as you arrive…

The same mechanism works across different model families, though the specific inductive bias differs. For neural networks trained with SGD, there is empirical and theoretical evidence that implicitly favors low-norm solutions — this is known as implicit regularization. For random forests, the inductive bias comes from averaging: each individual interpolating tree may be rough, but averaging many such trees produces a smooth interpolating function.

Remarkably, for kernel machines, three seemingly different approaches — explicit norm minimization, SGD from zero initialization, and averaging Gaussian process trajectories — all converge to the same minimum-norm interpolating solution. This convergence suggests that the double descent phenomenon is not an artifact of a particular algorithm, but a fundamental property of how model capacity interacts with data.

Evidence: double descent is universal

A central contribution of this paper is showing that double descent is not peculiar to one model type — it appears across fundamentally different learning methods.

Random Fourier Features (RFF): Two-layer neural networks with fixed random first-layer weights. On MNIST (n=10,000n=10{,}000), the test risk curve shows a clean U-shape for N<nN < n, peaks sharply at N=nN = n, and descends again for N>nN > n. The minimum-norm kernel machine solution (h^n,∞\hat{h}_{n,\infty}) outperforms every finite-NN model. The same pattern repeats on CIFAR-10, SVHN, TIMIT, and 20-Newsgroups.

Fully connected neural networks: Single hidden-layer networks trained with SGD on MNIST. The interpolation threshold appears at N≈n⋅KN \approx n \cdot K (where KK is the number of classes). Despite the non-convexity of the optimization landscape, double descent is clearly visible, though noisier than with RFF.

Random forests and boosted trees: Using the number of leaves per tree and the number of trees as a proxy for capacity, double descent appears. Averaging many interpolating trees produces smoother solutions with lower test risk than any individual tree.

Open in Lab
Double descent across model families: toggle between RFF, Neural Networks, and Random Forests to see the same phenomenon emerge in each.
The demo wakes as you arrive…

Theory: the approximation guarantee

The authors provide a theoretical foundation for why minimum-norm interpolation works. Suppose the data is generated by a target function h∗h^* in the RKHS H∞\mathcal{H}_\infty, with no noise. The training points x1,…,xnx_1, \ldots, x_n are drawn uniformly from a compact domain Ω\Omega. Then any interpolating function h∈H∞h \in \mathcal{H}_\infty satisfies a uniform approximation bound:

sup⁡x∈Ω∣h(x)−h∗(x)∣<A e−B(nlog⁡n)1/d(∥h∗∥H∞+∥h∥H∞)\sup_{x \in \Omega} |h(x) - h^*(x)| < A \, e^{-B \left(\frac{n}{\log n}\right)^{1/d}} \bigl(\|h^*\|_{\mathcal{H}_\infty} + \|h\|_{\mathcal{H}_\infty}\bigr)
Approximation theorem — interpolation quality improves exponentially with data — The error between any interpolating function hh and the true function h∗h^* decays exponentially with the data density. The bound scales with the sum of their norms, so the minimum-norm interpolator h^n,∞\hat{h}_{n,\infty} achieves the tightest bound. This formalizes the intuition: among all perfect fits, the smoothest one generalizes best.

Why was double descent historically overlooked?

If double descent is so universal, why did the field miss it for decades? The authors identify several practical and cultural reasons.

First, classical statistics typically works with fixed, small feature sets — there is no way to increase capacity past the interpolation threshold. Second, in non-parametric statistics, regularization is almost always used, and regularization masks the interpolation peak by preventing exact fitting. Third, RFF models were originally proposed as a cheap alternative to kernel machines, so they were typically used with N≪nN \ll n — well below the interpolation threshold.

For neural networks specifically, the non-convexity of training makes the curve noisy and hard to observe. Early stopping — stopping training when validation error starts rising — has a strong regularizing effect that hides the peak. And the peak itself occurs in a narrow range of parameters that is easy to miss when sampling architectures coarsely.

Open in Lab
See how regularization hides the peak: toggle regularization on/off and watch the interpolation spike appear or vanish.
The demo wakes as you arrive…

Why this matters: practical implications

The double descent framework has three immediate practical consequences.

First, model selection changes. The classical rule "pick the sweet spot" only applies in the under-parameterized regime. In modern practice, if you can afford to train a model past the interpolation threshold, making it larger improves generalization. This gives theoretical backing to the empirical deep learning practice of using very large networks.

Second, optimization becomes easier. Over-parameterized models have favorable optimization landscapes: SGD converges to global minima more reliably when N≫nN \gg n. So larger models are both easier to train and better at generalization — a double benefit that classical theory never predicted.

Third, regularization should be reconsidered. While regularization helps in the classical regime, it can prevent a model from reaching the beneficial over-parameterized regime by keeping it from interpolating. The paper suggests that practitioners should be more careful about when and how much regularization to apply.

Minimal experiment: double descent with Random Fourier Featurespython

Simplified to show the idea — not the real implementation.

import numpy as np
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split

# Load a small dataset (digits, 10 classes)
X, y = load_digits(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, train_size=200, random_state=42)
n = X_train.shape[0]  # 200 training points

# One-hot encode targets
Y_train = np.eye(10)[y_train]  # shape (200, 10)

# Try different numbers of Random Fourier Features
for N in [50, 150, 200, 500, 2000]:
    # Generate random frequencies from N(0, sigma^-2 * I)
    sigma = 5.0
    W = np.random.randn(X_train.shape[1], N) / sigma
    # Compute features: cos and sin (2N real params)
    Z_train = np.hstack([np.cos(X_train @ W), np.sin(X_train @ W)])
    Z_test  = np.hstack([np.cos(X_test @ W),  np.sin(X_test @ W)])
    # Minimum-norm least squares solution
    coeffs, _, _, _ = np.linalg.lstsq(Z_train, Y_train, rcond=None)
    # Predict and evaluate
    preds = Z_test @ coeffs
    acc = np.mean(preds.argmax(axis=1) == y_test)
    norm = np.linalg.norm(coeffs)
    print(f"N={N:5d}  acc={acc:.3f}  norm={norm:.1f}  {'← interpolation threshold' if N == n else ''}")

Legacy: from double descent to scaling laws

  1. 1992

    Bias–variance trade-off formalized

    Geman, Bienenstock & Doursat formally articulated the bias–variance decomposition for neural networks, establishing the U-shaped curve as conventional wisdom.

  2. 2017

    "Rethinking Generalization" (Zhang et al.)

    Showed that neural networks can memorize random labels yet still generalize on real data, exposing the gap between classical theory and modern practice.

  3. 2019

    This paper (Belkin et al.) — Double Descent

    Unified the classical U-curve and modern interpolation into the double descent framework. Demonstrated the phenomenon across neural nets, RFF, random forests, and boosting on multiple datasets.

  4. 2020

    Deep Double Descent (Nakkiran et al.)

    Extended double descent to epoch-wise and dataset-size-wise axes. Showed that test error can exhibit double descent as a function of training time or amount of data, not just model size.

  5. 2021

    Grokking (Power et al.)

    Discovered that models can suddenly transition from memorization to generalization long after fitting the training data. Double descent provides context for understanding this delayed generalization.

  6. 2020

    Scaling laws for neural language models

    Kaplan et al. showed smooth power-law relationships between model size, data, compute, and loss. The double descent framework helped explain why simply making models bigger keeps working.

The double descent curve fundamentally changed how the field thinks about the relationship between model size and generalization. Before this paper, "bigger model" was associated with "more overfitting." After it, the community understood that there are two regimes — and the modern regime, where larger models find smoother interpolations, explains the empirical success of deep learning. This insight is now woven into the foundation of the scaling-laws era, where the guiding principle is: if resources allow, make the model bigger.

CitationBelkin, Hsu, Ma, Mandal. Reconciling Modern Machine-Learning Practice and the Classical Bias–Variance Trade-Off. Proceedings of the National Academy of Sciences (PNAS), 2019.

Terms in this paper