Learning Theory2019intermediate13 min read
Reconciling Modern Machine-Learning Practice and the Classical Bias–Variance Trade-Off
كيف نوفّق بين نجاح النماذج الكبيرة والفهم الكلاسيكي لمقايضة الانحياز والتباين
Belkin, M. · Hsu, D. · Ma, S. · Mandal, S. — PNAS
The problem
By 2018, a glaring contradiction sat at the heart of machine learning. The classical – trade-off — a cornerstone taught in every textbook — says that models should balance complexity: too simple means , too complex means . The "sweet spot" in the middle gives the best test performance. Yet practitioners were routinely enormous neural networks to perfectly fit every training point (zero training error) and getting excellent test accuracy. Classically, this should be a disaster. The gap between theory and practice had become untenable.
The contribution
The paper introduces the "double descent" risk curve — a unified framework that extends the classical U-shaped bias–variance curve beyond the threshold. The key insight: when is pushed far enough past the point where it perfectly fits (interpolates) training data, test risk decreases again, often falling below the classical "sweet spot." The authors demonstrate this phenomenon across neural networks, random Fourier features, random forests, and boosted trees on multiple datasets (MNIST, CIFAR-10, SVHN, TIMIT, 20-Newsgroups). They identify the mechanism: larger function classes contain smoother, lower- interpolating solutions — a form of Occam's razor where the simplest perfect fit improves as the search space grows.
The impact
This paper named and formalized a phenomenon that reshaped how the field thinks about model selection and . It inspired Nakkiran et al. (2020) to discover "deep double descent" (epoch-wise and dataset-size-wise variants), influenced the theoretical study of overparameterized models through the lens, and contributed to the understanding of grokking. The paper gave practitioners theoretical permission to use very large models without guilt — if the model is big enough to interpolate, making it bigger helps rather than hurts. This insight is now foundational to the scaling-laws era of modern AI.
Imagine hiring chefs for a cooking competition. A chef with too few skills undercooks every dish (underfitting). A chef with just enough skill to replicate every recipe in the cookbook — but no room to improvise — panics under pressure and produces bizarre dishes (overfitting at the interpolation threshold). But a master chef who knows thousands of recipes can reproduce the cookbook perfectly and improvise gracefully, because their vast repertoire lets them default to the simplest, most natural interpretation of each dish.
This paper discovered that machine learning models behave the same way. Push a model past the danger zone of "just barely enough capacity," and it finds smoother, simpler solutions — not despite its size, but because of it.
The classical story: the U-shaped curve
Every machine learning textbook tells the same story. You have training data, and you want to find a predictor from some function class that performs well on new, unseen data. The standard approach is (ERM): pick the function that minimizes the training loss.
The classical wisdom says: if is too small, every function in it underfits — none can capture the true pattern. If is too large, the ERM solution overfits — it memorizes noise and spurious patterns. The textbook conclusion: find the sweet spot in between. This produces the famous U-shaped test risk curve, where performance first improves then degrades as model capacity grows.
This thinking guided decades of practice. , , cross-validated model selection — all are tools for finding that sweet spot. The unspoken rule was: a model with zero training error is overfit and will generalize poorly.
The modern contradiction: interpolation works
Then came deep learning. Practitioners discovered that the recipe for success was not finding the sweet spot — it was building networks so large that they could easily memorize every training example with zero loss. This is called interpolation: the model passes through every training point exactly, like a curve drawn through every dot on a scatter plot.
According to classical theory, this should be catastrophic overfitting. The model has memorized the training data including all its noise — surely it will fail on new data. But empirically, these interpolating models achieved state-of-the-art test accuracy. Zhang et al. (2017) demonstrated this dramatically: neural networks that could memorize completely random labels still generalized well when trained on real data.
A best practice emerged in deep learning: the network should be large enough to achieve zero training loss effortlessly. This directly contradicted the textbook rule. Something fundamental was missing from the classical picture.
The resolution: the double descent curve
Belkin et al. proposed a unified framework that resolves the contradiction. The classical U-shaped curve is not wrong — it is incomplete. It only shows what happens in the under-parameterized regime, where the model has fewer parameters than training examples. If you keep increasing model capacity past the interpolation threshold, a second phenomenon emerges.
The full picture is the double descent curve. It has three distinct regions. On the left, the classical regime: underfitting gives way to a sweet spot, then overfitting increases as capacity approaches (the number of training examples). At the interpolation threshold (, where is the number of parameters), the model just barely has enough capacity to fit all training points — and test risk peaks sharply. This is the worst possible model: all its capacity is consumed by the constraint of fitting the data exactly, with no room for any useful .
But past this threshold, in the over-parameterized regime (), something remarkable happens: test risk decreases again. The model has so many parameters that it can fit the training data in many different ways — and among all interpolating solutions, the learning algorithm finds progressively smoother, simpler ones. More capacity means more choices, and more choices means better solutions.
The mechanism: why bigger is smoother
All interpolating models (those to the right of the threshold) have zero training error. So what distinguishes a good interpolator from a bad one? The answer lies in the inductive bias — the implicit preference of the learning algorithm for certain solutions over others.
For Random Fourier Features (RFF), the learning algorithm finds the minimum-norm solution: among all coefficient vectors that achieve zero training loss, it picks the one with the smallest norm. This is the mathematical formulation of Occam's razor — choose the simplest explanation compatible with the data.
The key mathematical fact: as the number of features grows beyond , the minimum-norm interpolating function approximates the minimum-norm function in the full Reproducing Hilbert Space (RKHS). This RKHS-norm solution is the smoothest possible interpolator. More features means a closer approximation to this ideal smooth function.
The same mechanism works across different model families, though the specific inductive bias differs. For neural networks trained with SGD, there is empirical and theoretical evidence that implicitly favors low-norm solutions — this is known as implicit regularization. For random forests, the inductive bias comes from averaging: each individual interpolating tree may be rough, but averaging many such trees produces a smooth interpolating function.
Remarkably, for kernel machines, three seemingly different approaches — explicit norm minimization, SGD from zero initialization, and averaging Gaussian process trajectories — all converge to the same minimum-norm interpolating solution. This convergence suggests that the double descent phenomenon is not an artifact of a particular algorithm, but a fundamental property of how model capacity interacts with data.
Evidence: double descent is universal
A central contribution of this paper is showing that double descent is not peculiar to one model type — it appears across fundamentally different learning methods.
Random Fourier Features (RFF): Two-layer neural networks with fixed random first-layer weights. On MNIST (), the test risk curve shows a clean U-shape for , peaks sharply at , and descends again for . The minimum-norm kernel machine solution () outperforms every finite- model. The same pattern repeats on CIFAR-10, SVHN, TIMIT, and 20-Newsgroups.
Fully connected neural networks: Single hidden-layer networks trained with SGD on MNIST. The interpolation threshold appears at (where is the number of classes). Despite the non-convexity of the optimization landscape, double descent is clearly visible, though noisier than with RFF.
Random forests and boosted trees: Using the number of leaves per tree and the number of trees as a proxy for capacity, double descent appears. Averaging many interpolating trees produces smoother solutions with lower test risk than any individual tree.
Theory: the approximation guarantee
The authors provide a theoretical foundation for why minimum-norm interpolation works. Suppose the data is generated by a target function in the RKHS , with no noise. The training points are drawn uniformly from a compact domain . Then any interpolating function satisfies a uniform approximation bound:
Why was double descent historically overlooked?
If double descent is so universal, why did the field miss it for decades? The authors identify several practical and cultural reasons.
First, classical statistics typically works with fixed, small feature sets — there is no way to increase capacity past the interpolation threshold. Second, in non-parametric statistics, regularization is almost always used, and regularization masks the interpolation peak by preventing exact fitting. Third, RFF models were originally proposed as a cheap alternative to kernel machines, so they were typically used with — well below the interpolation threshold.
For neural networks specifically, the non-convexity of training makes the curve noisy and hard to observe. Early stopping — stopping training when validation error starts rising — has a strong regularizing effect that hides the peak. And the peak itself occurs in a narrow range of parameters that is easy to miss when sampling architectures coarsely.
Why this matters: practical implications
The double descent framework has three immediate practical consequences.
First, model selection changes. The classical rule "pick the sweet spot" only applies in the under-parameterized regime. In modern practice, if you can afford to train a model past the interpolation threshold, making it larger improves generalization. This gives theoretical backing to the empirical deep learning practice of using very large networks.
Second, optimization becomes easier. Over-parameterized models have favorable optimization landscapes: SGD converges to global minima more reliably when . So larger models are both easier to train and better at generalization — a double benefit that classical theory never predicted.
Third, regularization should be reconsidered. While regularization helps in the classical regime, it can prevent a model from reaching the beneficial over-parameterized regime by keeping it from interpolating. The paper suggests that practitioners should be more careful about when and how much regularization to apply.
Simplified to show the idea — not the real implementation.
import numpy as np
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
# Load a small dataset (digits, 10 classes)
X, y = load_digits(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, train_size=200, random_state=42)
n = X_train.shape[0] # 200 training points
# One-hot encode targets
Y_train = np.eye(10)[y_train] # shape (200, 10)
# Try different numbers of Random Fourier Features
for N in [50, 150, 200, 500, 2000]:
# Generate random frequencies from N(0, sigma^-2 * I)
sigma = 5.0
W = np.random.randn(X_train.shape[1], N) / sigma
# Compute features: cos and sin (2N real params)
Z_train = np.hstack([np.cos(X_train @ W), np.sin(X_train @ W)])
Z_test = np.hstack([np.cos(X_test @ W), np.sin(X_test @ W)])
# Minimum-norm least squares solution
coeffs, _, _, _ = np.linalg.lstsq(Z_train, Y_train, rcond=None)
# Predict and evaluate
preds = Z_test @ coeffs
acc = np.mean(preds.argmax(axis=1) == y_test)
norm = np.linalg.norm(coeffs)
print(f"N={N:5d} acc={acc:.3f} norm={norm:.1f} {'← interpolation threshold' if N == n else ''}")Legacy: from double descent to scaling laws
1992
Bias–variance trade-off formalized
Geman, Bienenstock & Doursat formally articulated the bias–variance decomposition for neural networks, establishing the U-shaped curve as conventional wisdom.
2017
"Rethinking Generalization" (Zhang et al.)
Showed that neural networks can memorize random labels yet still generalize on real data, exposing the gap between classical theory and modern practice.
2019
This paper (Belkin et al.) — Double Descent
Unified the classical U-curve and modern interpolation into the double descent framework. Demonstrated the phenomenon across neural nets, RFF, random forests, and boosting on multiple datasets.
2020
Deep Double Descent (Nakkiran et al.)
Extended double descent to epoch-wise and dataset-size-wise axes. Showed that test error can exhibit double descent as a function of training time or amount of data, not just model size.
2021
Grokking (Power et al.)
Discovered that models can suddenly transition from memorization to generalization long after fitting the training data. Double descent provides context for understanding this delayed generalization.
2020
Scaling laws for neural language models
Kaplan et al. showed smooth power-law relationships between model size, data, compute, and loss. The double descent framework helped explain why simply making models bigger keeps working.
The double descent curve fundamentally changed how the field thinks about the relationship between model size and generalization. Before this paper, "bigger model" was associated with "more overfitting." After it, the community understood that there are two regimes — and the modern regime, where larger models find smoother interpolations, explains the empirical success of deep learning. This insight is now woven into the foundation of the scaling-laws era, where the guiding principle is: if resources allow, make the model bigger.
CitationBelkin, Hsu, Ma, Mandal. Reconciling Modern Machine-Learning Practice and the Classical Bias–Variance Trade-Off. Proceedings of the National Academy of Sciences (PNAS), 2019.
Terms in this paper
- Bias-Variance Tradeoffالموازنة بين الانحياز والتباعد
- Overfittingفرط التخصيص
- Generalizationالتعميم
- Interpolationالاستيفاء
- Regularizationالضبط الهيكلي
- Neural Networkالشبكة العصبية
- Decision Treeشجرة القرار الإحصائية
- Ensembleالنماذج التجميعية الهجينة
- Kernelالنواة الحسابية
- Modelالنموذج
- Empirical Risk Minimizationتقليل المخاطر التجريبية
- Inductive Biasالانحياز الاستقرائي المسبق
- Neural Tangent Kernelنواة المماس العصبي
- Normالمعيار التفاضلي