Core ML1970foundational9 min read

Ridge Regression: Biased Estimation for Nonorthogonal Problems

انحدار الحَرْف: التقدير المنحاز للمسائل غير المتعامدة

Hoerl, A. E. · Kennard, R. W. — Technometrics

The problem

Ordinary (OLS) is the workhorse of , but when predictor variables are correlated (), OLS coefficients become wildly unstable — tiny changes in the data produce huge swings in estimates. Mathematically, the matrix X'X is nearly singular, its small eigenvalues inflate the of β̂ to useless levels, and the resulting may fit the training data perfectly yet predict new data terribly.

The contribution

Add a small positive constant k to the diagonal of X'X before inverting: β̂_ridge = (X'X + kI)⁻¹ X'y. This "ridge" along the diagonal stabilizes the inverse, shrinks every coefficient toward zero, and introduces a controlled amount of . The key theorem shows that there always exists a k > 0 for which the of β̂_ridge is strictly less than that of OLS — the gain in stability more than compensates for the introduced bias.

The impact

Ridge regression introduced the idea that a biased estimator can outperform an unbiased one, launching the entire field of . It directly inspired the Lasso (L1 penalty), Elastic Net, and every modern regularized model. The bias–variance tradeoff it formalized is now a foundational concept taught in every course.

Imagine a panel of expert judges scoring a diving competition. If two judges always give nearly identical scores (they're correlated), averaging all judges equally makes the final score swing wildly with tiny disagreements between that correlated pair.

OLS is that naive average — it trusts every judge equally, even the ones whose opinions are redundant. One noisy measurement from the correlated pair can hijack the entire result.

Ridge regression adds a gentle pull toward moderate scores for every judge. The correlated pair's extreme influence gets dampened. The final score might not be perfectly centered (it's slightly biased), but it's far more stable — and stability is what you need when predicting the next dive.

The problem: when predictors walk in lockstep

Ordinary Least Squares finds the coefficient vector β̂ that minimizes the sum of squared residuals. When predictors are nearly independent (orthogonal), OLS works beautifully. But in practice — housing prices predicted from square footage and number of rooms, or gene expression levels from overlapping pathways — predictors are correlated.

Correlation means the columns of the design matrix X point in similar directions. The matrix X'X becomes ill-conditioned: some of its eigenvalues shrink toward zero. Inverting a matrix with near-zero eigenvalues is like dividing by a number very close to zero — the result explodes. OLS amplifies along these fragile directions, producing coefficients that are mathematically unbiased but practically useless because they jump wildly from sample to sample.

Open in Lab
Drag the correlation slider to see how OLS coefficients become unstable as predictors grow more correlated. Ridge coefficients stay calm.
The demo wakes as you arrive…

Eigenvalues: the amplifiers of noise

To understand why correlated predictors cause trouble, think of X'X through its eigendecomposition. Each λj\lambda_j of X'X represents the amount of "signal energy" along its corresponding direction. OLS inverts X'X, so it divides by each eigenvalue. Directions with large λj\lambda_j are fine — dividing by a big number keeps things small. But directions with tiny λj\lambda_j get amplified enormously: Var(β^j)∝1/λj\text{Var}(\hat{\beta}_j) \propto 1/\lambda_j.

The (VIF) quantifies this: VIFj=1/(1−Rj2)\text{VIF}_j = 1/(1 - R_j^2) where Rj2R_j^2 is how well predictor jj is predicted by all other predictors. A VIF above 10 signals dangerous multicollinearity — the coefficient's variance is inflated tenfold.

Open in Lab
Watch how shrinking eigenvalues of X'X cause OLS variance to explode. The ridge addition (kI) puts a floor under every eigenvalue.
The demo wakes as you arrive…

The idea: add a ridge to the diagonal

The fix is elegant: before inverting X'X, add a small positive constant kk to every diagonal element. This is like raising the floor under all eigenvalues — even the tiny ones now have a minimum value of kk, so inverting them no longer explodes.

Geometrically, OLS finds the β that minimizes ∥y−Xβ∥2\|y - X\beta\|^2. Ridge regression adds a penalty: minimize ∥y−Xβ∥2+k∥β∥2\|y - X\beta\|^2 + k\|\beta\|^2. The penalty k∥β∥2k\|\beta\|^2 says "coefficients should not be too large" — it shrinks them toward zero by an amount controlled by kk. This is the earliest form of what we now call regularization.

β^ridge=(XTX+kI)−1XTy\hat{\beta}_{\text{ridge}} = (X^T X + kI)^{-1} X^T y
The ridge estimator — the core of the entire paper — X'X = the normal equations matrix · kI = a small identity "ridge" on the diagonal · the result: coefficients that are biased toward zero but far more stable than OLS

Think of it as a tug of war: one rope pulls β toward the least-squares solution (fit the data), the other pulls it toward zero (keep coefficients small). The parameter kk controls how hard the "stay small" rope pulls. When k=0k = 0, you get OLS. As kk grows, coefficients shrink more, bias increases, but variance drops — and there is a sweet spot where the total error (MSE) is minimized.

Open in Lab
Drag k from 0 to see the coefficient path shrink toward zero. Watch bias grow and variance drop — the MSE curve reveals the sweet spot.
The demo wakes as you arrive…

The theorem: bias can beat unbiasedness

The Gauss-Markov theorem guarantees OLS is the Best Linear Unbiased Estimator (BLUE). But "best among unbiased" is not the same as "best overall." The total prediction error (MSE) decomposes into two parts: MSE=Bias2+Variance\text{MSE} = \text{Bias}^2 + \text{Variance}. OLS has zero bias, but its variance can be enormous when eigenvalues are small. Ridge adds bias — the coefficient estimates are systematically pulled toward zero — but the variance drops so dramatically that the total MSE is lower.

Hoerl and Kennard's key result: for every finite β, there exists a k>0k > 0 such that MSE(β^ridge)<MSE(β^OLS)\text{MSE}(\hat{\beta}_{\text{ridge}}) < \text{MSE}(\hat{\beta}_{\text{OLS}}). In other words, accepting some bias always pays off when predictors are correlated.

MSE(β^ridge)=σ2∑j=1pλj(λj+k)2+k2∑j=1pαj2(λj+k)2\text{MSE}(\hat{\beta}_{\text{ridge}}) = \sigma^2 \sum_{j=1}^{p} \frac{\lambda_j}{(\lambda_j + k)^2} + k^2 \sum_{j=1}^{p} \frac{\alpha_j^2}{(\lambda_j + k)^2}
MSE decomposition — variance (first sum) + bias² (second sum) — λⱼ = eigenvalues of X'X · αⱼ = true coefficients in the eigenbasis · k shrinks the variance term (through (λⱼ+k)²) but grows the bias term (through k²) — the MSE curve is U-shaped with a minimum at the optimal k
Open in Lab
Watch the bias² and variance curves as k changes. Their sum (MSE) has a minimum — that's the optimal ridge parameter.
The demo wakes as you arrive…

The ridge trace: choosing k visually

Hoerl and Kennard proposed a diagnostic plot called the ridge trace: plot each coefficient β^j\hat{\beta}_j as a function of kk. At k=0k = 0 (OLS), coefficients may be large, erratic, or have wrong signs. As kk increases, coefficients shrink, stabilize, and often settle into physically reasonable values. The analyst picks the smallest kk at which the coefficients have stabilized.

Modern practice often replaces the ridge trace with — split the data, try many kk values, and pick the one that minimizes prediction error on held-out data. But the ridge trace remains valuable for understanding which coefficients are unstable and why.

Open in Lab
Watch coefficients stabilize as k increases. The vertical dashed line marks the optimal k from cross-validation.
The demo wakes as you arrive…

Geometry: the circle meets the ellipse

There is a beautiful geometric way to see ridge regression. OLS minimizes ∥y−Xβ∥2\|y - X\beta\|^2 — the contours of this loss in β-space are ellipses centered at the OLS solution. The ridge penalty k∥β∥2k\|\beta\|^2 says "stay inside a ball of radius tt around the origin." The ridge solution is where the smallest loss ellipse just touches the constraint ball.

Compare with Lasso: Lasso uses an L1 constraint (a diamond), whose corners sit on the axes — that's why Lasso sets some coefficients exactly to zero (feature selection). Ridge uses an L2 constraint (a circle/sphere), which has no corners — it shrinks coefficients smoothly but never to exact zero.

Open in Lab
The ellipse is the OLS loss contour, the circle is the ridge constraint, the diamond is Lasso. Drag k to see where the ellipse touches each shape.
The demo wakes as you arrive…

The same idea in code

Ridge regression from scratchpython

Simplified to show the idea — not the real implementation.

import numpy as np

def ridge(X, y, k):
    """Ridge regression: (X'X + kI)^{-1} X'y"""
    p = X.shape[1]
    return np.linalg.solve(X.T @ X + k * np.eye(p), X.T @ y)

def ols(X, y):
    """Ordinary Least Squares — ridge with k=0"""
    return ridge(X, y, k=0)

def ridge_mse(X, y, beta_true, k):
    """Compute MSE = bias² + variance for a given k"""
    p = X.shape[1]
    W = np.linalg.inv(X.T @ X + k * np.eye(p)) @ X.T
    beta_hat = W @ y
    bias_sq = np.sum((beta_hat - beta_true) ** 2)
    # Variance trace: σ² * trace(W @ W')
    return bias_sq  # simplified — full version adds variance term

# Try it: correlated predictors
rng = np.random.default_rng(42)
n, p = 100, 5
X = rng.standard_normal((n, p))
X[:, 1] = X[:, 0] + 0.01 * rng.standard_normal(n)  # near-collinear
beta_true = np.array([1, 2, 0.5, -1, 0.3])
y = X @ beta_true + 0.5 * rng.standard_normal(n)

print("OLS:  ", ols(X, y).round(2))        # unstable
print("Ridge:", ridge(X, y, k=1.0).round(2))  # stable

Connections: from Bayesian priors to deep learning

Ridge regression sits at a crossroads of many ideas. From a Bayesian perspective, it is equivalent to placing a Gaussian β∼N(0,σ2/k⋅I)\beta \sim N(0, \sigma^2/k \cdot I) on the coefficients and computing the mode. The penalty kk encodes a prior belief that coefficients should be moderate.

From a signal processing perspective, ridge is a Wiener filter — it optimally balances noise suppression against signal distortion.

In , weight decay (adding λ∥θ∥2\lambda \|\theta\|^2 to the loss) is exactly ridge regression applied to parameters. Every time you set weight_decay=0.01 in an , you are using the idea Hoerl and Kennard published in 1970.

Why it mattered: the regularization family tree

Ridge regression didn't just solve one practical problem — it introduced a principle that reshaped all of statistics and machine learning. The idea that you can systematically improve prediction by constraining model complexity became the foundation of modern model selection and the entire regularization toolkit.

  1. 1970

    Ridge Regression

    Hoerl & Kennard show that adding kI to X'X always improves MSE over OLS when predictors are correlated. L2 regularization is born.

  2. 1996

    Lasso (Tibshirani)

    Replace L2 penalty with L1. Some coefficients shrink exactly to zero — regularization now performs automatic feature selection.

  3. 2005

    Elastic Net

    Zou & Hastie combine L1 and L2 penalties, getting both sparsity (feature selection) and group stability (handling correlated features).

  4. 2012

    Deep Learning + Weight Decay

    AlexNet and successors adopt weight decay (L2 on all parameters) as a standard practice. Ridge's idea scales to billions of parameters.

  5. 2017

    AdamW — Decoupled Weight Decay

    Loshchilov & Hutter show that proper weight decay in Adam requires decoupling it from the gradient update. The ridge principle gets a modern correction.

Ridge vs Lasso: when to use which

Lasso (L1 penalty) and ridge (L2 penalty) both shrink coefficients, but they behave differently:

  • Ridge shrinks all coefficients proportionally. Use it when you believe all predictors contribute and the problem is multicollinearity (e.g., spectroscopy, genomics with pathway correlations).

  • Lasso drives some coefficients to exactly zero, performing feature selection. Use it when you suspect only a few predictors matter and want a sparse model.

  • Elastic Net combines both: α∥β∥1+(1−α)∥β∥22\alpha \|\beta\|_1 + (1-\alpha)\|\beta\|_2^2. Use it when you want both sparsity and grouping of correlated features.

Ridge is the "keep everyone but quiet them down" strategy; Lasso is "fire the redundant ones." In practice, try both and let cross-validation decide.

Open in Lab
Compare ridge (L2) and Lasso (L1) coefficient paths side by side. Notice Lasso sets coefficients to zero while ridge only approaches zero.
The demo wakes as you arrive…

CitationHoerl, Kennard. Ridge Regression: Biased Estimation for Nonorthogonal Problems. Technometrics, 1970.

Terms in this paper