Core ML1996intermediate10 min read

Regression Shrinkage and Selection via the Lasso

تقليص المعاملات وانتقاء المتغيرات بأسلوب Lasso

Tibshirani, R. — Journal of the Royal Statistical Society Series B

The problem

Ordinary fits data well but has no mechanism to control complexity. When there are many predictors — especially more than observations — OLS overfits wildly, producing coefficients that are large and unstable. Ridge (L2 penalty) shrinks coefficients toward zero but never sets any exactly to zero, so every predictor stays in the model and interpretation suffers. There was no method that could simultaneously shrink coefficients and perform automatic variable selection in a single convex .

The contribution

The Lasso (Least Absolute and Selection Operator): a new method that replaces Ridge's squared L2 penalty with an L1 penalty — the sum of absolute values of coefficients. This seemingly small change has a dramatic geometric consequence: the L1 constraint region is a diamond with sharp corners on the coordinate axes. When the elliptical contours of the least-squares touch this diamond, they almost always hit a corner, which means one or more coefficients land on exactly zero. The result is a sparse model that performs simultaneous shrinkage and automatic in a single convex optimization.

The impact

The Lasso became one of the most cited papers in statistics and machine learning. It launched an entire family of sparse methods: Elastic Net, Group Lasso, Adaptive Lasso, Fused Lasso, and Graphical Lasso. It made automatic selection a first-class citizen in regularized optimization, enabling practical modeling in genomics, finance, signal processing, and any domain where the number of variables exceeds the number of samples. The L1 penalty is now a standard tool in nearly every machine learning library.

Imagine you run a kitchen and have 50 spice jars on the rack, but most dishes only need 3–5 spices. Least squares grabs a pinch from every jar — even ones that add nothing — and the dish ends up muddled. Ridge regression keeps all 50 jars on the counter but uses smaller pinches, so the flavors are milder but still cluttered.

Lasso is a ruthless sous-chef: it looks at each spice jar and if the spice doesn't clearly improve the dish, it puts the jar back on the shelf — cap on, fully removed. Only the spices that genuinely earn their place stay in the recipe. The secret is the L1 budget: instead of limiting total squared (a smooth bowl), it limits total absolute weight (a diamond with sharp corners) — and those corners are where entire spices get zeroed out.

The problem: too many predictors, too little discipline

In linear regression we model an outcome yy as a weighted sum of pp predictors x1,x2,…,xpx_1, x_2, \dots, x_p plus noise. Ordinary Least Squares (OLS) finds the weights β^\hat{\beta} that minimize the sum of squared residuals. When pp is small relative to nn (the number of observations), this works well.

But in many real problems — genomics with thousands of genes, finance with hundreds of indicators, text with millions of word features — pp can rival or exceed nn. OLS then has infinitely many perfect-fit solutions, all of them useless for prediction: the model memorizes noise instead of learning signal. Even when p<np < n, including irrelevant predictors inflates and makes the model fragile.

Open in Lab
Drag the slider to increase predictors (p). Watch how OLS training error drops to zero while test error explodes.
The demo wakes as you arrive…

Ridge regression: shrinkage without selection

Before the Lasso, the standard remedy for was Ridge regression (Hoerl & Kennard, 1970). Ridge adds a penalty proportional to the sum of squared coefficients:

β^ridge=arg⁡min⁡β{∑i=1n(yi−∑jxijβj)2+λ∑j=1pβj2}\hat{\beta}^{\text{ridge}} = \arg\min_{\beta} \left\{ \sum_{i=1}^{n} \left( y_i - \sum_{j} x_{ij} \beta_j \right)^2 + \lambda \sum_{j=1}^{p} \beta_j^2 \right\}
Ridge regression objective — L2 penalty — The λΣβ²_j term pushes all coefficients toward zero. Larger λ = more shrinkage. But no coefficient ever reaches exactly zero — Ridge keeps every predictor in the model.

Ridge stabilizes predictions and reduces variance, but it never produces a sparse model. If you started with 1,000 gene expression features, Ridge gives you 1,000 shrunken coefficients — none zero, none dropped. Interpretation becomes a wall of small numbers with no clear answer to "which genes matter?"

The Lasso: one change, everything changes

Tibshirani's insight was deceptively simple: replace the squared L2 penalty with an L1 penalty — the sum of absolute values instead of squared values:

β^lasso=arg⁡min⁡β{12n∑i=1n(yi−∑jxijβj)2+λ∑j=1p∣βj∣}\hat{\beta}^{\text{lasso}} = \arg\min_{\beta} \left\{ \frac{1}{2n} \sum_{i=1}^{n} \left( y_i - \sum_{j} x_{ij} \beta_j \right)^2 + \lambda \sum_{j=1}^{p} |\beta_j| \right\}
The Lasso objective — L1 penalty — Replace β²_j with |β_j|. The absolute value has a kink at zero — its derivative is ±1 everywhere except at zero, where it is undefined. This non-smoothness at zero is exactly what forces coefficients to land on zero.

An equivalent way to state the Lasso is as a constrained optimization: minimize the sum of squared errors subject to ∑j∣βj∣≤t\sum_j |\beta_j| \leq t. This constraint defines the feasible region — and its shape is the key to everything.

The geometry: why diamonds create zeros

Think of a two-coefficient problem on a flat plane. The least-squares loss defines elliptical contours centered on the OLS solution — each contour is a set of (β1,β2)(\beta_1, \beta_2) values with the same total error.

Ridge's constraint β12+β22≤t\beta_1^2 + \beta_2^2 \leq t is a circle. When the smallest ellipse that touches the circle makes contact, it almost always touches the smooth curve at a point where both coefficients are nonzero.

The Lasso's constraint ∣β1∣+∣β2∣≤t|\beta_1| + |\beta_2| \leq t is a diamond — a square rotated 45°, with sharp corners sitting right on the axes. When the loss ellipse expands until it just touches this diamond, it preferentially hits a corner, and at a corner one coordinate is exactly zero. That coefficient has been eliminated. In higher dimensions the diamond becomes a cross-polytope with far more corners than faces, so the probability of landing on a corner (a sparse solution) is even higher.

Open in Lab
Drag λ to see how the constraint region (diamond vs circle) intersects the loss contours. Notice the Lasso solution snaps to the corner.
The demo wakes as you arrive…

Soft thresholding: the analytic view

When the design matrix columns are orthonormal (each predictor is uncorrelated with the others and has unit variance), the Lasso solution has a beautiful closed form called :

β^jlasso=sign(β^jOLS)⋅max⁡(∣β^jOLS∣−λ, 0)\hat{\beta}_j^{\text{lasso}} = \text{sign}(\hat{\beta}_j^{\text{OLS}}) \cdot \max\left(|\hat{\beta}_j^{\text{OLS}}| - \lambda, \ 0\right)
Soft thresholding — each coefficient individually — If the OLS estimate is smaller than λ in absolute value, the Lasso sets it to exactly zero. Otherwise it shrinks it toward zero by the fixed amount λ. Compare with Ridge which scales every coefficient by a factor 1/(1+λ) — it never reaches zero.
Open in Lab
Drag λ to see how soft thresholding zeros out small coefficients while shrinking large ones.
The demo wakes as you arrive…

The regularization path: watching features enter and leave

As λ\lambda decreases from a large value (where all coefficients are zero) toward zero (the OLS solution), coefficients "enter" the model one by one. The regularization path — the curve of each coefficient as a function of λ\lambda — is piecewise linear for the Lasso, a remarkable property discovered by Efron et al. (2004) through the LARS algorithm.

This path reveals the story of variable importance: the first coefficient to become nonzero is the strongest predictor, the next is the second-strongest given the first, and so on. then picks the λ\lambda that gives the best out-of-sample performance, yielding a principled sparse model.

Open in Lab
Drag the λ slider to watch coefficients enter and leave the model along the regularization path.
The demo wakes as you arrive…

Solving the Lasso: coordinate descent

The L1 penalty makes the objective non-smooth — there is no simple closed-form solution when predictors are correlated. The modern workhorse algorithm is : optimize one coefficient at a time while holding the others fixed.

For a single coefficient the subproblem is one-dimensional, and the solution is just the soft-thresholding formula applied to the partial residual. Cycling through all pp coefficients repeatedly converges to the global optimum because the objective is convex. Coordinate descent is remarkably fast — it handles millions of predictors in seconds — and is the engine behind the popular glmnet library (Friedman, Hastie, Tibshirani 2010).

Coordinate descent for the Lassopython

Simplified to show the idea — not the real implementation.

import numpy as np

def soft_threshold(z, lam):
    """Soft thresholding: S(z, λ) = sign(z) * max(|z| - λ, 0)"""
    return np.sign(z) * np.maximum(np.abs(z) - lam, 0.0)

def lasso_cd(X, y, lam, max_iter=1000, tol=1e-6):
    """Coordinate descent for Lasso.
    X: (n, p) design matrix (standardized)
    y: (n,) response (centered)
    """
    n, p = X.shape
    beta = np.zeros(p)
    r = y.copy()               # residual = y - X @ beta

    for iteration in range(max_iter):
        beta_old = beta.copy()
        for j in range(p):
            # Add back j-th contribution
            r += X[:, j] * beta[j]
            # Compute OLS estimate for j-th predictor
            z_j = X[:, j] @ r / n
            # Apply soft thresholding
            beta[j] = soft_threshold(z_j, lam)
            # Update residual
            r -= X[:, j] * beta[j]

        # Check convergence
        if np.max(np.abs(beta - beta_old)) < tol:
            break
    return beta

The Bayesian lens: Laplace priors

The L1 penalty has a beautiful probabilistic interpretation. Placing an independent Laplace p(βj)∝exp⁡(−λ∣βj∣)p(\beta_j) \propto \exp(-\lambda|\beta_j|) on each coefficient and computing the MAP (maximum a posteriori) estimate is mathematically identical to solving the Lasso. The Laplace distribution is sharply peaked at zero with heavy tails — it encodes the belief that most coefficients should be zero, but the few that are nonzero can be large. Compare this with Ridge's Gaussian prior, which is smoothly peaked at zero and decays quickly — it believes coefficients are small but never exactly zero.

Open in Lab
Compare the Laplace prior (L1 / Lasso) with the Gaussian prior (L2 / Ridge). Notice how the Laplace peak at zero is sharper.
The demo wakes as you arrive…

Side by side: OLS vs Ridge vs Lasso

The three methods form a spectrum of discipline. OLS is fully permissive — use any coefficients you want. Ridge adds a smooth budget that shrinks everything proportionally. Lasso adds a sharp budget that eliminates the weakest contributors entirely. The interactive comparison below makes this concrete.

Open in Lab
Toggle between OLS, Ridge, and Lasso to see coefficient profiles. Lasso zeros out irrelevant features.
The demo wakes as you arrive…

Limitations and extensions

The Lasso is not perfect. It has two well-known limitations:

  • Correlated predictors. When two features are highly correlated, Lasso arbitrarily picks one and zeros out the other. It doesn't "share credit" — it selects a representative from each correlated cluster rather than keeping the whole group.

  • The p > n case. When predictors outnumber observations, Lasso can select at most nn variables, even if more are truly relevant.

These limitations motivated important extensions. The Elastic Net (Zou & Hastie, 2005) combines L1 and L2 penalties, inheriting Lasso's sparsity while handling correlated features gracefully. The Adaptive Lasso (Zou, 2006) assigns different penalty weights to each coefficient, achieving oracle properties. The Group Lasso (Yuan & Lin, 2006) selects or eliminates entire groups of related variables together.

The Lasso family tree

  1. 1970

    Ridge Regression

    Hoerl & Kennard introduced L2 regularization for ill-conditioned regression. Shrinks all coefficients but zeros none.

  2. 1996

    The Lasso

    Tibshirani replaced L2 with L1, achieving simultaneous shrinkage and variable selection in one convex optimization.

  3. 2004

    LARS Algorithm

    Efron, Hastie, Johnstone, Tibshirani developed Least Angle Regression, efficiently computing the entire Lasso path and revealing its piecewise linearity.

  4. 2005

    Elastic Net

    Zou & Hastie combined L1 and L2 penalties, fixing Lasso's instability with correlated features while preserving sparsity.

  5. 2006

    Adaptive Lasso & Group Lasso

    Zou introduced adaptive weights for oracle properties. Yuan & Lin introduced group-level selection for structured variables.

  6. 2010

    glmnet

    Friedman, Hastie, Tibshirani released the coordinate descent implementation that made Lasso practical at scale — millions of predictors in seconds.

The same idea in code

Lasso with scikit-learn — from data to sparse modelpython

Simplified to show the idea — not the real implementation.

from sklearn.linear_model import LassoCV
from sklearn.preprocessing import StandardScaler
from sklearn.datasets import make_regression
import numpy as np

# Generate data: 200 samples, 50 features, only 5 are truly relevant
X, y, true_coef = make_regression(
    n_samples=200, n_features=50, n_informative=5,
    noise=10, coef=True, random_state=42
)

# Always standardize before regularizing
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)

# LassoCV automatically picks the best λ via cross-validation
model = LassoCV(cv=5, random_state=42)
model.fit(X_scaled, y)

nonzero = np.sum(model.coef_ != 0)
print(f"Features selected: {nonzero} / 50")      # ~5, matching truth
print(f"Best λ: {model.alpha_:.4f}")
print(f"R² score: {model.score(X_scaled, y):.3f}")

# Compare: the 5 largest coefficients match the 5 truly informative features

Test your understanding

CitationTibshirani, R.. Regression Shrinkage and Selection via the Lasso. Journal of the Royal Statistical Society Series B, 1996.

Terms in this paper