Core ML1996intermediate10 min read
Regression Shrinkage and Selection via the Lasso
تقليص المعاملات وانتقاء المتغيرات بأسلوب Lasso
Tibshirani, R. — Journal of the Royal Statistical Society Series B
The problem
Ordinary fits data well but has no mechanism to control complexity. When there are many predictors — especially more than observations — OLS overfits wildly, producing coefficients that are large and unstable. Ridge (L2 penalty) shrinks coefficients toward zero but never sets any exactly to zero, so every predictor stays in the model and interpretation suffers. There was no method that could simultaneously shrink coefficients and perform automatic variable selection in a single convex .
The contribution
The Lasso (Least Absolute and Selection Operator): a new method that replaces Ridge's squared L2 penalty with an L1 penalty — the sum of absolute values of coefficients. This seemingly small change has a dramatic geometric consequence: the L1 constraint region is a diamond with sharp corners on the coordinate axes. When the elliptical contours of the least-squares touch this diamond, they almost always hit a corner, which means one or more coefficients land on exactly zero. The result is a sparse model that performs simultaneous shrinkage and automatic in a single convex optimization.
The impact
The Lasso became one of the most cited papers in statistics and machine learning. It launched an entire family of sparse methods: Elastic Net, Group Lasso, Adaptive Lasso, Fused Lasso, and Graphical Lasso. It made automatic selection a first-class citizen in regularized optimization, enabling practical modeling in genomics, finance, signal processing, and any domain where the number of variables exceeds the number of samples. The L1 penalty is now a standard tool in nearly every machine learning library.
Imagine you run a kitchen and have 50 spice jars on the rack, but most dishes only need 3–5 spices. Least squares grabs a pinch from every jar — even ones that add nothing — and the dish ends up muddled. Ridge regression keeps all 50 jars on the counter but uses smaller pinches, so the flavors are milder but still cluttered.
Lasso is a ruthless sous-chef: it looks at each spice jar and if the spice doesn't clearly improve the dish, it puts the jar back on the shelf — cap on, fully removed. Only the spices that genuinely earn their place stay in the recipe. The secret is the L1 budget: instead of limiting total squared (a smooth bowl), it limits total absolute weight (a diamond with sharp corners) — and those corners are where entire spices get zeroed out.
The problem: too many predictors, too little discipline
In linear regression we model an outcome as a weighted sum of predictors plus noise. Ordinary Least Squares (OLS) finds the weights that minimize the sum of squared residuals. When is small relative to (the number of observations), this works well.
But in many real problems — genomics with thousands of genes, finance with hundreds of indicators, text with millions of word features — can rival or exceed . OLS then has infinitely many perfect-fit solutions, all of them useless for prediction: the model memorizes noise instead of learning signal. Even when , including irrelevant predictors inflates and makes the model fragile.
Ridge regression: shrinkage without selection
Before the Lasso, the standard remedy for was Ridge regression (Hoerl & Kennard, 1970). Ridge adds a penalty proportional to the sum of squared coefficients:
Ridge stabilizes predictions and reduces variance, but it never produces a sparse model. If you started with 1,000 gene expression features, Ridge gives you 1,000 shrunken coefficients — none zero, none dropped. Interpretation becomes a wall of small numbers with no clear answer to "which genes matter?"
The Lasso: one change, everything changes
Tibshirani's insight was deceptively simple: replace the squared L2 penalty with an L1 penalty — the sum of absolute values instead of squared values:
An equivalent way to state the Lasso is as a constrained optimization: minimize the sum of squared errors subject to . This constraint defines the feasible region — and its shape is the key to everything.
The geometry: why diamonds create zeros
Think of a two-coefficient problem on a flat plane. The least-squares loss defines elliptical contours centered on the OLS solution — each contour is a set of values with the same total error.
Ridge's constraint is a circle. When the smallest ellipse that touches the circle makes contact, it almost always touches the smooth curve at a point where both coefficients are nonzero.
The Lasso's constraint is a diamond — a square rotated 45°, with sharp corners sitting right on the axes. When the loss ellipse expands until it just touches this diamond, it preferentially hits a corner, and at a corner one coordinate is exactly zero. That coefficient has been eliminated. In higher dimensions the diamond becomes a cross-polytope with far more corners than faces, so the probability of landing on a corner (a sparse solution) is even higher.
Soft thresholding: the analytic view
When the design matrix columns are orthonormal (each predictor is uncorrelated with the others and has unit variance), the Lasso solution has a beautiful closed form called :
The regularization path: watching features enter and leave
As decreases from a large value (where all coefficients are zero) toward zero (the OLS solution), coefficients "enter" the model one by one. The regularization path — the curve of each coefficient as a function of — is piecewise linear for the Lasso, a remarkable property discovered by Efron et al. (2004) through the LARS algorithm.
This path reveals the story of variable importance: the first coefficient to become nonzero is the strongest predictor, the next is the second-strongest given the first, and so on. then picks the that gives the best out-of-sample performance, yielding a principled sparse model.
Solving the Lasso: coordinate descent
The L1 penalty makes the objective non-smooth — there is no simple closed-form solution when predictors are correlated. The modern workhorse algorithm is : optimize one coefficient at a time while holding the others fixed.
For a single coefficient the subproblem is one-dimensional, and the solution is just the soft-thresholding formula applied to the partial residual. Cycling through all coefficients repeatedly converges to the global optimum because the objective is convex. Coordinate descent is remarkably fast — it handles millions of predictors in seconds — and is the engine behind the popular glmnet library (Friedman, Hastie, Tibshirani 2010).
Simplified to show the idea — not the real implementation.
import numpy as np
def soft_threshold(z, lam):
"""Soft thresholding: S(z, λ) = sign(z) * max(|z| - λ, 0)"""
return np.sign(z) * np.maximum(np.abs(z) - lam, 0.0)
def lasso_cd(X, y, lam, max_iter=1000, tol=1e-6):
"""Coordinate descent for Lasso.
X: (n, p) design matrix (standardized)
y: (n,) response (centered)
"""
n, p = X.shape
beta = np.zeros(p)
r = y.copy() # residual = y - X @ beta
for iteration in range(max_iter):
beta_old = beta.copy()
for j in range(p):
# Add back j-th contribution
r += X[:, j] * beta[j]
# Compute OLS estimate for j-th predictor
z_j = X[:, j] @ r / n
# Apply soft thresholding
beta[j] = soft_threshold(z_j, lam)
# Update residual
r -= X[:, j] * beta[j]
# Check convergence
if np.max(np.abs(beta - beta_old)) < tol:
break
return betaThe Bayesian lens: Laplace priors
The L1 penalty has a beautiful probabilistic interpretation. Placing an independent Laplace on each coefficient and computing the MAP (maximum a posteriori) estimate is mathematically identical to solving the Lasso. The Laplace distribution is sharply peaked at zero with heavy tails — it encodes the belief that most coefficients should be zero, but the few that are nonzero can be large. Compare this with Ridge's Gaussian prior, which is smoothly peaked at zero and decays quickly — it believes coefficients are small but never exactly zero.
Side by side: OLS vs Ridge vs Lasso
The three methods form a spectrum of discipline. OLS is fully permissive — use any coefficients you want. Ridge adds a smooth budget that shrinks everything proportionally. Lasso adds a sharp budget that eliminates the weakest contributors entirely. The interactive comparison below makes this concrete.
Limitations and extensions
The Lasso is not perfect. It has two well-known limitations:
-
Correlated predictors. When two features are highly correlated, Lasso arbitrarily picks one and zeros out the other. It doesn't "share credit" — it selects a representative from each correlated cluster rather than keeping the whole group.
-
The p > n case. When predictors outnumber observations, Lasso can select at most variables, even if more are truly relevant.
These limitations motivated important extensions. The Elastic Net (Zou & Hastie, 2005) combines L1 and L2 penalties, inheriting Lasso's sparsity while handling correlated features gracefully. The Adaptive Lasso (Zou, 2006) assigns different penalty weights to each coefficient, achieving oracle properties. The Group Lasso (Yuan & Lin, 2006) selects or eliminates entire groups of related variables together.
The Lasso family tree
1970
Ridge Regression
Hoerl & Kennard introduced L2 regularization for ill-conditioned regression. Shrinks all coefficients but zeros none.
1996
The Lasso
Tibshirani replaced L2 with L1, achieving simultaneous shrinkage and variable selection in one convex optimization.
2004
LARS Algorithm
Efron, Hastie, Johnstone, Tibshirani developed Least Angle Regression, efficiently computing the entire Lasso path and revealing its piecewise linearity.
2005
Elastic Net
Zou & Hastie combined L1 and L2 penalties, fixing Lasso's instability with correlated features while preserving sparsity.
2006
Adaptive Lasso & Group Lasso
Zou introduced adaptive weights for oracle properties. Yuan & Lin introduced group-level selection for structured variables.
2010
glmnet
Friedman, Hastie, Tibshirani released the coordinate descent implementation that made Lasso practical at scale — millions of predictors in seconds.
The same idea in code
Simplified to show the idea — not the real implementation.
from sklearn.linear_model import LassoCV
from sklearn.preprocessing import StandardScaler
from sklearn.datasets import make_regression
import numpy as np
# Generate data: 200 samples, 50 features, only 5 are truly relevant
X, y, true_coef = make_regression(
n_samples=200, n_features=50, n_informative=5,
noise=10, coef=True, random_state=42
)
# Always standardize before regularizing
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)
# LassoCV automatically picks the best λ via cross-validation
model = LassoCV(cv=5, random_state=42)
model.fit(X_scaled, y)
nonzero = np.sum(model.coef_ != 0)
print(f"Features selected: {nonzero} / 50") # ~5, matching truth
print(f"Best λ: {model.alpha_:.4f}")
print(f"R² score: {model.score(X_scaled, y):.3f}")
# Compare: the 5 largest coefficients match the 5 truly informative featuresTest your understanding
CitationTibshirani, R.. Regression Shrinkage and Selection via the Lasso. Journal of the Royal Statistical Society Series B, 1996.
Terms in this paper
- Regularizationالضبط الهيكلي
- Sparsityالتناثر البنيوي للمصفوفات
- Feature Selectionانتقاء وتحديد السمات الأهم
- Shrinkageالانكماش
- Least Squaresالمربعات الصغرى
- Weight Decayاضمحلال الأوزان
- L2 Regularizationالضبط الهيكلي L2
- Regressionالانحدار الإحصائي
- Overfittingفرط التخصيص
- Bias-Variance Tradeoffالموازنة بين الانحياز والتباعد
- Cross-Validationالتحقق المتقاطع من الأداء
- Normالمعيار التفاضلي
- Convex Functionالدالة المحدَّبة
- Gradient Descentالانحدار التدريجي
- Coordinate Descentالنزول الإحداثي