Core ML1996beginner10 min read

Bagging Predictors

تكييس المتنبِّئات

Breiman, L. — Machine Learning

The problem

A single is expressive but fragile: remove a few examples or add some noise, and the tree can look completely different. This instability means high — the model overfits to the specific quirks of the training set rather than learning the true pattern. In 1996, there was no principled, general-purpose way to tame this variance without sacrificing the expressiveness that makes trees powerful.

The contribution

Bootstrap Aggregating (): draw B bootstrap samples from the training set (each the same size, sampled with replacement), train an independent predictor on each, then average () or vote (). Breiman proved that aggregation always reduces mean-squared error for unstable predictors, showed 20–47% drops in misclassification on real datasets, and introduced the estimate — a free validation score with no held-out set required.

The impact

Bagging is the intellectual ancestor of Random Forests (Breiman, 2001) and one half of the -variance management toolkit that includes . It demonstrated that combining weak, unstable learners could rival or beat carefully engineered single models. Today, every tree-based — , Extra-Trees, and even the early rounds of — traces its variance-reduction principle back to this paper.

Imagine you're trying to guess the number of jelly beans in a jar. One person's guess might be way off — maybe they misjudge the jar's depth, or the bean size. But if you ask fifty people and take the average, the individual errors — some too high, some too low — tend to cancel out, and the average lands remarkably close to the truth.

Bagging does the same thing with models. Instead of trusting one tree that might have memorized quirks in the data, it creates fifty slightly different training sets, builds a separate tree on each, and averages their answers. Each tree overfits in its own unique way, but the crowd's average smooths out the noise.

The problem: one tree, many possible shapes

Decision trees are powerful because they can capture complex, nonlinear decision boundaries without any assumptions about the data's shape. But this flexibility comes at a cost: instability. Remove a handful of training examples and the tree may split on entirely different variables, producing completely different predictions.

This is the variance component of the . A deep tree has low bias (it can fit any pattern) but high variance (it fits the noise too). helps, but the fundamental fragility remains — the tree is too sensitive to the particular training set it happened to see.

Open in Lab
Click "Resample" to draw a new bootstrap sample and watch how the decision boundary changes each time. This is the instability that bagging exploits.
The demo wakes as you arrive…

The bootstrap: making many datasets from one

Ideally, we would collect many independent training sets from the real world, train a model on each, and average. But we only have one training set. The bootstrap is a statistical trick that simulates having multiple datasets: draw NN samples with replacement from our NN training examples. Some examples appear multiple times, others are left out entirely.

On average, each bootstrap sample contains about 63.2% of the unique original examples (the rest are duplicates). The ~36.8% of examples not selected in a given draw are called out-of-bag (OOB) examples — and they become a free , as we'll see later.

Open in Lab
Click "Draw Sample" to generate a bootstrap sample. Green = selected (may appear multiple times). Gray = left out (out-of-bag).
The demo wakes as you arrive…

The algorithm: bootstrap, train, aggregate

Bagging is beautifully simple — three steps repeated BB times, then one aggregation:

Step 1 — Bootstrap. Draw a sample L(b)L^{(b)} of size NN with replacement from the training set LL.

Step 2 — Train. Fit a full, unpruned predictor φ(x,L(b))\varphi(x, L^{(b)}) on each bootstrap sample. No need to regularize — we want each tree to overfit so that the ensemble has low bias.

Step 3 — Aggregate. Combine the BB predictors: average for regression, majority vote for classification.

φB(x)=1B∑b=1Bφ(x,L(b))\varphi_B(x) = \frac{1}{B} \sum_{b=1}^{B} \varphi(x, L^{(b)})
Bagging prediction (regression) — average B bootstrap predictors — Each training set is created by bootstrap sampling from the original data. A separate model is trained on each bootstrap sample. The final regression prediction is the simple average of all model predictions. For classification, the average is replaced by a majority vote among the models.
Open in Lab
Watch the full bagging pipeline: bootstrap samples → independent trees → aggregated prediction. Add more trees and watch the ensemble prediction stabilize.
The demo wakes as you arrive…

Why it works: averaging reduces variance

The theoretical justification is elegant. Consider the ideal case where we have infinitely many independent training sets drawn from the true distribution PP. Define the aggregated predictor:

φA(x,P)=EL[φ(x,L)]\varphi_A(x, P) = E_L[\varphi(x, L)]
Ideal aggregated predictor — the expected model over all possible training sets — If we could average over all possible training sets drawn from the true distribution, this is the predictor we would get. Bagging approximates this using bootstrap samples.

Breiman proved a key inequality using Jensen's inequality: for any squared-error predictor, the aggregated predictor is always at least as good as the average single predictor:

EL[(Y−φ(X,L))2]  ≥  (Y−EL[φ(X,L)])2E_L\bigl[(Y - \varphi(X, L))^2\bigr] \;\geq\; \bigl(Y - E_L[\varphi(X, L)]\bigr)^2
Jensen's inequality — aggregation never hurts — The left side represents the expected prediction error of a single model trained on a random training set. The right side represents the prediction error of the averaged ensemble prediction. The difference between the two is exactly the variance of the model predictions across different training sets. The more unstable the base learner, the larger this variance becomes, and the more bagging can improve performance.
Open in Lab
Drag the "Instability" slider and watch how bagging affects bias and variance. High instability = big variance reduction. Low instability = no benefit.
The demo wakes as you arrive…

Free validation: the out-of-bag estimate

Each bootstrap sample leaves out about 36.8% of the training examples. Those left-out examples were never seen by that particular tree — so we can use them as an unbiased test set for that tree. For each training example xix_i, collect predictions only from the trees whose bootstrap sample did not include xix_i, and average those predictions. This gives the out-of-bag (OOB) error estimate — a nearly unbiased estimate of test error that requires no separate validation set or .

The OOB estimate is one of the most practical contributions of the bagging framework. It gives you a reliable error estimate "for free" during training, making bagging especially convenient when data is scarce.

Open in Lab
Each row is a bootstrap sample. Highlighted cells show which training examples were left out. Follow one example across rows to see which trees provide its OOB prediction.
The demo wakes as you arrive…

Experimental results: 20–47% error reduction

Breiman tested bagging with classification trees on seven datasets and regression trees on five datasets. The results were striking: bagging reduced misclassification rates by 20–47% across all classification benchmarks, and reduced by 22–46% for regression. The largest gains came from the datasets where single trees were most unstable.

Open in Lab
Breiman's original classification results. Toggle between classification and regression to see the error reduction.
The demo wakes as you arrive…

A revealing control experiment was bagging k-nearest neighbors — a stable method. As predicted by the theory, the misclassification rates were identical before and after bagging. The algorithm simply had no variance to reduce.

How many bootstrap replicates?

Breiman tested 10, 25, 50, and 100 bootstrap replicates on the waveform dataset. With just 10 trees, most of the improvement was already captured (21.8% vs 29.0% for a single tree). At 25 trees the error was 19.5%, and 50 or 100 trees added essentially nothing more (19.4%).

The practical message: bagging converges quickly. Unlike gradient boosting, where adding more rounds can overfit, bagging with more trees never hurts — it just stops helping after a point. The returns plateau but never reverse.

Open in Lab
Drag the slider to add trees and watch the ensemble error converge. Notice how most improvement happens in the first 10–25 trees.
The demo wakes as you arrive…

The algorithm in code

Bagging from scratch — classification and regressionpython

Simplified to show the idea — not the real implementation.

import numpy as np
from sklearn.tree import DecisionTreeClassifier, DecisionTreeRegressor
from collections import Counter

def bootstrap_sample(X, y, rng):
    """Draw N samples with replacement from (X, y)."""
    n = len(X)
    indices = rng.choice(n, size=n, replace=True)
    return X[indices], y[indices], indices

def bagging_predict_regression(X_train, y_train, X_test, B=50, seed=42):
    """Train B trees on bootstrap samples, average predictions."""
    rng = np.random.default_rng(seed)
    predictions = np.zeros((B, len(X_test)))

    for b in range(B):
        X_boot, y_boot, _ = bootstrap_sample(X_train, y_train, rng)
        tree = DecisionTreeRegressor()
        tree.fit(X_boot, y_boot)
        predictions[b] = tree.predict(X_test)

    return predictions.mean(axis=0)        # average over B trees

def bagging_predict_classification(X_train, y_train, X_test, B=50, seed=42):
    """Train B trees on bootstrap samples, majority vote."""
    rng = np.random.default_rng(seed)
    all_preds = []

    for b in range(B):
        X_boot, y_boot, _ = bootstrap_sample(X_train, y_train, rng)
        tree = DecisionTreeClassifier()
        tree.fit(X_boot, y_boot)
        all_preds.append(tree.predict(X_test))

    # Majority vote for each test example
    all_preds = np.array(all_preds)         # (B, n_test)
    final = []
    for i in range(len(X_test)):
        votes = Counter(all_preds[:, i])
        final.append(votes.most_common(1)[0][0])
    return np.array(final)

# That's bagging. Random Forest adds one twist: at each split,
# only a random subset of features is considered. This decorrelates
# the trees further, squeezing out even more variance reduction.

What bagging cannot do

Bagging reduces variance but does not reduce bias. If every individual tree systematically misses the true pattern in the same way (e.g., all trees underfit because the model class is too simple), averaging them will not fix the error.

This is the fundamental difference between bagging and boosting. Boosting focuses each new model on the mistakes of the previous ones, gradually reducing bias. Bagging builds independent models in parallel, reducing variance. Both are ensemble methods, but they attack different components of the error.

Another limitation: . A single tree is easy to inspect and explain. An ensemble of fifty trees is a black box. Breiman acknowledged this tradeoff: "What one loses, with the trees, is a simple and interpretable structure. What one gains is increased accuracy."

Legacy: from bagging to forests

Bagging established two principles that shaped the next two decades of machine learning. First, that combining many weak, learners can produce a strong one — not by making each one better, but by diversifying their errors so they cancel. Second, that the bootstrap is a powerful engine for ensemble diversity — each resampled dataset emphasizes different aspects of the data.

  1. 1984

    CART (Classification and Regression Trees)

    Breiman, Friedman, Olshen, and Stone introduce CART — the foundational tree algorithm that bagging would later tame.

  2. 1996

    Bagging Predictors

    Breiman proposes bootstrap aggregating. Proves variance reduction, demonstrates 20–47% error reduction on real datasets. Introduces the OOB estimate.

  3. 1997

    AdaBoost

    Freund and Schapire introduce AdaBoost — the complementary ensemble approach that reduces bias sequentially rather than variance in parallel.

  4. 2001

    Random Forests

    Breiman extends bagging with random feature selection at each split, decorrelating the trees for even greater variance reduction. Becomes one of the most widely used algorithms in machine learning.

  5. 2016

    XGBoost dominates Kaggle

    Gradient boosted trees — which owe their ensemble philosophy to bagging and boosting — become the default for tabular data competitions.

Random Forests are bagging plus one extra trick: at each split, only a random subset of features is considered. This decorrelates the trees, squeezing out variance that bagging alone leaves on the table. AdaBoost took the opposite path — sequential, bias-reducing ensembles — and led to the gradient boosting family. Together, these two branches gave us the most powerful algorithms for that exist today.

CitationBreiman, L.. Bagging Predictors. Machine Learning, 1996.

Terms in this paper