Classical MLintermediateGPU optional~30 minColab

Leakage, Staged

التسرّب، على مراحل

A leak changes what you believe, not what you ship

A hospital wants a model that flags, at discharge, the patients likely to be readmitted within 30 days. The care team can phone 20% of them. You have two years of admissions. You build an honest evaluation, then take four shortcuts that real projects take every week, one at a time. Each shortcut raises the cross-validated score. The admissions from the following year, from patients the model has never met, stay sealed until the end. When they are opened, three of the four shortcuts turn out to have changed nothing but the number on the report, and the fourth does something the report's own metric barely registers: the alert goes silent.

The goal

Score one model under five protocols, each adding a leak to the last, and compare every cross-validated score with performance on sealed, later admissions from new patients. Show that the honest protocol's estimate tracks reality, identify which leaks inflate only the estimate and which change the model, and measure what the leaked column does to the alert once it is live.

Colab opens a read-only copy. Save a copy to Drive to keep your edits.

The notebook needs a keyboard — best opened on a desktop.

The papers behind this

A model that leaks does not look broken. It looks excellent. Nothing errors, the folds agree with each other, and the score is the best the project has seen. That is why leakage is the most common way real projects fail: every warning sign looks like success.

This workshop stages that failure on purpose. It begins with an evaluation that is honest, then adds four leaks one at a time. Each leak is a single line of code that an analyst might write on an ordinary day. Every stage is scored twice: once by its own , and once on data that was sealed before any of it began.

The printed code hash should match the one on this page.

import azimuth_nb as azimuth

env = azimuth.setup(SLUG, lang=LANG, profile=PROFILE)
Workshop code
Leakage, Staged
Tesla T4 · 14.6 GB · 12.7 GB RAM · PyTorch 2.11.0+cu130
profile: free
ready · dev_patients=800, holdout_patients=1500, visits_mean=3.5, n_markers=400, k_select=20, folds=5, trees=300, alert_rate=0.2, shuffles=3, seed=17
code · 4ac5ee59710b85e9

The data is synthetic, generated so that the true answer is known. Each patient has a hidden frailty that raises the risk of every one of their admissions. It is never a column, but it makes one patient's visits resemble each other. The real signal sits in eight clinical columns. Around them sit hundreds of lab markers that carry no signal at all, the way a wide hospital extract does.

One more column comes from the warehouse: care-coordination calls logged in the 30 days after discharge. Readmitted patients receive many of these calls, because the readmission is what triggers them. At the moment the model has to decide, no call has happened yet.

The sealed holdout is built the way the model will meet the world: a later year, patients never seen before, and every column as it exists at discharge. Its readmission rate is close to the development rate (35.400% against 38.200%). The world does not change between the two periods, by design, so any collapse later in this workshop cannot be blamed on distribution shift.

Two cohorts with no patient in common. Note that the post-discharge column exists in both, but is zero in the holdout.

import numpy as np
import pandas as pd

cfg = env.cfg
CLINICAL = ["age", "chronic", "los", "n_meds", "lab_a", "lab_b", "emergency", "prior"]
MARKERS = [f"marker_{i:03d}" for i in range(cfg["n_markers"])]
POST_DISCHARGE = "followup_calls"


def cohort(n_patients, t_start, t_end, rng):
    """Admissions for n_patients over [t_start, t_end) months.

    Each patient carries a hidden frailty that raises the risk of every one of
    their admissions. It is never a column, but it makes one patient's visits
    resemble each other — which is what a careless split will exploit.
    """
    rows = []
    for pid in range(n_patients):
        frailty = rng.normal()
        age = rng.normal(65, 12)
        chronic = max(0.0, rng.normal(2 + 0.8 * frailty, 1.2))
        n_visits = 1 + rng.poisson(cfg["visits_mean"])
        times = np.sort(rng.uniform(t_start, t_end, n_visits))
        for j, t in enumerate(times):
            los = rng.gamma(2, 2 + 0.5 * max(frailty, -1.5))
            n_meds = rng.poisson(5 + chronic)
            lab_a = rng.normal(0.4 * frailty, 1)
            lab_b = rng.normal(0, 1)
            emergency = float(rng.random() < 0.3 + 0.1 * (frailty > 0))
            logit = (
                -1.3
                + 1.5 * frailty
                + 0.25 * (chronic - 2)
                + 0.12 * (los - 5)
                + 0.4 * emergency
                + 0.3 * lab_a
                + 0.15 * j
                + 0.01 * (age - 65)
            )
            readmitted = int(rng.random() < 1 / (1 + np.exp(-logit)))
            rows.append(
                {
                    "patient": pid,
                    "month": t,
                    "age": age,
                    "chronic": chronic,
                    "los": los,
                    "n_meds": n_meds,
                    "lab_a": lab_a,
                    "lab_b": lab_b,
                    "emergency": emergency,
                    "prior": float(j),
                    "readmitted": readmitted,
                }
            )
    df = pd.DataFrame(rows)
    for col in ["lab_a", "lab_b"]:  # labs are not always drawn
        df.loc[rng.random(len(df)) < 0.2, col] = np.nan
    markers = pd.DataFrame(rng.normal(size=(len(df), len(MARKERS))), columns=MARKERS)
    return pd.concat([df, markers], axis=1)


rng = np.random.default_rng(cfg["seed"])
dev = cohort(cfg["dev_patients"], 0, 24, rng)
holdout = cohort(cfg["holdout_patients"], 24, 36, rng)
holdout["patient"] += 1_000_000  # new people, not new visits by old ones

# The warehouse also holds care-coordination calls logged in the 30 days AFTER
# discharge. Readmitted patients get many of them, because the readmission is
# what triggers the calls. At the moment the model must decide, none have
# happened yet — so in the holdout, built as the model will meet the world,
# the column is zero.
dev[POST_DISCHARGE] = rng.poisson(0.2 + 1.5 * dev["readmitted"]).astype(float)
holdout[POST_DISCHARGE] = 0.0

y_dev = dev["readmitted"].to_numpy()
y_holdout = holdout["readmitted"].to_numpy()
groups = dev["patient"].to_numpy()

# Captured values are rounded where they are made, so the page quotes the
# same digits the notebook prints.
dev_rate_pct = round(100 * y_dev.mean(), 1)
holdout_rate_pct = round(100 * y_holdout.mean(), 1)
dev_admissions = len(dev)
holdout_admissions = len(holdout)

if env.lang == "ar":
    # Counts go after a colon: Arabic number-noun agreement depends on the
    # number, and these numbers come from the profile.
    print(
        f"التطوير — الأشهر 0–24 | المرضى: {cfg['dev_patients']} | "
        f"حالات الدخول: {dev_admissions} | نسبة العودة: {dev_rate_pct}٪"
    )
    print(
        f"البيانات المحجوزة — الأشهر 24–36 | مرضى جدد: {cfg['holdout_patients']} | "
        f"حالات الدخول: {holdout_admissions} | نسبة العودة: {holdout_rate_pct}٪"
    )
    print(
        f"الأعمدة — سريرية: {len(CLINICAL)} | مؤشّرات مخبرية: {len(MARKERS)} | "
        f"بعد الخروج: {POST_DISCHARGE}"
    )
else:
    print(
        f"development: {dev_admissions} admissions from {cfg['dev_patients']} patients, "
        f"months 0-24, readmission rate {dev_rate_pct}%"
    )
    print(
        f"holdout:     {holdout_admissions} admissions from {cfg['holdout_patients']} new "
        f"patients, months 24-36, readmission rate {holdout_rate_pct}%"
    )
    print(
        f"columns: {len(CLINICAL)} clinical, {len(MARKERS)} lab markers, "
        f"and one post-discharge field ({POST_DISCHARGE})"
    )
Workshop code
development: 3562 admissions from 800 patients, months 0-24, readmission rate 38.2%
holdout:     6736 admissions from 1500 new patients, months 24-36, readmission rate 35.4%
columns: 8 clinical, 400 lab markers, and one post-discharge field (followup_calls)

The honest protocol follows two rules. First, every step that learns from data (imputing, scaling, choosing 20 features) sits inside a pipeline, so it is refitted within each training fold and never sees the fold it is scored on. Second, folds are split by patient. At deployment the model meets new patients, so the evaluation must too.

The same function produces every stage. Each leak is one if branch, marked LEAK in the code, so you can see exactly which line changes.

Look at the gap, not the AUC. An honest estimate is not a high one; it is one that the sealed data agrees with.

import warnings

from sklearn.ensemble import HistGradientBoostingClassifier
from sklearn.feature_selection import SelectKBest, f_classif
from sklearn.impute import SimpleImputer
from sklearn.metrics import roc_auc_score
from sklearn.model_selection import GroupKFold, KFold, cross_val_predict
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler

warnings.filterwarnings("ignore", message="The groups parameter is ignored")

STAGES = {
    0: "honest",
    1: "+ scale before split",
    2: "+ select before split",
    3: "+ split rows, not patients",
    4: "+ post-discharge column",
}
if env.lang == "ar":
    STAGES = {
        0: "نزيه",
        1: "+ توحيد قبل التقسيم",
        2: "+ انتقاء قبل التقسيم",
        3: "+ تقسيم الزيارات لا المرضى",
        4: "+ عمود ما بعد الخروج",
    }


def boosted():
    return HistGradientBoostingClassifier(
        max_iter=cfg["trees"], learning_rate=0.05, random_state=cfg["seed"]
    )


def evaluate(stage, labels=None, make_model=boosted):
    """Score one protocol by its own cross-validation, then on the holdout.

    Leaks are cumulative: stage 3 carries the leaks of stages 1 and 2. Each one
    is a single line below, and each line is one an analyst might write.
    """
    y = y_dev if labels is None else labels
    cols = CLINICAL + MARKERS
    if stage >= 4:
        cols = [*cols, POST_DISCHARGE]  # LEAK 4: a field that is filled in later
    X, X_out = dev[cols].to_numpy(), holdout[cols].to_numpy()

    steps = []
    if stage >= 1:  # LEAK 1: impute and scale using every development row
        imputer, scaler = SimpleImputer(), StandardScaler()
        X = scaler.fit_transform(imputer.fit_transform(X))
        X_out = scaler.transform(imputer.transform(X_out))
    else:
        steps += [SimpleImputer(), StandardScaler()]

    if stage >= 2:  # LEAK 2: choose features by their link to the label, on every row
        keep = SelectKBest(f_classif, k=cfg["k_select"]).fit(X, y).get_support()
        if stage >= 4:
            keep[-1] = True  # the analyst keeps the field; it ranks first anyway
        X, X_out = X[:, keep], X_out[:, keep]
    else:
        steps += [SelectKBest(f_classif, k=cfg["k_select"])]

    if stage >= 3:  # LEAK 3: shuffle admissions, so a patient lands on both sides
        splitter = KFold(cfg["folds"], shuffle=True, random_state=cfg["seed"])
    else:
        splitter = GroupKFold(cfg["folds"])

    model = make_pipeline(*steps, make_model())
    oof = cross_val_predict(model, X, y, cv=splitter, groups=groups, method="predict_proba")[:, 1]
    # The care team can call one patient in five: the threshold is set on the
    # cross-validated scores, exactly as it would be before launch.
    threshold = np.quantile(oof, 1 - cfg["alert_rate"])

    model.fit(X, y)
    p_out = model.predict_proba(X_out)[:, 1]
    return {
        "cv_auc": roc_auc_score(y, oof),
        "holdout_auc": roc_auc_score(y_holdout, p_out),
        "cv_recall": (oof[y == 1] >= threshold).mean(),
        "holdout_recall": (p_out[y_holdout == 1] >= threshold).mean(),
        "holdout_flagged": (p_out >= threshold).mean(),
        "p_out": p_out,
    }


honest = evaluate(0)
cv_auc_honest = round(honest["cv_auc"], 3)
holdout_auc_honest = round(honest["holdout_auc"], 3)
honest_gap = round(honest["cv_auc"] - honest["holdout_auc"], 3)

if env.lang == "ar":
    print(f"التقييم النزيه — AUC بالتحقّق المتقاطع: {cv_auc_honest:.3f}")
    print(f"                  AUC على البيانات المحجوزة: {holdout_auc_honest:.3f}")
    print(f"                  الفرق: {honest_gap:+.3f}")
else:
    print(f"honest protocol — cross-validated AUC: {cv_auc_honest:.3f}")
    print(f"                  holdout AUC:         {holdout_auc_honest:.3f}")
    print(f"                  gap:                 {honest_gap:+.3f}")
Workshop code
honest protocol — cross-validated AUC: 0.749
                  holdout AUC:         0.749
                  gap:                 -0.001

Cross-validation reported 0.749 and the sealed holdout returned 0.749, a gap of -0.001. That agreement is the whole point of the baseline. It shows that this data can be evaluated honestly, so every gap that opens from here on is caused by a leak, not by the data.

The leaks are cumulative: each stage keeps the ones before it.

  1. Scale before splitting. Imputation and scaling are fitted on all development rows, then cross-validation runs on the transformed data.
  2. Select before splitting. The 20 features most associated with the label are chosen using every row, including the rows each fold will later be scored on.
  3. Split rows, not patients. Admissions are shuffled into folds, so the same patient appears on both sides of a split.
  4. Keep the post-discharge column. It ranks first in feature selection, so an analyst keeps it.

Predict before you run it: which of the four will raise the cross-validated score the most, and which will hurt the holdout?

Two lines, one per question. The orange line is what each protocol claimed. The blue line is what its model did on the sealed year. The axis starts at 0.5 because that is a coin flip.

results = {0: honest}
for stage in range(1, 5):
    results[stage] = evaluate(stage)

header = (
    ("المرحلة", "AUC تحقّق", "AUC محجوزة", "الفرق")
    if env.lang == "ar"
    else ("stage", "CV AUC", "holdout AUC", "gap")
)
print(f"{header[0]:<30}{header[1]:>12}{header[2]:>14}{header[3]:>9}")
for stage, r in results.items():
    print(
        f"{stage} {STAGES[stage]:<28}{r['cv_auc']:>12.3f}{r['holdout_auc']:>14.3f}"
        f"{r['cv_auc'] - r['holdout_auc']:>+9.3f}"
    )

cv_auc_leaky = round(results[4]["cv_auc"], 3)
holdout_auc_leaky = round(results[4]["holdout_auc"], 3)
leaky_gap = round(results[4]["cv_auc"] - results[4]["holdout_auc"], 3)
cv_auc_stage3 = round(results[3]["cv_auc"], 3)
cv_gain_stage1 = round(results[1]["cv_auc"] - results[0]["cv_auc"], 3)
cv_gain_stage2 = round(results[2]["cv_auc"] - results[1]["cv_auc"], 3)
cv_gain_stage3 = round(results[3]["cv_auc"] - results[2]["cv_auc"], 3)
cv_gain_stage4 = round(results[4]["cv_auc"] - results[3]["cv_auc"], 3)
holdout_drop_leaky = round(results[0]["holdout_auc"] - results[4]["holdout_auc"], 3)

import matplotlib.pyplot as plt
from matplotlib import font_manager


def ar(text):
    """Shape and order Arabic for matplotlib, which draws isolated glyphs in
    logical order. A no-op in the English build."""
    if env.lang != "ar":
        return text
    import arabic_reshaper
    from bidi.algorithm import get_display

    # Line by line: the bidi pass would otherwise reorder whole lines.
    return "\n".join(get_display(arabic_reshaper.reshape(line)) for line in text.split("\n"))


if env.lang == "ar":
    from pathlib import Path

    # matplotlib's bundled DejaVu has no Arabic glyphs. Amiri comes from the
    # apt line in the dependencies cell; Ubuntu releases name the file
    # differently, so match case-insensitively.
    amiri = [
        p for p in Path("/usr/share/fonts").rglob("*.ttf") if p.name.lower() == "amiri-regular.ttf"
    ]
    if amiri:
        font_manager.fontManager.addfont(str(amiri[0]))
        plt.rcParams["font.family"] = font_manager.FontProperties(fname=str(amiri[0])).get_name()
    else:
        print("تنبيه: لم يُعثر على خطّ Amiri؛ ستظهر النصوص العربية في الرسوم مربّعات فارغة.")

stages = list(results)
fig, ax = plt.subplots(figsize=(7.5, 4.2))
ax.plot(
    stages,
    [results[s]["cv_auc"] for s in stages],
    "o-",
    color="#c2410c",
    lw=2,
    label=ar("ما قاله التحقّق المتقاطع") if env.lang == "ar" else "what cross-validation said",
)
ax.plot(
    stages,
    [results[s]["holdout_auc"] for s in stages],
    "o-",
    color="#1d4ed8",
    lw=2,
    label=ar("ما حدث على بيانات لم تُرَ") if env.lang == "ar" else "what happened on unseen data",
)
# AUC's floor is 0.5, not 0: a coin flip. Anchoring there keeps "flat" honest.
ax.set_ylim(0.5, 1.0)
ax.axhline(0.5, color="#999", lw=0.8)
ax.set_xticks(stages)
if env.lang == "ar":
    ax.set_xticklabels(
        [
            ar(s)
            for s in [
                "نزيه",
                "+ توحيد\nقبل التقسيم",
                "+ انتقاء\nقبل التقسيم",
                "+ تقسيم الزيارات\nلا المرضى",
                "+ عمود\nبعد الخروج",
            ]
        ]
    )
    ax.set_ylabel("AUC")
    ax.set_title(ar("كل تسريب يرفع الدرجة المُعلنة، ولا يرفع الأداء الحقيقي"))
else:
    ax.set_xticklabels(
        [
            "honest",
            "+ scale\nbefore split",
            "+ select\nbefore split",
            "+ split rows,\nnot patients",
            "+ post-discharge\ncolumn",
        ]
    )
    ax.set_ylabel("AUC")
    ax.set_title("Every leak raises the reported score. None raises the real one.")
ax.legend(loc="upper left", frameon=False)
fig.tight_layout()
plt.show()
Workshop code
stage                               CV AUC   holdout AUC      gap
0 honest                             0.749         0.749   -0.001
1 + scale before split               0.750         0.749   +0.001
2 + select before split              0.764         0.749   +0.015
3 + split rows, not patients         0.771         0.749   +0.021
4 + post-discharge column            0.900         0.731   +0.169

The orange line rises at every stage. The blue line stays exactly flat for three stages and then drops.

The leak that tutorials warn about first, scaling before the split, moved the score by 0.002. That is because tree models split on the order of values, and scaling does not change the order. With a distance-based model the same line can matter, so it is still worth fixing; it is simply not where the damage is here.

Selecting features on every row added 0.014. Among 400 markers of pure noise, some correlate with these particular labels by chance, and selection finds exactly those. Inside each fold they look predictive, because they were chosen with that fold's labels in view. Splitting rows instead of patients added 0.006: the model scores admissions of patients whose other admissions it has trained on.

The post-discharge column added 0.129 on its own, lifting the reported AUC to 0.900. On the sealed year, that same model scored 0.731, which is 0.019 below the honest model. The protocol with the best score shipped the worst model.

Stages 0 to 3 each produce a model to ship. Before running this, decide how different you expect their predictions to be.

# Stages 0 to 3 differ only in how the score was ESTIMATED. The model each one
# would ship is trained on all development rows either way. Compare what those
# four models predict for the holdout patients.
reference = results[0]["p_out"]
same_model_max_diff = float(max(np.abs(results[s]["p_out"] - reference).max() for s in (1, 2, 3)))
cv_inflation_stage3 = round(results[3]["cv_auc"] - results[0]["cv_auc"], 3)

if env.lang == "ar":
    print(f"أكبر فرق بين تنبّؤات المراحل 0–3 على المرضى الجدد: {same_model_max_diff:.2e}")
    print(f"ومع ذلك ارتفع AUC المُعلَن من {cv_auc_honest:.3f} إلى {cv_auc_stage3:.3f}")
else:
    print(f"largest difference in holdout predictions across stages 0-3: {same_model_max_diff:.2e}")
    print(f"while the reported AUC rose from {cv_auc_honest:.3f} to {cv_auc_stage3:.3f}")
Workshop code
largest difference in holdout predictions across stages 0-3: 0.00e+00
while the reported AUC rose from 0.749 to 0.771

The largest difference between the predictions of those four models is 0. They are the same model. Whichever protocol you follow, the model you ship is fitted on every development row, so scaling and selecting on every row is exactly what happens at the end anyway. The first three leaks changed only the estimate, which rose by 0.022 without the model changing at all.

That gives two kinds of leak. A leak in the evaluation leaves the model untouched and corrupts what you believe about it. The damage arrives later, through decisions made with that belief: which model to ship, which features to keep, what performance to promise. A leak in the data changes the model itself. The post-discharge column is the second kind, and the next cell shows what it does once the model is live.

A risk score becomes an action only through a threshold. The care team can call 20% of patients, so the threshold is set where the cross-validated scores put 20% of admissions above it. This is exactly how it would be set before launch. Then the same threshold is applied to the sealed year.

Recall is the share of real readmissions the alert catches. Compare each orange bar with the blue bar beside it.

recall_cv_leaky = round(results[4]["cv_recall"], 3)
recall_holdout_leaky = round(results[4]["holdout_recall"], 3)
recall_cv_honest = round(results[0]["cv_recall"], 3)
recall_holdout_honest = round(results[0]["holdout_recall"], 3)
recall_lost = round(results[0]["holdout_recall"] - results[4]["holdout_recall"], 3)
# Percentages for prose: "caught 0.2% of readmissions" reads; "0.002" does not.
recall_cv_leaky_pct = round(100 * results[4]["cv_recall"], 1)
recall_holdout_leaky_pct = round(100 * results[4]["holdout_recall"], 1)
flagged_holdout_leaky_pct = round(100 * results[4]["holdout_flagged"], 1)
recall_cv_honest_pct = round(100 * results[0]["cv_recall"], 1)
recall_holdout_honest_pct = round(100 * results[0]["holdout_recall"], 1)
alert_rate_pct = round(100 * cfg["alert_rate"])

if env.lang == "ar":
    print(f"{'':<30}{'الالتقاط (تحقّق)':>16}{'الالتقاط بعد الإطلاق':>22}{'نسبة المُنبَّه عليهم':>22}")
else:
    print(f"{'':<30}{'recall (CV)':>14}{'recall (holdout)':>18}{'share alerted':>16}")
for stage in (0, 4):
    r = results[stage]
    print(
        f"{stage} {STAGES[stage]:<28}{r['cv_recall']:>14.1%}{r['holdout_recall']:>18.1%}"
        f"{r['holdout_flagged']:>16.1%}"
    )

fig, ax = plt.subplots(figsize=(6.5, 3.6))
x = np.arange(2)
cv_bars = [recall_cv_honest, recall_cv_leaky]
out_bars = [recall_holdout_honest, recall_holdout_leaky]
ax.bar(
    x - 0.18,
    cv_bars,
    0.36,
    color="#c2410c",
    label=ar("في التحقّق المتقاطع") if env.lang == "ar" else "in cross-validation",
)
ax.bar(
    x + 0.18,
    out_bars,
    0.36,
    color="#1d4ed8",
    label=ar("بعد الإطلاق") if env.lang == "ar" else "after launch",
)
ax.set_ylim(0, 1)  # recall starts at zero, and zero is part of the story
ax.set_xticks(x)
if env.lang == "ar":
    ax.set_xticklabels([ar("نزيه"), ar("مع العمود المُسرَّب")])
    ax.set_ylabel(ar("نسبة حالات العودة التي التُقطت"))
    ax.set_title(ar("العتبة نفسها، بعد الإطلاق"))
else:
    ax.set_xticklabels(["honest", "with the leaked column"])
    ax.set_ylabel("share of readmissions caught")
    ax.set_title("The same threshold, after launch")
ax.legend(frameon=False)
fig.tight_layout()
plt.show()
Workshop code
                                 recall (CV)  recall (holdout)   share alerted
0 honest                               36.1%             32.3%           16.8%
4 + post-discharge column              49.2%              0.2%            0.1%

In cross-validation, the leaky model caught 49.200% of readmissions, more than the honest model's 36.100%. After launch it caught 0.200%, and alerted on 0.100% of patients. The honest model still caught 32.300%.

This is the part AUC did not show. AUC only asks whether readmitted patients are ranked above the others. The leaky model still ranks them nearly as well, which is why its holdout AUC dropped by only 0.019. What changed is the level of every score. The model learned that calls after discharge mean risk, and at discharge every patient has zero calls. So every score moved down, and the threshold, set on scores where the column was filled in, now sits above almost all of them.

On the dashboard this looks like a small drop in AUC. On the ward, the alert has gone silent.

A different model family, the same two protocols. Look at whether the gaps follow the model or the protocol.

# Is this a property of gradient boosting? Swap in a random forest and repeat
# the two ends of the ladder.
from sklearn.ensemble import RandomForestClassifier


def forest():
    return RandomForestClassifier(n_estimators=cfg["trees"], n_jobs=-1, random_state=cfg["seed"])


rf_honest = evaluate(0, make_model=forest)
rf_leaky = evaluate(4, make_model=forest)
rf_gap_honest = round(rf_honest["cv_auc"] - rf_honest["holdout_auc"], 3)
rf_gap_leaky = round(rf_leaky["cv_auc"] - rf_leaky["holdout_auc"], 3)
rf_cv_auc_leaky = round(rf_leaky["cv_auc"], 3)
rf_recall_holdout_leaky_pct = round(100 * rf_leaky["holdout_recall"], 1)

if env.lang == "ar":
    print(f"{'غابة عشوائية':<30}{'AUC تحقّق':>12}{'AUC محجوزة':>14}{'الالتقاط بعد الإطلاق':>22}")
else:
    print(f"{'random forest':<30}{'CV AUC':>12}{'holdout AUC':>14}{'recall after launch':>22}")
for stage, r in ((0, rf_honest), (4, rf_leaky)):
    print(
        f"{stage} {STAGES[stage]:<28}{r['cv_auc']:>12.3f}{r['holdout_auc']:>14.3f}"
        f"{r['holdout_recall']:>22.1%}"
    )
Workshop code
random forest                       CV AUC   holdout AUC   recall after launch
0 honest                             0.765         0.761                 34.9%
4 + post-discharge column            0.908         0.748                  0.0%

The random forest repeats the pattern. Its honest gap is 0.003 and its leaky gap is 0.160, with a cross-validated AUC of 0.908 and 0% recall after launch. Changing the model does not remove the leak, because the leak is in the evaluation and the data, not in the learner.

You will not usually have a sealed future year. One cheap test needs no extra data: shuffle the labels and rerun the whole protocol. With the labels shuffled, nothing can be predicted, so an honest protocol must score about 0.5. A score reliably above that was produced by the procedure, not learned from the patients. The cell averages 3 shuffles for each stage.

The first column should sit near 0.5. Compare stage 2 with stage 4: their reported AUCs are far apart.

# A cheap test anyone can run: destroy the labels, rerun the WHOLE protocol.
# With no signal left, an honest protocol must score about 0.5. Anything
# reliably above it was learned from the procedure, not from the patients.
audit = {}
for stage in STAGES:
    scores = []
    for k in range(cfg["shuffles"]):
        shuffled = np.random.default_rng(cfg["seed"] + 100 + k).permutation(y_dev)
        scores.append(evaluate(stage, labels=shuffled)["cv_auc"])
    audit[stage] = float(np.mean(scores))

head = (
    ("المرحلة", "AUC على تسميات مخلوطة", "AUC المُعلَن")
    if env.lang == "ar"
    else ("stage", "AUC on shuffled labels", "reported AUC")
)
print(f"{head[0]:<30}{head[1]:>24}{head[2]:>22}")
for stage in STAGES:
    print(f"{stage} {STAGES[stage]:<28}{audit[stage]:>24.3f}{results[stage]['cv_auc']:>22.3f}")

audit_auc_honest = round(audit[0], 3)
audit_auc_stage2 = round(audit[2], 3)
audit_auc_leaky = round(audit[4], 3)
Workshop code
stage                           AUC on shuffled labels          reported AUC
0 honest                                         0.496                 0.749
1 + scale before split                           0.496                 0.750
2 + select before split                          0.531                 0.764
3 + split rows, not patients                     0.526                 0.771
4 + post-discharge column                        0.527                 0.900

The honest protocol scored 0.496 on shuffled labels, as it should. Selecting before the split pushed that to 0.531: the procedure found signal in noise, and the audit caught it.

Stage 4 scored 0.527, about the same as stage 2. The audit did not see the post-discharge column, even though that column was worth 0.129 of reported AUC. The reason is that the column copies the real outcomes. Once the labels are shuffled, the column no longer matches them, so it has nothing to leak. Row-level shuffling also hides the patient overlap from stage 3, because shuffled labels no longer belong to patients. An audit for that leak has to keep each patient's labels together while it shuffles.

So the shuffle audit tests the procedure. To test the columns, you need a different screen.

Exercise

This screen scores every column on its own against the label and flags any column above SUSPICIOUS_AUC. Run it once as it is. Then lower the threshold until a second column is flagged. Is that column a leak? What would you need to know about each flagged column to decide, and where would you find it?

# YOUR TURN.
#
# The shuffle audit tests the procedure. This screen tests the columns: score
# each one ALONE against the label. A single field that nearly matches the
# whole model is a question to ask whoever owns that field, not a gift.
SUSPICIOUS_AUC = 0.80

solo = {}
for col in CLINICAL + MARKERS + [POST_DISCHARGE]:
    values = dev[col].fillna(dev[col].median()).to_numpy()
    a = roc_auc_score(y_dev, values)
    solo[col] = max(a, 1 - a)  # direction does not matter for a screen

ranked = sorted(solo.items(), key=lambda kv: kv[1], reverse=True)
for col, score in ranked[:8]:
    flag = (
        "  <-- " + ("اسأل عنه" if env.lang == "ar" else "ask about it")
        if score >= SUSPICIOUS_AUC
        else ""
    )
    print(f"{col:<18}{score:.3f}{flag}")

solo_auc_post_discharge = round(solo[POST_DISCHARGE], 3)
solo_auc_best_clinical = round(max(solo[c] for c in CLINICAL), 3)
solo_auc_best_marker = round(max(solo[c] for c in MARKERS), 3)
Workshop code

A hint is available in the notebook — env.hint(1)

On its own, the post-discharge column scored 0.846, higher than the entire honest model (0.749). A single field that beats the whole model is the clearest warning a dataset can give. The best clinical column scored 0.738, and the best of the 400 noise markers scored 0.532.

The number alone cannot settle it. A strong legitimate feature and a leak can score alike. What settles it is a fact outside the data: when each field is written, compared with when the decision is made. Every column needs that timestamp, and finding it usually means asking the people who own the field.

The control comes first: an honest model that never learned would make every comparison after it meaningless.

# The control first: if the honest model never learned to rank patients, every
# comparison below is a comparison against nothing.
honest_learns_ok = env.check("honest-model-learns", holdout_auc_honest)
honest_gap_ok = env.check("honest-cv-tracks-holdout", honest_gap)
recall_lost_ok = env.check("leak-costs-recall", recall_lost)
Workshop code
✓ Honest model: AUC on admissions from new patients: 0.749 (needs ≥ 0.68)
✓ Honest protocol: cross-validated AUC minus holdout AUC: -0.001 (needs ≥ -0.05 and ≤ 0.05)
✓ Readmissions caught after launch: honest model minus leaky model: 0.321 (needs ≥ 0.2)

Every leak in this workshop raised the reported score, and none of them improved the model. That is the shape leakage takes in practice. It rarely looks like a bug; it looks like the best result so far.

The evaluation is a system in its own right, and it needs the same care as the model:

  • Split by whatever is new at deployment. Here that meant patients and time; elsewhere it may be customers, sites, or devices.
  • Put every step that learns from data inside the pipeline, so it is refitted in each fold.
  • Give every column a timestamp and compare it with the moment of decision.
  • Seal a holdout built the way the world will deliver data, and open it once.
  • Shuffle the labels to test the procedure, and screen columns one at a time to test the data.
  • Check the decision, not only the ranking. AUC barely moved here, while the alert stopped working.
receipt = env.receipt()
Workshop code
Workshop complete.

Completion code: AZ-██████████
Paste it on the workshop's page on Azimuth to record it.
Last verified: 2026-10-02 · Tesla T4 · PyTorch 2.11.0+cu130 · Python 3.13.15 · 8894810

Terms in this workshop