Core ML2016intermediate10 min read

"Why Should I Trust You?": Explaining the Predictions of Any Classifier

«لماذا أثق بك؟»: تفسير تنبؤات أي مُصنِّف

Ribeiro, M. T. · Singh, S. · Guestrin, C. — KDD

The problem

Machine learning models — deep neural networks, random forests, gradient boosting — achieve high but operate as black boxes. A practitioner cannot tell whether a model learned genuine patterns or just memorized artifacts in the training data. Without explanations, deploying a model in medicine, law, or finance is a leap of faith. Existing interpretable models like decision trees sacrifice too much accuracy, while complex models sacrifice all transparency.

The contribution

LIME (Local Interpretable Model-agnostic Explanations): a technique that explains any classifier's prediction by fitting a simple, interpretable model — typically a linear model — in the local neighborhood of the prediction. It perturbs the input, queries the black-box model on those perturbations, weights them by proximity, and learns which features matter most for that specific prediction. SP-LIME extends this to select a representative set of explanations that give a global understanding of the model.

The impact

LIME established the paradigm of local post-hoc explainability. It became the default tool for practitioners needing to explain individual predictions, and directly inspired SHAP, Anchors, and Grad-CAM. The paper's framing of trust — distinguishing between trusting a single prediction versus trusting the entire model — reshaped how the ML community thinks about deploying models responsibly.

Imagine a judge who always gives the correct verdict but never explains the reasoning. You'd have no way to know whether the verdict came from solid evidence or a lucky guess.

LIME acts like a courtroom stenographer for algorithms: after the black-box model makes a prediction, LIME replays the case with small variations — removing one piece of evidence at a time — and records how the verdict changes. The features whose removal flips the verdict the most are the ones the model was really relying on.

Now you can read the reasoning, even though the judge never wrote it down.

The problem: black boxes demand blind trust

A deep neural network classifying X-ray images might achieve 95% accuracy on the . But accuracy alone doesn't tell you why it's right. Maybe it learned that sick patients' X-rays tend to have a particular hospital's watermark — a spurious correlation. Deploy that model in another hospital and it collapses.

Ribeiro et al. identify two distinct levels of trust:

  • Trusting a prediction: should I act on this specific output? A doctor needs to know which regions of the X-ray drove the diagnosis before ordering treatment.

  • Trusting a model: should I deploy the model itself? A team needs a global picture of what the model has learned, not just one explanation.

Without tools that address both levels, practitioners are stuck choosing between accuracy (black boxes) and transparency (simple models that often underperform).

Open in Lab
Algorithm 1 uses genuine features; Algorithm 2 uses a data leak. Both score well — but only explanations reveal the difference.
The demo wakes as you arrive…

What makes a good explanation?

The authors argue that a useful explanation must satisfy three properties simultaneously:

  • Interpretable: it uses a representation humans understand — words present or absent, image patches visible or hidden — not raw pixel values or dimensions.

  • Locally faithful: the explanation accurately reflects the model's behavior in the neighborhood of the prediction. Global fidelity is impossible for complex models, but local fidelity is achievable and sufficient.

  • Model-agnostic: the same explanation technique must work for any classifier — neural networks, random forests, SVMs — without accessing internal weights or gradients. Only the input-output behavior is used.

Open in Lab
Click anywhere on the nonlinear boundary. A local linear model (dashed) fits the curve well nearby but diverges far away.
The demo wakes as you arrive…

The LIME algorithm: perturb, observe, explain

LIME follows a four-step pipeline to explain any single prediction:

Step 1 — Choose an interpretable representation. For text, this is a binary vector indicating which words are present or absent. For images, it's a binary vector over superpixels (contiguous patches). The model internally may use complex embeddings, but the explanation lives in this simpler space.

Step 2 — Generate perturbations. LIME creates many variations of the input by randomly toggling features on and off. For text, it randomly removes words; for images, it randomly grays out superpixels.

Step 3 — Query the black box. Each perturbed input is fed into the original model to get its prediction. This is purely input-output — no internal access needed.

Step 4 — Fit a local interpretable model. A weighted linear model is trained on the perturbations. Perturbations closer to the original input (measured by an exponential kernel) get higher weight. The coefficients of this linear model become the explanation: positive means the pushes toward the predicted class, negative means it pushes away.

Open in Lab
Step through LIME's four stages on a text example. Watch perturbations get created, scored, weighted, and fitted.
The demo wakes as you arrive…

The formal objective: balancing fidelity and simplicity

Before seeing the formula, understand what LIME is optimizing for. It wants to find the simplest possible model gg that still accurately mimics the black box ff in the local neighborhood of the point xx. There are two competing goals:

  • Fidelity: gg should agree with ff on nearby points. Measured by a L\mathcal{L} that compares their outputs.

  • Simplicity: gg should use as few features as possible so a human can read it. Measured by a complexity penalty Ω(g)\Omega(g).

The explanation is the model gg that achieves the best trade-off between these two goals.

ξ(x)=arg min⁡g∈G  L(f,g,πx)+Ω(g)\xi(x) = \underset{g \in G}{\operatorname{arg\,min}} \; \mathcal{L}(f, g, \pi_x) + \Omega(g)
LIME objective — the core optimization — G = family of interpretable models (e.g. linear models) · L = fidelity loss — how well g mimics f near x · π_x = proximity weighting — closer perturbations matter more · Ω(g) = complexity penalty — fewer features = simpler explanation

Think of it as a magnifying glass: LIME zooms into a tiny region around xx (the proximity kernel πx\pi_x controls the zoom level), fits a simple straight line inside that region (minimizing L\mathcal{L}), and keeps that line as short as possible (minimizing Ω\Omega). The line's slope tells you which features matter.

The proximity kernel: how close is close?

Not all perturbations are equally informative. A slightly modified input tells you about the model's behavior right here; a wildly different input tells you about somewhere else entirely. LIME uses an exponential kernel to weight perturbations by their distance from the original input. The kernel width σ\sigma acts like a zoom dial: smaller σ\sigma means a tighter, more local explanation; larger σ\sigma includes more distant perturbations.

πx(z)=exp⁡ ⁣(−D(x,z)2σ2)\pi_x(z) = \exp\!\left(-\frac{D(x, z)^2}{\sigma^2}\right)
Exponential proximity kernel — D(x, z) = distance between original input x and perturbation z (cosine for text, Euclidean for tabular) · σ = kernel width (controls locality) · Result: perturbations near x get weight ≈ 1, distant ones get weight ≈ 0
Open in Lab
Adjust σ and watch how the weighting changes. Small σ: only very close perturbations matter. Large σ: the explanation considers a wider neighborhood.
The demo wakes as you arrive…

LIME for text, images, and tabular data

The power of LIME lies in its model-agnostic design: the same core algorithm adapts to any data type by changing only the interpretable representation and strategy:

Text: the interpretable representation is a binary vector of word presence/absence. Perturbation means randomly removing words from the document. The explanation highlights which words push the prediction toward or away from a class.

Images: the image is first segmented into superpixels (contiguous patches of similar color). The interpretable representation is a binary vector over these superpixels. Perturbation means graying out random superpixels. The explanation highlights which regions of the image matter most.

Tabular data: features are perturbed by sampling from the training data distribution. For continuous features, values are drawn from a normal distribution fit to the training data. The explanation shows which features and their directions drove the prediction.

Open in Lab
Switch between Text, Image, and Tabular modes to see how LIME adapts its perturbation and explanation strategy.
The demo wakes as you arrive…

SP-LIME: from local explanations to global understanding

A single LIME explanation tells you about one prediction. But to trust or debug an entire model, you need a representative sample of explanations. Picking random instances might give you redundant explanations that all highlight the same features.

SP-LIME (Submodular Pick LIME) solves this with a greedy submodular optimization. It selects a small budget of BB instances whose explanations together cover the widest variety of important features. The key insight is submodularity: adding an explanation that covers new features is valuable, but adding one that repeats already-covered features has diminishing returns.

The result is a compact, non-redundant summary of the model's behavior — like choosing the best photos for a portfolio rather than dumping the entire camera roll.

Open in Lab
Watch SP-LIME greedily pick instances that maximize feature coverage. Each new pick adds features not yet covered.
The demo wakes as you arrive…

LIME in code

Simplified LIME for text classificationpython

Simplified to show the idea — not the real implementation.

import numpy as np
from sklearn.linear_model import Ridge

def lime_explain_text(text, predict_fn, n_samples=500, n_features=6):
    """Explain a text prediction using LIME.
       text: original document (list of words)
       predict_fn: black-box model, takes list of strings → probabilities
    """
    original_pred = predict_fn([' '.join(text)])[0]

    # Step 1: Generate binary perturbations (1 = word present, 0 = removed)
    n_words = len(text)
    samples = np.random.binomial(1, 0.5, size=(n_samples, n_words))
    samples[0] = np.ones(n_words)  # always include the original

    # Step 2: Build perturbed texts by removing words where sample == 0
    perturbed_texts = []
    for s in samples:
        words = [text[i] for i in range(n_words) if s[i] == 1]
        perturbed_texts.append(' '.join(words) if words else ' ')

    # Step 3: Get black-box predictions for all perturbations
    predictions = predict_fn(perturbed_texts)

    # Step 4: Compute proximity weights (exponential kernel on cosine distance)
    distances = np.sqrt(np.sum((samples - samples[0]) ** 2, axis=1))
    sigma = 0.75 * np.sqrt(n_words)
    weights = np.exp(-(distances ** 2) / (sigma ** 2))

    # Step 5: Fit a weighted linear model
    model = Ridge(alpha=1.0)
    model.fit(samples, predictions, sample_weight=weights)

    # Return top features with their importance
    importances = list(zip(text, model.coef_))
    importances.sort(key=lambda x: abs(x[1]), reverse=True)
    return importances[:n_features]

Strengths and limitations

LIME's elegance comes from its simplicity, but that simplicity also introduces trade-offs worth understanding:

  • Strength: truly model-agnostic. LIME only needs to call the model — no gradients, no architecture knowledge. It works on neural networks, random forests, SVMs, and even proprietary APIs where you cannot see the internals.

  • Strength: human-readable. The output is a ranked list of features with positive/negative contributions — something a non-expert can evaluate.

  • Limitation: instability. Because perturbations are sampled randomly, running LIME twice on the same input can give different explanations. The randomness is inherent to the sampling approach.

  • Limitation: kernel width sensitivity. The choice of σ\sigma significantly affects results. Too small and the explanation uses too few perturbations; too large and it loses locality.

  • Limitation: linearity assumption. The local model is linear, but even locally, some decision boundaries may have curvature that a line cannot capture.

Impact: LIME reshaped explainable AI

  1. 2016

    LIME published

    Ribeiro, Singh, and Guestrin introduce LIME at KDD. First practical framework for explaining any classifier's predictions through local perturbation and surrogate models.

  2. 2017

    SHAP unifies explanations

    Lundberg and Lee show that LIME is a special case of Shapley additive explanations. SHAP provides deterministic, theoretically grounded feature attributions. LIME's speed remains an advantage for large-scale applications.

  3. 2017

    Grad-CAM for visual explanations

    Selvaraju et al. use gradient information to produce visual explanations for CNNs. Unlike LIME, Grad-CAM is model-specific (needs gradients) but faster for image models.

  4. 2018

    Anchors — high-precision rules

    Ribeiro et al. extend LIME with Anchors: if-then rules that guarantee a prediction with high probability, addressing LIME's linearity limitation with more expressive rules.

  5. 2019

    LIME in production and regulation

    LIME becomes a standard tool in regulated industries. GDPR's "right to explanation" drives adoption in finance and healthcare, where model decisions must be justifiable.

  6. 2020

    Explainability becomes mainstream

    Major cloud platforms (AWS SageMaker, Google Cloud AI, Azure ML) integrate LIME and SHAP into their MLOps pipelines, making explainability a default rather than an afterthought.

LIME's legacy extends beyond its algorithm. It framed explainability as a practical engineering problem, not an abstract philosophical one. The distinction between trusting a prediction and trusting a model became foundational vocabulary. And by showing that local perturbation-based explanations work remarkably well, LIME opened the door for an entire field: SHAP, Grad-CAM, Anchors, Integrated Gradients, and dozens of methods that followed.

CitationRibeiro, Singh, Guestrin. "Why Should I Trust You?": Explaining the Predictions of Any Classifier. KDD, 2016.

Terms in this paper