Core ML2017intermediate9 min read

A Unified Approach to Interpreting Model Predictions

إطار موحَّد لتفسير تنبؤات النماذج

Lundberg, S. · Lee, S. — NeurIPS

The problem

Complex models — deep networks, -boosted trees, ensembles — achieve high but are opaque: users cannot tell why a particular prediction was made. Many interpretation methods existed (LIME, DeepLIFT, layer-wise relevance propagation, Shapley sampling) but they were developed independently, with no unifying theory. It was unclear when one method was better than another, and some methods violated basic fairness properties without users knowing it.

The contribution

SHAP (SHapley Additive exPlanations) unifies six existing interpretation methods under one class — methods — and proves that from theory are the unique solution satisfying three desirable properties: , missingness, and consistency. The paper also introduces (a model-agnostic estimator connecting LIME and Shapley values) and DeepSHAP (a fast approximation for deep networks).

The impact

SHAP became the de facto standard for model . The open-source shap library is used across healthcare, finance, autonomous systems, and regulatory compliance. It established that explanation methods should be evaluated by axiomatic properties, not just visual appeal, and bridged the gap between game theory and practical machine learning.

A company's quarterly profit just came in at record highs. The board wants to know: how much credit does each department deserve? You can't just look at final numbers — the marketing team's campaign only worked because engineering shipped the product on time.

Shapley values solve this exact problem. They consider every possible combination of departments, measure each one's in every context, and produce a provably fair credit assignment.

SHAP does this for AI: for any single prediction, it tells you exactly how much each input feature pushed the answer up or down, with mathematical guarantees that the credits add up and are fair.

The problem: powerful models, opaque decisions

Modern machine learning has a trust problem. A deep can predict whether a patient has cancer with 95% accuracy, but if the doctor asks why — which features in the scan mattered? — the model stays silent. A bank's gradient-boosted tree can deny a loan instantly, but regulators demand to know which factors drove the denial.

By 2017, several interpretation methods existed — LIME, DeepLIFT, layer-wise relevance propagation, Shapley sampling values — but they were developed in isolation. Each method told a different story about the same prediction, and users had no way to judge which story was more trustworthy. Worse, some methods violated basic fairness properties: a feature that contributed nothing might still receive credit, or two equally important features might get different attributions.

Open in Lab
Before SHAP, interpretation methods were scattered islands. Click each method to see its assumptions. SHAP shows they are all special cases of one framework.
The demo wakes as you arrive…

The key insight: fair credit from game theory

In 1953, mathematician Lloyd Shapley asked: if a group of players cooperate to win a game, how should the prize be divided fairly? His answer — now called Shapley values — considers every possible ordering in which players could have joined the game, measures each player's marginal contribution in each ordering, and averages over all orderings.

SHAP transplants this idea into machine learning. The "game" is a single prediction. The "players" are input features. The "prize" is the difference between the model's prediction for this specific input and the average prediction across all inputs. Each feature's Shapley value tells you: on average, across all possible subsets of other features, how much did including this feature change the prediction?

Open in Lab
Choose a feature and step through every coalition. Watch how its marginal contribution changes depending on which other features are already present.
The demo wakes as you arrive…

The formula: Shapley values for predictions

Before seeing the math, build a mental picture: imagine a courtroom where the prediction is the verdict. Each feature is a witness. To judge one witness's contribution, you consider every possible subset of other witnesses who might have testified before this one, measure how the verdict changes when this witness adds their testimony, and average over all such orderings. That average is the Shapley value.

ϕi=∑S⊆F∖{i}∣S∣!  (∣F∣−∣S∣−1)!∣F∣![f(S∪{i})−f(S)]\phi_i = \sum_{S \subseteq F \setminus \{i\}} \frac{|S|!\;(|F|-|S|-1)!}{|F|!} \bigl[f(S \cup \{i\}) - f(S)\bigr]
Shapley value for feature i — S = a subset of features not including i · F = all features · f(S) = model prediction using only features in S · The weighting ensures every ordering of players is equally likely · The bracketed term is the marginal contribution of feature i to coalition S.

The computational challenge is clear: with MM features, there are 2M2^M possible subsets. For 20 features that's over a million coalitions. For 1000 features — typical in a neural network — exact computation is impossible. That is why SHAP introduces smart approximations.

Three properties that force one answer

The key theoretical result of the paper is a uniqueness theorem. SHAP defines three properties that any good explanation should satisfy, then proves that only Shapley values satisfy all three simultaneously:

  • Local accuracy — the feature attributions must sum to the actual prediction for this specific input minus the average prediction. No credit is left unassigned; none is fabricated.

  • Missingness — if a feature is absent from the input (its simplified value is 0), it must receive zero attribution. You don't credit a player who never entered the game.

  • Consistency — if a model changes so that a feature's marginal contribution increases or stays the same in every possible context, that feature's attribution must not decrease. Rewarding a feature less for contributing more would be absurd.

Open in Lab
Toggle each property on/off and see which explanation methods survive. Only SHAP satisfies all three.
The demo wakes as you arrive…

The unifying class: additive feature attribution

SHAP's first contribution is recognizing that six seemingly different interpretation methods — LIME, DeepLIFT, layer-wise relevance propagation, Shapley regression values, Shapley sampling values, and quantitative input influence — all belong to the same class. Each one explains a prediction using a linear model of binary variables:

g(z′)=ϕ0+∑i=1Mϕizi′g(z') = \phi_0 + \sum_{i=1}^{M} \phi_i z'_i

where z′∈{0,1}Mz' \in \{0,1\}^M indicates which features are "present" and ϕi\phi_i is the attribution assigned to feature ii. This is the explanation model — a simple surrogate that locally approximates the original model's behavior. The methods differ only in how they choose the ϕi\phi_i values.

Open in Lab
See how feature attributions stack to form the prediction. Drag features to toggle them on/off and watch the explanation model update.
The demo wakes as you arrive…

KernelSHAP: the bridge between LIME and Shapley values

LIME explains a prediction by fitting a weighted linear model on perturbed inputs. But LIME's choice of weights and are heuristic — they lack theoretical justification. SHAP's key insight is that a specific choice of loss function, weighting , and recovers exact Shapley values.

Think of it as calibrating a scale: LIME uses rough weights and gets approximate readings. KernelSHAP uses mathematically precise weights — the Shapley kernel — and gets the provably fair answer. The Shapley kernel assigns each SS a weight:

π(∣S∣)=(M−1)(M∣S∣)⋅∣S∣⋅(M−∣S∣)\pi(|S|) = \frac{(M - 1)}{\binom{M}{|S|} \cdot |S| \cdot (M - |S|)}

This kernel gives extreme weight to coalitions where almost all features are present or almost none are. These are the most informative comparisons — they isolate one feature's contribution most cleanly.

πx′(z′)=(M−1)(M∣z′∣) ∣z′∣ (M−∣z′∣)\pi_{x'}(z') = \frac{(M-1)}{\binom{M}{|z'|}\,|z'|\,(M-|z'|)}
The Shapley kernel — the key to KernelSHAP — This kernel gives infinite weight to the empty and full coalitions (all features absent or all present), which forces the linear model to match the model's output exactly at the endpoints. Intermediate coalitions get weights that ensure Shapley-fair averaging.
Open in Lab
Drag the coalition size slider and see how the Shapley kernel assigns weights. Extreme sizes get the most weight.
The demo wakes as you arrive…

DeepSHAP: fast explanations for deep networks

KernelSHAP works for any model but requires many model evaluations. For deep networks with thousands of parameters, this is slow. DeepSHAP adapts DeepLIFT's trick: instead of sampling coalitions, it pushes attribution backward through the network, layer by layer, using a modified chain rule.

Imagine water flowing backward through a pipe network: at each junction, the flow splits proportionally to how much each incoming pipe contributed to the output flow. DeepSHAP does the same with prediction credit, but uses Shapley-calibrated splitting rules so the final attributions approximately satisfy the three desirable properties.

Under the assumption that features are independent and the network is approximately linear, DeepSHAP recovers the exact Shapley values. In practice, the approximation is remarkably good even when these assumptions are imperfect.

The same idea in code

A minimal Shapley value calculatorpython

Simplified to show the idea — not the real implementation.

import numpy as np
from itertools import combinations
from math import factorial

def shapley_values(model_fn, x, baseline, M):
    """Compute exact Shapley values for a single prediction.
    model_fn: takes a masked input, returns a scalar prediction
    x:        the input to explain (1D array of M features)
    baseline: reference input (e.g. mean of training data)
    M:        number of features
    """
    phi = np.zeros(M)
    for i in range(M):
        for size in range(0, M):
            for S in combinations([j for j in range(M) if j != i], size):
                S = set(S)
                # Build input with only coalition S active
                z_without = baseline.copy()
                for j in S:
                    z_without[j] = x[j]
                # Add feature i to the coalition
                z_with = z_without.copy()
                z_with[i] = x[i]
                # Marginal contribution of feature i
                marginal = model_fn(z_with) - model_fn(z_without)
                # Shapley weight: |S|! * (M-|S|-1)! / M!
                weight = factorial(len(S)) * factorial(M - len(S) - 1) / factorial(M)
                phi[i] += weight * marginal
    return phi

# Usage: phi = shapley_values(model.predict, x_test[0], x_train.mean(0), M=10)
# phi[i] = how much feature i pushed this prediction above/below the average

Reading SHAP explanations: force and summary plots

SHAP produces two key visualizations. A force plot explains a single prediction: features pushing the prediction higher appear in red on one side, features pushing it lower appear in blue on the other, and together they stretch or compress the bar from the base value to the final prediction.

A summary plot (beeswarm) explains the model globally: each dot is one prediction, features are ranked by importance, and the color shows whether the feature value was high or low. Patterns emerge: "high income always pushes the loan prediction up" or "high age has mixed effects depending on other features."

Open in Lab
Click different test samples to see how features push each prediction. Red features push the prediction up; blue features push it down.
The demo wakes as you arrive…

Why it mattered

  1. 2016

    LIME — local surrogates

    Ribeiro et al. propose LIME: fit a simple model around each prediction to explain it. Powerful idea, but the choice of kernel and perturbation method is ad-hoc.

  2. 2017

    SHAP — the unification

    Lundberg & Lee show that LIME, DeepLIFT, and four other methods are all additive feature attribution methods, and that Shapley values are the unique solution satisfying three desirable properties.

  3. 2020

    TreeSHAP — exact and fast for trees

    Lundberg et al. exploit tree structure to compute exact Shapley values in polynomial time. XGBoost and LightGBM integrate it natively.

  4. 2021

    Regulatory adoption

    EU AI Act drafts cite interpretability as a requirement. SHAP becomes the most widely used tool for model explanation in compliance workflows.

  5. 2023

    SHAP everywhere

    The shap library exceeds 20k GitHub stars. SHAP values are standard in healthcare ML, credit scoring, autonomous driving, and scientific discovery.

SHAP's legacy is not just a better explanation method — it established that interpretability must be held to mathematical standards, just like prediction accuracy. The question is no longer "does this explanation look reasonable?" but "does it satisfy provable fairness properties?" That shift — from aesthetics to axiomatics — is what made model interpretability trustworthy.

CitationLundberg, Lee. A Unified Approach to Interpreting Model Predictions. NeurIPS, 2017.

Terms in this paper