Core ML2017intermediate9 min read
A Unified Approach to Interpreting Model Predictions
إطار موحَّد لتفسير تنبؤات النماذج
Lundberg, S. · Lee, S. — NeurIPS
The problem
Complex models — deep networks, -boosted trees, ensembles — achieve high but are opaque: users cannot tell why a particular prediction was made. Many interpretation methods existed (LIME, DeepLIFT, layer-wise relevance propagation, Shapley sampling) but they were developed independently, with no unifying theory. It was unclear when one method was better than another, and some methods violated basic fairness properties without users knowing it.
The contribution
SHAP (SHapley Additive exPlanations) unifies six existing interpretation methods under one class — methods — and proves that from theory are the unique solution satisfying three desirable properties: , missingness, and consistency. The paper also introduces (a model-agnostic estimator connecting LIME and Shapley values) and DeepSHAP (a fast approximation for deep networks).
The impact
SHAP became the de facto standard for model . The open-source shap library is used across healthcare, finance, autonomous systems, and regulatory compliance. It established that explanation methods should be evaluated by axiomatic properties, not just visual appeal, and bridged the gap between game theory and practical machine learning.
A company's quarterly profit just came in at record highs. The board wants to know: how much credit does each department deserve? You can't just look at final numbers — the marketing team's campaign only worked because engineering shipped the product on time.
Shapley values solve this exact problem. They consider every possible combination of departments, measure each one's in every context, and produce a provably fair credit assignment.
SHAP does this for AI: for any single prediction, it tells you exactly how much each input feature pushed the answer up or down, with mathematical guarantees that the credits add up and are fair.
The problem: powerful models, opaque decisions
Modern machine learning has a trust problem. A deep can predict whether a patient has cancer with 95% accuracy, but if the doctor asks why — which features in the scan mattered? — the model stays silent. A bank's gradient-boosted tree can deny a loan instantly, but regulators demand to know which factors drove the denial.
By 2017, several interpretation methods existed — LIME, DeepLIFT, layer-wise relevance propagation, Shapley sampling values — but they were developed in isolation. Each method told a different story about the same prediction, and users had no way to judge which story was more trustworthy. Worse, some methods violated basic fairness properties: a feature that contributed nothing might still receive credit, or two equally important features might get different attributions.
The key insight: fair credit from game theory
In 1953, mathematician Lloyd Shapley asked: if a group of players cooperate to win a game, how should the prize be divided fairly? His answer — now called Shapley values — considers every possible ordering in which players could have joined the game, measures each player's marginal contribution in each ordering, and averages over all orderings.
SHAP transplants this idea into machine learning. The "game" is a single prediction. The "players" are input features. The "prize" is the difference between the model's prediction for this specific input and the average prediction across all inputs. Each feature's Shapley value tells you: on average, across all possible subsets of other features, how much did including this feature change the prediction?
The formula: Shapley values for predictions
Before seeing the math, build a mental picture: imagine a courtroom where the prediction is the verdict. Each feature is a witness. To judge one witness's contribution, you consider every possible subset of other witnesses who might have testified before this one, measure how the verdict changes when this witness adds their testimony, and average over all such orderings. That average is the Shapley value.
The computational challenge is clear: with features, there are possible subsets. For 20 features that's over a million coalitions. For 1000 features — typical in a neural network — exact computation is impossible. That is why SHAP introduces smart approximations.
Three properties that force one answer
The key theoretical result of the paper is a uniqueness theorem. SHAP defines three properties that any good explanation should satisfy, then proves that only Shapley values satisfy all three simultaneously:
-
Local accuracy — the feature attributions must sum to the actual prediction for this specific input minus the average prediction. No credit is left unassigned; none is fabricated.
-
Missingness — if a feature is absent from the input (its simplified value is 0), it must receive zero attribution. You don't credit a player who never entered the game.
-
Consistency — if a model changes so that a feature's marginal contribution increases or stays the same in every possible context, that feature's attribution must not decrease. Rewarding a feature less for contributing more would be absurd.
The unifying class: additive feature attribution
SHAP's first contribution is recognizing that six seemingly different interpretation methods — LIME, DeepLIFT, layer-wise relevance propagation, Shapley regression values, Shapley sampling values, and quantitative input influence — all belong to the same class. Each one explains a prediction using a linear model of binary variables:
where indicates which features are "present" and is the attribution assigned to feature . This is the explanation model — a simple surrogate that locally approximates the original model's behavior. The methods differ only in how they choose the values.
KernelSHAP: the bridge between LIME and Shapley values
LIME explains a prediction by fitting a weighted linear model on perturbed inputs. But LIME's choice of weights and are heuristic — they lack theoretical justification. SHAP's key insight is that a specific choice of loss function, weighting , and recovers exact Shapley values.
Think of it as calibrating a scale: LIME uses rough weights and gets approximate readings. KernelSHAP uses mathematically precise weights — the Shapley kernel — and gets the provably fair answer. The Shapley kernel assigns each a weight:
This kernel gives extreme weight to coalitions where almost all features are present or almost none are. These are the most informative comparisons — they isolate one feature's contribution most cleanly.
DeepSHAP: fast explanations for deep networks
KernelSHAP works for any model but requires many model evaluations. For deep networks with thousands of parameters, this is slow. DeepSHAP adapts DeepLIFT's trick: instead of sampling coalitions, it pushes attribution backward through the network, layer by layer, using a modified chain rule.
Imagine water flowing backward through a pipe network: at each junction, the flow splits proportionally to how much each incoming pipe contributed to the output flow. DeepSHAP does the same with prediction credit, but uses Shapley-calibrated splitting rules so the final attributions approximately satisfy the three desirable properties.
Under the assumption that features are independent and the network is approximately linear, DeepSHAP recovers the exact Shapley values. In practice, the approximation is remarkably good even when these assumptions are imperfect.
The same idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
from itertools import combinations
from math import factorial
def shapley_values(model_fn, x, baseline, M):
"""Compute exact Shapley values for a single prediction.
model_fn: takes a masked input, returns a scalar prediction
x: the input to explain (1D array of M features)
baseline: reference input (e.g. mean of training data)
M: number of features
"""
phi = np.zeros(M)
for i in range(M):
for size in range(0, M):
for S in combinations([j for j in range(M) if j != i], size):
S = set(S)
# Build input with only coalition S active
z_without = baseline.copy()
for j in S:
z_without[j] = x[j]
# Add feature i to the coalition
z_with = z_without.copy()
z_with[i] = x[i]
# Marginal contribution of feature i
marginal = model_fn(z_with) - model_fn(z_without)
# Shapley weight: |S|! * (M-|S|-1)! / M!
weight = factorial(len(S)) * factorial(M - len(S) - 1) / factorial(M)
phi[i] += weight * marginal
return phi
# Usage: phi = shapley_values(model.predict, x_test[0], x_train.mean(0), M=10)
# phi[i] = how much feature i pushed this prediction above/below the averageReading SHAP explanations: force and summary plots
SHAP produces two key visualizations. A force plot explains a single prediction: features pushing the prediction higher appear in red on one side, features pushing it lower appear in blue on the other, and together they stretch or compress the bar from the base value to the final prediction.
A summary plot (beeswarm) explains the model globally: each dot is one prediction, features are ranked by importance, and the color shows whether the feature value was high or low. Patterns emerge: "high income always pushes the loan prediction up" or "high age has mixed effects depending on other features."
Why it mattered
2016
LIME — local surrogates
Ribeiro et al. propose LIME: fit a simple model around each prediction to explain it. Powerful idea, but the choice of kernel and perturbation method is ad-hoc.
2017
SHAP — the unification
Lundberg & Lee show that LIME, DeepLIFT, and four other methods are all additive feature attribution methods, and that Shapley values are the unique solution satisfying three desirable properties.
2020
TreeSHAP — exact and fast for trees
Lundberg et al. exploit tree structure to compute exact Shapley values in polynomial time. XGBoost and LightGBM integrate it natively.
2021
Regulatory adoption
EU AI Act drafts cite interpretability as a requirement. SHAP becomes the most widely used tool for model explanation in compliance workflows.
2023
SHAP everywhere
The shap library exceeds 20k GitHub stars. SHAP values are standard in healthcare ML, credit scoring, autonomous driving, and scientific discovery.
SHAP's legacy is not just a better explanation method — it established that interpretability must be held to mathematical standards, just like prediction accuracy. The question is no longer "does this explanation look reasonable?" but "does it satisfy provable fairness properties?" That shift — from aesthetics to axiomatics — is what made model interpretability trustworthy.
CitationLundberg, Lee. A Unified Approach to Interpreting Model Predictions. NeurIPS, 2017.
Terms in this paper
- Interpretabilityالقابلية للتفسير
- Shapley Valuesقيم شابلي
- Feature Importanceأهمية السمات
- Additive Feature Attributionالإسناد الجمعي للسمات
- Kernelالنواة الحسابية
- Loss functionدالة الخسارة
- Gradientالتدرج التفاضلي
- Deep Learningالتعلم العميق
- Decision Treeشجرة القرار الإحصائية
- Random Forestالغابة العشوائية خوارزمية
- Cooperative Gameاللعبة التعاونية