When Model Attribution Has Only One Solution, Ask Which Axiom It Violates
A risk-control model rejects a loan, and customer service hands over a bar chart: this bar is the longest, so it's the culprit.
A risk-control model rejects a loan, and customer service hands over a bar chart: this bar is the longest, so it's the culprit. The 2017 paper asked a harder question — on what grounds does that chart count as an explanation of the model?
The authors put it plainly: within the class of additive feature attribution methods, there is exactly one solution that satisfies local accuracy, missingness, and consistency simultaneously — the Shapley value. Any approach that doesn't follow Shapley must violate at least one of these axioms.
Don't rush to treat the bar chart as a verdict. The original user study involved only 30 and 52 participants, and the experiments focused on simple models and settings like MNIST; exact values are computationally infeasible, and the approximations actually used rely on assumptions that features are mutually independent or that the model is approximately linear.
Its reach is also limited: explanation methods outside the framework are out of scope, and for tasks beyond text and images, as well as unverified model types, no transferable conclusions are offered.
“A Unified Approach to Interpreting Model Predictions” (2017) | Next review 2027-09-20
Sources:https://arxiv.org/abs/1705.07874