Papers › Order in the Court: Explainable AI Methods Prone to Disagreement

Order in the Court: Explainable AI Methods Prone to Disagreement

7 May 2021arXiv:2105.03287archive 2025-07-28

Michael Neely, Stefan F. Schouten, Maurits J. R. Bleeker, Ana Lucic

By computing the rank correlation between attention weights and feature-additive explanation methods, previous analyses either invalidate or support the role of attention-based explanations as a faithful and plausible measure of salience. To investigate whether this approach is appropriate, we compare LIME, Integrated Gradients, DeepLIFT, Grad-SHAP, Deep-SHAP, and attention-based explanations, applied to two neural architectures trained on single- and pair-sequence language tasks. In most cases, we find that none of our chosen methods agree. Based on our empirical observations and theoretical objections, we conclude that rank correlation does not measure the quality of feature-additive methods. Practitioners should instead use the numerous and rigorous diagnostic methods proposed by the community.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

sfschouten/court-of-xai officialmentioned in papermentioned on GitHubpytorchMIT report
grobruegge/vitexplcomp mentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Diagnostic

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

LIME

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections