Methods › General › Interpretability

Interpretability

17 methods 1,371 papers tagged archive 2025-07-28

Interpretability Methods seek to explain the predictions made by neural networks by introducing mechanisms to enduce or enforce interpretability. For example, LIME approximates the neural network with a locally interpretable model. Below you can find a continuously updating list of interpretability methods.

Methods

All 17 methods in this collection, most-tagged first. Year is the archive's introduced_year; the archive stores 2000 when it has none, shown here as “–”. Papers counts distinct papers the archive tags with the method. Click a heading to sort.

SHAP Shapley Additive Explanations – 550
LIME Local Interpretable Model-Agnostic Explanations – 378
CAM Class-activation map – 267
Monte Carlo Dropout – 202
Visual Analytics – 78
NAM Neural Additive Model – 18
Network Dissection – 11
CDEP Contextual Decomposition Explanation Penalization – 2
DecomCAM Decomposition-Integration Class Activation Map – 2
Universal Probing Massively multilingual probing based on Universal Dependencies – 2
Agglomerative Contextual Decomposition – 1
Disentangled Attribution Curves – 1
Hierarchical Network Dissection – 1
Symbolic Deep Learning – 1
Syntax Heat Parse Tree – 1
TE2Rules Tree Ensemble to Rules – 1
WinTSR Windowed Temporal Saliency Rescaling – 1