Methods › General › Interpretability
Interpretability
Interpretability Methods seek to explain the predictions made by neural networks by introducing mechanisms to enduce or enforce interpretability. For example, LIME approximates the neural network with a locally interpretable model. Below you can find a continuously updating list of interpretability methods.
Methods
All 17 methods in this collection, most-tagged first. Year is the archive's introduced_year; the archive stores 2000 when it has none, shown here as “–”. Papers counts distinct papers the archive tags with the method. Click a heading to sort.
| SHAP Shapley Additive Explanations | – | 550 |
| LIME Local Interpretable Model-Agnostic Explanations | – | 378 |
| CAM Class-activation map | – | 267 |
| Monte Carlo Dropout | – | 202 |
| Visual Analytics | – | 78 |
| NAM Neural Additive Model | – | 18 |
| Network Dissection | – | 11 |
| CDEP Contextual Decomposition Explanation Penalization | – | 2 |
| DecomCAM Decomposition-Integration Class Activation Map | – | 2 |
| Universal Probing Massively multilingual probing based on Universal Dependencies | – | 2 |
| Agglomerative Contextual Decomposition | – | 1 |
| Disentangled Attribution Curves | – | 1 |
| Hierarchical Network Dissection | – | 1 |
| Symbolic Deep Learning | – | 1 |
| Syntax Heat Parse Tree | – | 1 |
| TE2Rules Tree Ensemble to Rules | – | 1 |
| WinTSR Windowed Temporal Saliency Rescaling | – | 1 |