{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/evaluating-the-visualization-of-what-a-deep","title":"Evaluating the visualization of what a Deep Neural Network has learned","arxiv_id":"1509.06321","date":"2015-09-21","proceeding":null,"authors":["Wojciech Samek","Alexander Binder","Grégoire Montavon","Sebastian Bach","Klaus-Robert Müller"],"abstract":"Deep Neural Networks (DNNs) have demonstrated impressive performance in\ncomplex machine learning tasks such as image classification or speech\nrecognition. However, due to their multi-layer nonlinear structure, they are\nnot transparent, i.e., it is hard to grasp what makes them arrive at a\nparticular classification or recognition decision given a new unseen data\nsample. Recently, several approaches have been proposed enabling one to\nunderstand and interpret the reasoning embodied in a DNN for a single test\nimage. These methods quantify the ''importance'' of individual pixels wrt the\nclassification decision and allow a visualization in terms of a heatmap in\npixel/input space. While the usefulness of heatmaps can be judged subjectively\nby a human, an objective quality measure is missing. In this paper we present a\ngeneral methodology based on region perturbation for evaluating ordered\ncollections of pixels such as heatmaps. We compare heatmaps computed by three\ndifferent methods on the SUN397, ILSVRC2012 and MIT Places data sets. Our main\nresult is that the recently proposed Layer-wise Relevance Propagation (LRP)\nalgorithm qualitatively and quantitatively provides a better explanation of\nwhat made a DNN arrive at a particular classification decision than the\nsensitivity-based approach or the deconvolution method. We provide theoretical\narguments to explain this result and discuss its practical implications.\nFinally, we investigate the use of heatmaps for unsupervised assessment of\nneural network performance.","url_abs":"http://arxiv.org/abs/1509.06321v1","url_pdf":"http://arxiv.org/pdf/1509.06321v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"evaluating-the-visualization-of-what-a-deep","repo_url":"https://github.com/icrto/xML","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"classification-1","task_name":"Classification"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"image-classification","task_name":"image-classification"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[{"method_slug":"heatmap","method_name":"Heatmap"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1509.06321","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}