{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/assessing-the-un-trustworthiness-of-saliency","title":"Assessing the (Un)Trustworthiness of Saliency Maps for Localizing Abnormalities in Medical Imaging","arxiv_id":"2008.02766","date":"2020-08-06","proceeding":null,"authors":["Nishanth Arun","Nathan Gaw","Praveer Singh","Ken Chang","Mehak Aggarwal","Bryan Chen","Katharina Hoebel","Sharut Gupta","Jay Patel","Mishka Gidwani","Julius Adebayo","Matthew D. Li","Jayashree Kalpathy-Cramer"],"abstract":"Saliency maps have become a widely used method to make deep learning models more interpretable by providing post-hoc explanations of classifiers through identification of the most pertinent areas of the input medical image. They are increasingly being used in medical imaging to provide clinically plausible explanations for the decisions the neural network makes. However, the utility and robustness of these visualization maps has not yet been rigorously examined in the context of medical imaging. We posit that trustworthiness in this context requires 1) localization utility, 2) sensitivity to model weight randomization, 3) repeatability, and 4) reproducibility. Using the localization information available in two large public radiology datasets, we quantify the performance of eight commonly used saliency map approaches for the above criteria using area under the precision-recall curves (AUPRC) and structural similarity index (SSIM), comparing their performance to various baseline measures. Using our framework to quantify the trustworthiness of saliency maps, we show that all eight saliency map techniques fail at least one of the criteria and are, in most cases, less trustworthy when compared to the baselines. We suggest that their usage in the high-risk domain of medical imaging warrants additional scrutiny and recommend that detection or segmentation models be used if localization is the desired output of the network. Additionally, to promote reproducibility of our findings, we provide the code we used for all tests performed in this work at this link: https://github.com/QTIM-Lab/Assessing-Saliency-Maps.","url_abs":"https://arxiv.org/abs/2008.02766v2","url_pdf":"https://arxiv.org/pdf/2008.02766v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"assessing-the-un-trustworthiness-of-saliency","repo_url":"https://github.com/QTIM-Lab/Assessing-Saliency-Maps","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"ssim","task_name":"SSIM"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2008.02766","atlas_url":"https://app.syntology.ai/?focus=2008.02766","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2008.02766"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/QTIM-Lab/Assessing-Saliency-Maps","reach":null}],"summary":{"ran_honours":2},"by_repo_kind":{"official":{"samples":2,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":2,"samples":[{"code_sha256_prefix":"aef8f64128072ffc","entry":"soft_dice","repo":"QTIM-Lab/Assessing-Saliency-Maps","repo_kind":"official","path":"scripts/create_maps_rsna.py","file_url":"https://github.com/QTIM-Lab/Assessing-Saliency-Maps/blob/HEAD/scripts/create_maps_rsna.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"aef8f64128072ffc"}},{"code_sha256_prefix":"eda4d44490dfc546","entry":"soft_dice","repo":"QTIM-Lab/Assessing-Saliency-Maps","repo_kind":"official","path":"scripts/create_maps_siim.py","file_url":"https://github.com/QTIM-Lab/Assessing-Saliency-Maps/blob/HEAD/scripts/create_maps_siim.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"eda4d44490dfc546"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}