{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/how-can-we-fool-lime-and-shap-adversarial","title":"Fooling LIME and SHAP: Adversarial Attacks on Post hoc Explanation Methods","arxiv_id":"1911.02508","date":"2019-11-06","proceeding":null,"authors":["Dylan Slack","Sophie Hilgard","Emily Jia","Sameer Singh","Himabindu Lakkaraju"],"abstract":"As machine learning black boxes are increasingly being deployed in domains such as healthcare and criminal justice, there is growing emphasis on building tools and techniques for explaining these black boxes in an interpretable manner. Such explanations are being leveraged by domain experts to diagnose systematic errors and underlying biases of black boxes. In this paper, we demonstrate that post hoc explanations techniques that rely on input perturbations, such as LIME and SHAP, are not reliable. Specifically, we propose a novel scaffolding technique that effectively hides the biases of any given classifier by allowing an adversarial entity to craft an arbitrary desired explanation. Our approach can be used to scaffold any biased classifier in such a way that its predictions on the input data distribution still remain biased, but the post hoc explanations of the scaffolded classifier look innocuous. Using extensive evaluation with multiple real-world datasets (including COMPAS), we demonstrate how extremely biased (racist) classifiers crafted by our framework can easily fool popular explanation techniques such as LIME and SHAP into generating innocuous explanations which do not reflect the underlying biases.","url_abs":"https://arxiv.org/abs/1911.02508v2","url_pdf":"https://arxiv.org/pdf/1911.02508v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"how-can-we-fool-lime-and-shap-adversarial","repo_url":"https://github.com/dylan-slack/Fooling-LIME-SHAP","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"how-can-we-fool-lime-and-shap-adversarial","repo_url":"https://github.com/MachineLearningJournalClub/LearningNLP","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[],"methods":[{"method_slug":"lime","method_name":"LIME"},{"method_slug":"shap","method_name":"SHAP"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1911.02508","atlas_url":"https://app.syntology.ai/?focus=1911.02508","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1911.02508"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/MachineLearningJournalClub/LearningNLP","reach":{"status":"ok","spdx":"NOASSERTION"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/dylan-slack/Fooling-LIME-SHAP","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"unverified":6},"by_repo_kind":{"official":{"samples":6,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"a46063cbe60e6db0","entry":"get_and_preprocess_cc","repo":"dylan-slack/Fooling-LIME-SHAP","repo_kind":"official","path":"get_data.py","file_url":"https://github.com/dylan-slack/Fooling-LIME-SHAP/blob/HEAD/get_data.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"a46063cbe60e6db0"}},{"code_sha256_prefix":"a84d2da9f9de4008","entry":"get_and_preprocess_compas_data","repo":"dylan-slack/Fooling-LIME-SHAP","repo_kind":"official","path":"get_data.py","file_url":"https://github.com/dylan-slack/Fooling-LIME-SHAP/blob/HEAD/get_data.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"a84d2da9f9de4008"}},{"code_sha256_prefix":"b0c2d8e2e77deadf","entry":"get_and_preprocess_german","repo":"dylan-slack/Fooling-LIME-SHAP","repo_kind":"official","path":"get_data.py","file_url":"https://github.com/dylan-slack/Fooling-LIME-SHAP/blob/HEAD/get_data.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"b0c2d8e2e77deadf"}},{"code_sha256_prefix":"e480a9fcf9f84a0f","entry":"get_rank_map","repo":"dylan-slack/Fooling-LIME-SHAP","repo_kind":"official","path":"utils.py","file_url":"https://github.com/dylan-slack/Fooling-LIME-SHAP/blob/HEAD/utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"e480a9fcf9f84a0f"}},{"code_sha256_prefix":"abfefac5568cb4e2","entry":"one_hot_encode","repo":"dylan-slack/Fooling-LIME-SHAP","repo_kind":"official","path":"utils.py","file_url":"https://github.com/dylan-slack/Fooling-LIME-SHAP/blob/HEAD/utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"abfefac5568cb4e2"}},{"code_sha256_prefix":"c58a875000ccc10e","entry":"rank_features","repo":"dylan-slack/Fooling-LIME-SHAP","repo_kind":"official","path":"utils.py","file_url":"https://github.com/dylan-slack/Fooling-LIME-SHAP/blob/HEAD/utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"c58a875000ccc10e"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}