{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/towards-ground-truth-evaluation-of-visual","title":"Ground Truth Evaluation of Neural Network Explanations with CLEVR-XAI","arxiv_id":"2003.07258","date":"2020-03-16","proceeding":null,"authors":["Leila Arras","Ahmed Osman","Wojciech Samek"],"abstract":"The rise of deep learning in today's applications entailed an increasing need in explaining the model's decisions beyond prediction performances in order to foster trust and accountability. Recently, the field of explainable AI (XAI) has developed methods that provide such explanations for already trained neural networks. In computer vision tasks such explanations, termed heatmaps, visualize the contributions of individual pixels to the prediction. So far XAI methods along with their heatmaps were mainly validated qualitatively via human-based assessment, or evaluated through auxiliary proxy tasks such as pixel perturbation, weak object localization or randomization tests. Due to the lack of an objective and commonly accepted quality measure for heatmaps, it was debatable which XAI method performs best and whether explanations can be trusted at all. In the present work, we tackle the problem by proposing a ground truth based evaluation framework for XAI methods based on the CLEVR visual question answering task. Our framework provides a (1) selective, (2) controlled and (3) realistic testbed for the evaluation of neural network explanations. We compare ten different explanation methods, resulting in new insights about the quality and properties of XAI methods, sometimes contradicting with conclusions from previous comparative studies. The CLEVR-XAI dataset and the benchmarking code can be found at https://github.com/ahmedmagdiosman/clevr-xai.","url_abs":"https://arxiv.org/abs/2003.07258v2","url_pdf":"https://arxiv.org/pdf/2003.07258v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"towards-ground-truth-evaluation-of-visual","repo_url":"https://github.com/ahmedmagdiosman/clevr-xai","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"towards-ground-truth-evaluation-of-visual","repo_url":"https://github.com/ahmedmagdiosman/simply-clevr-dataset","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"benchmarking","task_name":"Benchmarking"},{"task_slug":"xai","task_name":"Explainable Artificial Intelligence (XAI)"},{"task_slug":"feature-importance","task_name":"Feature Importance"},{"task_slug":"object-localization","task_name":"Object Localization"},{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"visual-question-answering-1","task_name":"Visual Question Answering"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"}],"methods":[],"datasets_introduced":[{"slug":"simply-clevr","name":"simply-CLEVR","full_name":null}],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2003.07258","atlas_url":"https://app.syntology.ai/?focus=2003.07258","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2003.07258"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ahmedmagdiosman/simply-clevr-dataset","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ahmedmagdiosman/clevr-xai","reach":null}],"summary":{"ran_fixture":1,"ran_draft_wrong":2},"by_repo_kind":{"official":{"samples":3,"ran":3,"repositories":2}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"4e2e73983db94402","entry":"find_filter_options","repo":"ahmedmagdiosman/clevr-xai","repo_kind":"official","path":"question_generation/generate_questions.py","file_url":"https://github.com/ahmedmagdiosman/clevr-xai/blob/HEAD/question_generation/generate_questions.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"4e2e73983db94402"}},{"code_sha256_prefix":"4fd7f6e374f964cd","entry":"find_relate_filter_options","repo":"ahmedmagdiosman/clevr-xai","repo_kind":"official","path":"question_generation/generate_questions.py","file_url":"https://github.com/ahmedmagdiosman/clevr-xai/blob/HEAD/question_generation/generate_questions.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"4fd7f6e374f964cd"}},{"code_sha256_prefix":"399457377d04215f","entry":"node_shallow_copy","repo":"ahmedmagdiosman/simply-clevr-dataset","repo_kind":"official","path":"question_generation/generate_questions.py","file_url":"https://github.com/ahmedmagdiosman/simply-clevr-dataset/blob/HEAD/question_generation/generate_questions.py","link_basis":"plan_row","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"399457377d04215f"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}