{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/bagel-a-benchmark-for-assessing-graph-neural","title":"BAGEL: A Benchmark for Assessing Graph Neural Network Explanations","arxiv_id":"2206.13983","date":"2022-06-28","proceeding":null,"authors":["Mandeep Rathee","Thorben Funke","Avishek Anand","Megha Khosla"],"abstract":"The problem of interpreting the decisions of machine learning is a well-researched and important. We are interested in a specific type of machine learning model that deals with graph data called graph neural networks. Evaluating interpretability approaches for graph neural networks (GNN) specifically are known to be challenging due to the lack of a commonly accepted benchmark. Given a GNN model, several interpretability approaches exist to explain GNN models with diverse (sometimes conflicting) evaluation methodologies. In this paper, we propose a benchmark for evaluating the explainability approaches for GNNs called Bagel. In Bagel, we firstly propose four diverse GNN explanation evaluation regimes -- 1) faithfulness, 2) sparsity, 3) correctness. and 4) plausibility. We reconcile multiple evaluation metrics in the existing literature and cover diverse notions for a holistic evaluation. Our graph datasets range from citation networks, document graphs, to graphs from molecules and proteins. We conduct an extensive empirical study on four GNN models and nine post-hoc explanation approaches for node and graph classification tasks. We open both the benchmarks and reference implementations and make them available at https://github.com/Mandeep-Rathee/Bagel-benchmark.","url_abs":"https://arxiv.org/abs/2206.13983v1","url_pdf":"https://arxiv.org/pdf/2206.13983v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"bagel-a-benchmark-for-assessing-graph-neural","repo_url":"https://github.com/mandeep-rathee/bagel-benchmark","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"machine-learning","task_name":"BIG-bench Machine Learning"},{"task_slug":"graph-classification","task_name":"Graph Classification"},{"task_slug":"graph-neural-network","task_name":"Graph Neural Network"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2206.13983","atlas_url":"https://app.syntology.ai/?focus=2206.13983","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2206.13983"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/mandeep-rathee/bagel-benchmark","reach":null}],"summary":{"ran_violates":2,"ran_honours":1,"unverified":1},"by_repo_kind":{"official":{"samples":4,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"45e4a164344e6932","entry":"attr_mask","repo":"mandeep-rathee/bagel-benchmark","repo_kind":"official","path":"bagel_benchmark/explainers/grad_explainer_node.py","file_url":"https://github.com/mandeep-rathee/bagel-benchmark/blob/HEAD/bagel_benchmark/explainers/grad_explainer_node.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"45e4a164344e6932"}},{"code_sha256_prefix":"d9b06a75f9da9100","entry":"prediction_val","repo":"mandeep-rathee/bagel-benchmark","repo_kind":"official","path":"bagel_benchmark/explainers/grad_explainer_node.py","file_url":"https://github.com/mandeep-rathee/bagel-benchmark/blob/HEAD/bagel_benchmark/explainers/grad_explainer_node.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"d9b06a75f9da9100"}},{"code_sha256_prefix":"451c91cf42e54e33","entry":"test","repo":"mandeep-rathee/bagel-benchmark","repo_kind":"official","path":"bagel_benchmark/graph_classification/utils_movie_reviews.py","file_url":"https://github.com/mandeep-rathee/bagel-benchmark/blob/HEAD/bagel_benchmark/graph_classification/utils_movie_reviews.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"451c91cf42e54e33"}},{"code_sha256_prefix":"5719af7e6ac163ca","entry":"normalize","repo":"mandeep-rathee/bagel-benchmark","repo_kind":"official","path":"bagel_benchmark/explainers/grad_explainer_node.py","file_url":"https://github.com/mandeep-rathee/bagel-benchmark/blob/HEAD/bagel_benchmark/explainers/grad_explainer_node.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"5719af7e6ac163ca"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}