{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/pitfalls-of-graph-neural-network-evaluation","title":"Pitfalls of Graph Neural Network Evaluation","arxiv_id":"1811.05868","date":"2018-11-14","proceeding":null,"authors":["Oleksandr Shchur","Maximilian Mumme","Aleksandar Bojchevski","Stephan Günnemann"],"abstract":"Semi-supervised node classification in graphs is a fundamental problem in graph mining, and the recently proposed graph neural networks (GNNs) have achieved unparalleled results on this task. Due to their massive success, GNNs have attracted a lot of attention, and many novel architectures have been put forward. In this paper we show that existing evaluation strategies for GNN models have serious shortcomings. We show that using the same train/validation/test splits of the same datasets, as well as making significant changes to the training procedure (e.g. early stopping criteria) precludes a fair comparison of different architectures. We perform a thorough empirical evaluation of four prominent GNN models and show that considering different splits of the data leads to dramatically different rankings of models. Even more importantly, our findings suggest that simpler GNN architectures are able to outperform the more sophisticated ones if the hyperparameters and the training procedure are tuned fairly for all models.","url_abs":"https://arxiv.org/abs/1811.05868v2","url_pdf":"https://arxiv.org/pdf/1811.05868v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"pitfalls-of-graph-neural-network-evaluation","repo_url":"https://github.com/toinesayan/node-classification-and-label-dependencies","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"pitfalls-of-graph-neural-network-evaluation","repo_url":"https://github.com/shchur/gnn-benchmark","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"graph-mining","task_name":"Graph Mining"},{"task_slug":"graph-neural-network","task_name":"Graph Neural Network"},{"task_slug":"node-classification","task_name":"Node Classification"}],"methods":[{"method_slug":"early-stopping","method_name":"Early Stopping"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1811.05868","atlas_url":"https://app.syntology.ai/?focus=1811.05868","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1811.05868"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/shchur/gnn-benchmark","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/toinesayan/node-classification-and-label-dependencies","reach":null}],"summary":{"ran_draft_wrong":2,"unverified":1},"by_repo_kind":{"listed":{"samples":3,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"46126de00c7ef161","entry":"add_to_full_config","repo":"shchur/gnn-benchmark","repo_kind":"listed","path":"scripts/create_jobs.py","file_url":"https://github.com/shchur/gnn-benchmark/blob/HEAD/scripts/create_jobs.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"46126de00c7ef161"}},{"code_sha256_prefix":"163ee624c9962868","entry":"generate_multiple_splits","repo":"shchur/gnn-benchmark","repo_kind":"listed","path":"scripts/create_jobs.py","file_url":"https://github.com/shchur/gnn-benchmark/blob/HEAD/scripts/create_jobs.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"163ee624c9962868"}},{"code_sha256_prefix":"c06dcd9a49f4851b","entry":"load_search_config","repo":"shchur/gnn-benchmark","repo_kind":"listed","path":"scripts/create_jobs.py","file_url":"https://github.com/shchur/gnn-benchmark/blob/HEAD/scripts/create_jobs.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"c06dcd9a49f4851b"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}