Papers › A Fair Comparison of Graph Neural Networks for Graph Classification

A Fair Comparison of Graph Neural Networks for Graph Classification

20 Dec 2019ICLR 2020 1arXiv:1912.09893archive 2025-07-28

Federico Errica, Marco Podda, Davide Bacciu, Alessio Micheli

Experimental reproducibility and replicability are critical topics in machine learning. Authors have often raised concerns about their lack in scientific publications to improve the quality of the field. Recently, the graph representation learning field has attracted the attention of a wide research community, which resulted in a large stream of works. As such, several Graph Neural Network models have been developed to effectively tackle graph classification. However, experimental procedures often lack rigorousness and are hardly reproducible. Motivated by this, we provide an overview of common practices that should be avoided to fairly compare with the state of the art. To counter this troubling trend, we ran more than 47000 experiments in a controlled and uniform framework to re-evaluate five popular models across nine common benchmarks. Moreover, by comparing GNNs with structure-agnostic baselines we provide convincing evidence that, on some datasets, structural information has not been exploited yet. We believe that this work can contribute to the development of the graph learning field, by providing a much needed grounding for rigorous evaluations of graph classification models.

PaperPDFConference PDFCode

In Syntology View this paper on Syntology: its repositories, every harvested function with whether it ran, its licence and the call to fetch it.

Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

diningphil/gnn-comparison officialmentioned in paperpytorchGPL-3.0 report
diningphil/CGMM mentioned on GitHubpytorchBSD-3-Clause report
diningphil/icgmm mentioned on GitHubpytorchBSD-3-Clause report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

General ClassificationGraph ClassificationGraph LearningGraph Neural NetworkGraph Representation LearningRepresentation Learning

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Graph Classification COLLAB GraphSAGE Accuracy 73.9% #26 of 39 Archive leaderboard report
Graph Classification D&D DGCNN Accuracy 76.6% #36 of 53 Archive leaderboard report
Graph Classification ENZYMES GIN Accuracy 59.6% #29 of 54 Archive leaderboard report
Graph Classification ENZYMES GraphSAGE Accuracy 58.2% #34 of 54 Archive leaderboard report
Graph Classification IMDb-B GraphSAGE Accuracy 68.8% #47 of 51 Archive leaderboard report
Graph Classification IMDb-M GraphSAGE Accuracy 47.6% #32 of 36 Archive leaderboard report
Graph Classification NCI1 GIN Accuracy 80% #38 of 69 Archive leaderboard report
Graph Classification NCI1 DGCNN Accuracy 76.4% #47 of 69 Archive leaderboard report
Graph Classification PROTEINS DiffPool Accuracy 73.7% #85 of 103 Archive leaderboard report
Graph Classification PROTEINS GraphSAGE Accuracy 73% #90 of 103 Archive leaderboard report
Graph Classification REDDIT-B GraphSAGE Accuracy 84.3 #10 of 12 Archive leaderboard report
Graph Classification REDDIT-MULTI-5k GraphSAGE Accuracy 50 #1 of 1 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Graph Neural Network

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections