Home › Datasets › task › Node Classification

Node Classification datasets

archive 2025-07-28

75 datasets carry the task tag "Node Classification" (the task itself: Node Classification), ordered by the archive's paper count. Page 1 of 2: 48 shown of 75. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Node Classification datasets 1–48 of 75

The PubMed dataset consists of 19717 scientific publications from PubMed database pertaining to diabetes classified into one of three classes.
1,236 papers · 19 benchmarks
OGB (Open Graph Benchmark)
The Open Graph Benchmark (OGB) is a collection of realistic, large-scale, and diverse benchmark datasets for machine learning on graphs.
1,000 papers · 17 benchmarks
The Flickr30k dataset contains 31,000 images collected from Flickr, together with 5 reference sentences provided by human annotators.
880 papers · 9 benchmarks
The Reddit dataset is a graph dataset from Reddit posts made in the month of September, 2014.
699 papers · 8 benchmarks
The Cora dataset consists of 2708 scientific publications classified into one of seven classes.
602 papers · 18 benchmarks
The CiteSeer dataset consists of 3312 scientific publications classified into one of six classes.
381 papers · 13 benchmarks
PPI (Protein-Protein Interactions (PPI))
protein roles—in terms of their cellular functions from gene ontology—in various protein-protein interaction (PPI) graphs, with each graph corresponding to a different human tissue [41].
309 papers · 2 benchmarks
In particular, MUTAG is a collection of nitroaromatic compounds and the goal is to predict their mutagenicity on Salmonella typhimurium.
274 papers · 3 benchmarks
MoleculeNet is a large scale benchmark for molecular machine learning.
240 papers · 1 benchmark
DBLP (Citation Network Dataset)
The DBLP is a citation network dataset.
218 papers · 4 benchmarks
Wiki Squirrel (Wikipedia Squirrel)
The data was collected from the English Wikipedia (December 2018).
208 papers · 1 benchmark
PASCAL VOC (PASCAL Visual Object Classes Challenge)
The PASCAL Visual Object Classes (VOC) 2012 dataset contains 20 object categories including vehicles, household, animals, and other: aeroplane, bicycle, boat, bus, car, motorbike, train, bottle, chair, dining table, potted plant, sofa,…
198 papers · 18 benchmarks
NELL (Never Ending Language Learning)
NELL is a dataset built from the Web via an intelligent agent called Never-Ending Language Learner.
177 papers · 2 benchmarks
CLUSTER is a node classification tasks generated with Stochastic Block Models, which is widely used to model communities in social networks by modulating the intra- and extra-communities connections, thereby controlling the difficulty of…
159 papers · 1 benchmark
PATTERN is a node classification tasks generated with Stochastic Block Models, which is widely used to model communities in social networks by modulating the intra- and extra-communities connections, thereby controlling the difficulty of…
153 papers · 1 benchmark
AVA (Atomic Visual Actions)
AVA is a project that provides audiovisual annotations of video for improving our understanding of human activity.
113 papers · 7 benchmarks
WebKB is a dataset that includes web pages from computer science departments of various universities.
113 papers · 2 benchmarks
Wiki-CS is a Wikipedia-based dataset for benchmarking Graph Neural Networks.
111 papers · 1 benchmark
The Yelp Dataset is a valuable resource for academic research, teaching, and learning.
86 papers · 15 benchmarks
The AMiner Dataset is a collection of different relational datasets.
80 papers · 1 benchmark
The Long Range Graph Benchmark (LRGB) is a collection of 5 graph learning datasets that arguably require long-range reasoning to achieve strong performance in a given task.
76 papers · 4 benchmarks
Friendster is an on-line gaming network.
65 papers · 0 benchmarks
Roman-empire is a word dependency graph based on the Roman Empire article from the English Wikipedia.
49 papers · 1 benchmark
Questions is an interaction graph of users of a question-answering website based on data provided by Yandex Q.
46 papers · 1 benchmark
Tolokers is a crowdsourcing platform workers network based on data provided by Toloka.
38 papers · 1 benchmark
BioGRID (Biological General Repository for Interaction Datasets)
BioGRID is a biomedical interaction repository with data compiled through comprehensive curation efforts.
34 papers · 2 benchmarks
OGB-LSC (OGB Large-Scale Challenge)
OGB Large-Scale Challenge (OGB-LSC) is a collection of three real-world datasets for advancing the state-of-the-art in large-scale graph ML.
34 papers · 3 benchmarks
minesweeper is a synthetic graph emulating the eponymous game.
34 papers · 1 benchmark
amazon-ratings is a product co-purchasing network based on data from SNAP datasets
33 papers · 1 benchmark
node classification on twitch-gamers
24 papers · 2 benchmarks
Node classification on Chameleon with the fixed 48%/32%/20% splits provided by Geom-GCN.
20 papers · 2 benchmarks
Node classification on Film with 60%/20%/20% random splits for training/validation/test.
19 papers · 1 benchmark
Node classification on Squirrel with the fixed 48%/32%/20% splits provided by Geom-GCN.
19 papers · 2 benchmarks
Node classification on Squirrel with 60%/20%/20% random splits for training/validation/test.
19 papers · 1 benchmark
Node classification on PubMed with 60%/20%/20% random splits for training/validation/test.
18 papers · 1 benchmark
Yeast dataset consists of a protein-protein interaction network.
18 papers · 0 benchmarks
Node classification on Chameleon with 60%/20%/20% random splits for training/validation/test.
17 papers · 2 benchmarks
Node classification on Cornell with the fixed 48%/32%/20% splits provided by Geom-GCN.
16 papers · 2 benchmarks
Node classification on Cornell with 60%/20%/20% random splits for training/validation/test.
16 papers · 2 benchmarks
Node classification on Citeseer with the fixed 48%/32%/20% splits provided by Geom-GCN.
15 papers · 1 benchmark
Node classification on Cora with the fixed 48%/32%/20% splits provided by Geom-GCN.
15 papers · 1 benchmark
Node classification on PubMed with the fixed 48%/32%/20% splits provided by Geom-GCN.
15 papers · 1 benchmark
Node classification on Wisconsin with the fixed 48%/32%/20% splits provided by Geom-GCN.
15 papers · 2 benchmarks
Node classification on Film with the fixed 48%/32%/20% splits provided by Geom-GCN.
14 papers · 2 benchmarks
Node classification on Texas with the fixed 48%/32%/20% splits provided by Geom-GCN.
14 papers · 2 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.