Home › Datasets › task › Node Classification
Node Classification datasets
archive 2025-07-28
75 datasets carry the task tag "Node Classification" (the task itself: Node Classification), ordered by the archive's paper count. Page 2 of 2: 27 shown of 75. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Node Classification datasets 49–75 of 75
Yelp-Fraud (Multi-relational Graph Dataset for Yelp Spam Review Detection)
Yelp-Fraud is a multi-relational graph dataset built upon the Yelp spam review dataset, which can be used in evaluating graph-based node classification, fraud detection, and anomaly detection models.
13 papers · 3 benchmarks
Aesthetic Visual Analysis is a dataset for aesthetic image assessment that contains over 250,000 images along with a rich variety of meta-data including a large number of aesthetic scores for each image, semantic labels for over 60…
12 papers · 1 benchmark
AMZ Computers is a co-purchase graph extracted from Amazon, where nodes represent products, edges represent the co-purchased relations of products, and features are bag-of-words vectors extracted from product reviews.
9 papers · 1 benchmark
Leonardo Filipe Rodrigues Ribeiro, Pedro H.
9 papers · 1 benchmark
Amazon-Fraud (Multi-relational Graph Dataset for Amazon Fraudulent Account Detection)
Amazon-Fraud is a multi-relational graph dataset built upon the Amazon review dataset, which can be used in evaluating graph-based node classification, fraud detection, and anomaly detection models.
8 papers · 3 benchmarks
Wiki (Web Traffic Time Series Forecasting)
Context There's a story behind every dataset and here's your opportunity to share yours.
8 papers · 3 benchmarks
MAG-Scholar-C is constructed by Bojchevski et al.
5 papers · 1 benchmark
MuMiN is a misinformation graph dataset containing rich social media data (tweets, replies, users, images, articles, hashtags), spanning 21 million tweets belonging to 26 thousand Twitter threads, each of which have been semantically…
5 papers · 0 benchmarks
BGP (Border Gateway Protocol (BGP) Network)
Border Gateway Protocol (BGP) Network describes the Internet's inter-domain structure, where nodes represent the autonomous systems and edges are the business relationships between nodes.
3 papers · 1 benchmark
The data was collected from the music streaming service Deezer (November 2017).
3 papers · 0 benchmarks
Placenta is a benchmark dataset for node classification in an underexplored domain: predicting microanatomical tissue structures from cell graphs in placenta histology whole slide images.
3 papers · 1 benchmark
City-Networks, a transductive learning dataset for testing long-range dependencies in Graph Neural Networks (GNNs).
2 papers · 4 benchmarks
This data was collected by performing a breadth-first search on the user-product-review graph until termination, meaning that it is a fairly comprehensive collection of English-language product data.
2 papers · 1 benchmark
A new fraud detection dataset FDCompCN for detecting financial statement fraud of companies in China.
2 papers · 1 benchmark
HeriGraph (Multimodal Machine Learning Datasets on Graphs of Heritage Values and Attributes)
The dataset contains constructed multi-modal features (visual and textual), pseudo-labels (on heritage values and attributes), and graph structures (with temporal, social, and spatial links) constructed using User-Generated Content data…
2 papers · 0 benchmarks
NBA: This is extended from a Kaggle dataset containing around 400 NBA basketball players.
2 papers · 1 benchmark
Classifying all cells in an organ is a relevant and difficult problem from plant developmental biology.
1 paper · 1 benchmark
FinDKG: The Global Financial Dynamic Knowledge Graph Dataset FinDKG is an open-source dataset focused on creating a temporally-resolved Financial Dynamic Knowledge Graph.
1 paper · 0 benchmarks
The MAPLE benchmark constructed by us contains 20 datasets across 19 fields for scientific literature tagging.
1 paper · 0 benchmarks
This is the large version of the MuMiN dataset.
1 paper · 1 benchmark
This is the medium version of the MuMiN dataset.
1 paper · 1 benchmark
This is the small version of the MuMiN dataset.
1 paper · 1 benchmark
SAGC-A68 (A space access graph dataset for the classification of spaces and space elements in apartment buildings)
The analysis of building models for usable area, building safety, and energy efficiency requires accurate classification data of spaces and space elements.
1 paper · 0 benchmarks
Twitter-HyDrug is a real-world hypergraph data that describes the drug trafficking communities on Twitter.
1 paper · 1 benchmark
This benchmark hypergraph dataset, Twitter-HyDrug-UR, is derived from Twitter-HyDrug by HyGCL-DC.
1 paper · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.