Home › Datasets › modality › Graphs
Graphs datasets
archive 2025-07-28
282 datasets carry the modality tag "Graphs", ordered by the archive's paper count. Page 1 of 6: 48 shown of 282. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Graphs datasets 1–48 of 282
The PubMed dataset consists of 19717 scientific publications from PubMed database pertaining to diabetes classified into one of three classes.
1,236 papers · 19 benchmarks
OGB (Open Graph Benchmark)
The Open Graph Benchmark (OGB) is a collection of realistic, large-scale, and diverse benchmark datasets for machine learning on graphs.
1,000 papers · 17 benchmarks
The Reddit dataset is a graph dataset from Reddit posts made in the month of September, 2014.
699 papers · 8 benchmarks
The FB15k dataset contains knowledge base relation triples and textual mentions of Freebase entity pairs.
641 papers · 6 benchmarks
The Cora dataset consists of 2708 scientific publications classified into one of seven classes.
602 papers · 18 benchmarks
DBpedia (from "DB" for "database") is a project aiming to extract structured content from the information created in the Wikipedia project.
597 papers · 4 benchmarks
The WN18 dataset has 18 relations scraped from WordNet for roughly 41,000 synsets, resulting in 141,442 triplets.
485 papers · 3 benchmarks
FB15k-237 is a link prediction dataset created from FB15k.
452 papers · 3 benchmarks
FrameNet is a linguistic knowledge graph containing information about lexical and predicate argument semantics of the English language.
444 papers · 0 benchmarks
WN18RR is a link prediction dataset created from WN18, which is a subset of WordNet.
394 papers · 3 benchmarks
The CiteSeer dataset consists of 3312 scientific publications classified into one of six classes.
381 papers · 13 benchmarks
PROTEINS is a dataset of proteins that are classified as enzymes or non-enzymes.
371 papers · 1 benchmark
YAGO (Yet Another Great Ontology)
Yet Another Great Ontology (YAGO) is a Knowledge Graph that augments WordNet with common knowledge facts extracted from Wikipedia, converting WordNet from a primarily linguistic resource to a common knowledge base.
359 papers · 5 benchmarks
IMDB-BINARY is a movie collaboration dataset that consists of the ego-networks of 1,000 actors/actresses who played roles in movies in IMDB.
326 papers · 2 benchmarks
In particular, MUTAG is a collection of nitroaromatic compounds and the goal is to predict their mutagenicity on Salmonella typhimurium.
274 papers · 3 benchmarks
The NCI1 dataset comes from the cheminformatics domain, where each input graph is used as representation of a chemical compound: each vertex stands for an atom of the molecule, and edges between vertices represent bonds between atoms.
260 papers · 2 benchmarks
COLLAB is a scientific collaboration dataset.
259 papers · 2 benchmarks
IMDB-MULTI is a relational dataset that consists of a network of 1000 actors or actresses who played roles in movies in IMDB.
243 papers · 3 benchmarks
MoleculeNet is a large scale benchmark for molecular machine learning.
240 papers · 1 benchmark
DBLP (Citation Network Dataset)
The DBLP is a citation network dataset.
218 papers · 4 benchmarks
The data was collected from the English Wikipedia (December 2018).
208 papers · 1 benchmark
ENZYMES is a dataset of 600 protein tertiary structures obtained from the BRENDA enzyme database.
202 papers · 1 benchmark
SNAP (Stanford Large Network Dataset Collection)
SNAP is a collection of large network datasets.
174 papers · 0 benchmarks
CLUSTER is a node classification tasks generated with Stochastic Block Models, which is widely used to model communities in social networks by modulating the intra- and extra-communities connections, thereby controlling the difficulty of…
159 papers · 1 benchmark
The Materials Project is a collection of chemical compounds labelled with different attributes.
157 papers · 1 benchmark
PATTERN is a node classification tasks generated with Stochastic Block Models, which is widely used to model communities in social networks by modulating the intra- and extra-communities connections, thereby controlling the difficulty of…
153 papers · 1 benchmark
REDDIT-BINARY consists of graphs corresponding to online discussions on Reddit.
150 papers · 1 benchmark
WebKB is a dataset that includes web pages from computer science departments of various universities.
113 papers · 2 benchmarks
PTC (Predictive Toxicology Challenge)
PTC is a collection of 344 chemical compounds represented as graphs which report the carcinogenicity for rats.
111 papers · 1 benchmark
Wiki-CS is a Wikipedia-based dataset for benchmarking Graph Neural Networks.
111 papers · 1 benchmark
Tudataset: A collection of benchmark datasets for learning with graphs
96 papers · 1 benchmark
Orkut is a social network dataset consisting of friendship social network and ground-truth communities from Orkut.com on-line social network where users form friendship each other.
85 papers · 0 benchmarks
The AMiner Dataset is a collection of different relational datasets.
80 papers · 1 benchmark
MOSES (Molecular sets (MOSES))
The set is based on the ZINC Clean Leads collection.
79 papers · 0 benchmarks
Reddit-5K is a relational dataset extracted from Reddit.
78 papers · 1 benchmark
RadGraph (RadGraph: Extracting Clinical Entities and Relations from Radiology Reports)
RadGraph is a dataset of entities and relations in radiology reports based on our novel information extraction schema, consisting of 600 reports with 30K radiologist annotations and 221K reports with 10.5M automatically generated…
78 papers · 0 benchmarks
The Long Range Graph Benchmark (LRGB) is a collection of 5 graph learning datasets that arguably require long-range reasoning to achieve strong performance in a given task.
76 papers · 4 benchmarks
The Reuters-21578 dataset is a collection of documents with news articles.
66 papers · 5 benchmarks
Friendster is an on-line gaming network.
65 papers · 0 benchmarks
GAP (GAP Benchmark Suite)
GAP is a graph processing benchmark suite with the goal of helping to standardize graph processing evaluations.
60 papers · 1 benchmark
Node classification on Penn94
60 papers · 2 benchmarks
The Slashdot dataset is a relational dataset obtained from Slashdot.
57 papers · 2 benchmarks
The Epinions dataset is built form a who-trust-whom online social network of a general consumer review site Epinions.com.
54 papers · 2 benchmarks
This corpus includes annotations of cancer-related PubMed articles, covering 3 full papers (PMID:24651010, PMID:11777939, PMID:15630473) as well as the result sections of 46 additional PubMed papers.
53 papers · 1 benchmark
MMKG is a collection of three knowledge graphs for link prediction and entity matching research.
49 papers · 3 benchmarks
Roman-empire is a word dependency graph based on the Roman Empire article from the English Wikipedia.
49 papers · 1 benchmark
Questions is an interaction graph of users of a question-answering website based on data provided by Yandex Q.
46 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.