Home › Datasets › modality › Graphs

Graphs datasets

archive 2025-07-28

282 datasets carry the modality tag "Graphs", ordered by the archive's paper count. Page 2 of 6: 48 shown of 282. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Graphs datasets 49–96 of 282

Dataset contains 33,010 molecule-description pairs split into 80\%/10\%/10\% train/val/test splits.
43 papers · 4 benchmarks
This dataset contains product reviews and metadata from Amazon, including 142.8 million reviews spanning May 1996 - July 2014.
41 papers · 5 benchmarks
The Argoverse 2 Motion Forecasting Dataset is a curated collection of 250,000 scenarios for training and validation.
39 papers · 0 benchmarks
Tolokers is a crowdsourcing platform workers network based on data provided by Toloka.
38 papers · 1 benchmark
Arxiv HEP-TH (high energy physics theory) citation graph is from the e-print arXiv and covers all the citations within a dataset of 27,770 papers with 352,807 edges.
35 papers · 5 benchmarks
STRING is a collection of protein-protein interaction (PPI) networks.
35 papers · 0 benchmarks
BioGRID (Biological General Repository for Interaction Datasets)
BioGRID is a biomedical interaction repository with data compiled through comprehensive curation efforts.
34 papers · 2 benchmarks
CSL is a synthetic dataset introduced in Murphy et al.
34 papers · 2 benchmarks
The Ciao dataset contains rating information of users given to items, and also contain item category information.
34 papers · 1 benchmark
OGB-LSC (OGB Large-Scale Challenge)
OGB Large-Scale Challenge (OGB-LSC) is a collection of three real-world datasets for advancing the state-of-the-art in large-scale graph ML.
34 papers · 3 benchmarks
minesweeper is a synthetic graph emulating the eponymous game.
34 papers · 1 benchmark
Decagon (Bio-decagon)
Bio-decagon is a dataset for polypharmacy side effect identification problem framed as a multirelational link prediction problem in a two-layer multimodal graph/network of two node types: drugs and proteins.
33 papers · 1 benchmark
EmailEU is a directed temporal network constructed from email exchanges in a large European research institution for a 803-day period.
33 papers · 0 benchmarks
The Pinterest dataset contains more than 1 million images associated to Pinterest users’ who have “pinned” them.
33 papers · 1 benchmark
amazon-ratings is a product co-purchasing network based on data from SNAP datasets
33 papers · 1 benchmark
Worldtree is a corpus of explanation graphs, explanatory role ratings, and associated tablestore.
32 papers · 0 benchmarks
LDC2017T10 (Abstract Meaning Representation (AMR) Annotation Release 2.0)
Abstract Meaning Representation (AMR) Annotation Release 2.0 was developed by the Linguistic Data Consortium (LDC), SDL/Language Weaver, Inc., the University of Colorado's Computational Language and Educational Research group and the…
27 papers · 1 benchmark
Reddit12k contains 11929 graphs each corresponding to an online discussion thread where nodes represent users, and an edge represents the fact that one of the two users responded to the comment of the other user.
24 papers · 1 benchmark
node classification on twitch-gamers
24 papers · 2 benchmarks
A social network of LastFM users which was collected from the public API in March 2020.
21 papers · 0 benchmarks
Node classification on Chameleon with the fixed 48%/32%/20% splits provided by Geom-GCN.
20 papers · 2 benchmarks
GEOM-DRUGS is a dataset of 430,000 large organic molecules of up to 180 atoms from Axelrod and Gómez-Bombarelli, Nature Scientific Data, 2022.
20 papers · 1 benchmark
AGENDA (Abstract GENeration DAtaset)
Abstract GENeration DAtaset (AGENDA) is a dataset of knowledge graphs paired with scientific abstracts.
19 papers · 1 benchmark
Node classification on Deezer Europe with 50%/25%/25% random splits for training/validation/test.
19 papers · 1 benchmark
Node classification on Film with 60%/20%/20% random splits for training/validation/test.
19 papers · 1 benchmark
Node classification on Squirrel with the fixed 48%/32%/20% splits provided by Geom-GCN.
19 papers · 2 benchmarks
Node classification on Squirrel with 60%/20%/20% random splits for training/validation/test.
19 papers · 1 benchmark
UMLS (Unified Medical Language System)
The Unified Medical Language System (UMLS) is a comprehensive resource that integrates and disseminates essential terminology, classification standards, and coding systems.
19 papers · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
18 papers · 1 benchmark
Node classification on PubMed with 60%/20%/20% random splits for training/validation/test.
18 papers · 1 benchmark
Node classification on Wisconsin with 60%/20%/20% random splits for training/validation/test.
18 papers · 1 benchmark
Yeast dataset consists of a protein-protein interaction network.
18 papers · 0 benchmarks
Regression dataset for molecular docking scores (predicted molecule-protein binding affinity).
18 papers · 4 benchmarks
Node classification on Chameleon with 60%/20%/20% random splits for training/validation/test.
17 papers · 2 benchmarks
MalNet is a large public graph database, representing a large-scale ontology of software function call graphs.
17 papers · 2 benchmarks
Node classification on Cornell with the fixed 48%/32%/20% splits provided by Geom-GCN.
16 papers · 2 benchmarks
Node classification on Cornell with 60%/20%/20% random splits for training/validation/test.
16 papers · 2 benchmarks
SketchGraphs is a dataset of 15 million sketches extracted from real-world CAD models intended to facilitate research in both ML-aided design and geometric program induction.
16 papers · 0 benchmarks
Node classification on Texas with 60%/20%/20% random splits for training/validation/test.
16 papers · 1 benchmark
This dataset is a Wikipedia dump, split by relations to perform Few-Shot Knowledge Graph Completion.
16 papers · 0 benchmarks
BeerAdvocate is a dataset that consists of beer reviews from beeradvocate.
15 papers · 1 benchmark
Node classification on Citeseer with the fixed 48%/32%/20% splits provided by Geom-GCN.
15 papers · 1 benchmark
Node classification on Cora with the fixed 48%/32%/20% splits provided by Geom-GCN.
15 papers · 1 benchmark
Node classification on PubMed with the fixed 48%/32%/20% splits provided by Geom-GCN.
15 papers · 1 benchmark
Node classification on Wisconsin with the fixed 48%/32%/20% splits provided by Geom-GCN.
15 papers · 2 benchmarks
Node classification on Film with the fixed 48%/32%/20% splits provided by Geom-GCN.
14 papers · 2 benchmarks
Linux (Linux Program Dependence Graphs)
The LINUX dataset consists of 48,747 Program Dependence Graphs (PDG) generated from the Linux kernel.
14 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.