Home › Datasets › modality › Graphs
Graphs datasets
archive 2025-07-28
282 datasets carry the modality tag "Graphs", ordered by the archive's paper count. Page 3 of 6: 48 shown of 282. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Graphs datasets 97–144 of 282
Node classification on Texas with the fixed 48%/32%/20% splits provided by Geom-GCN.
14 papers · 2 benchmarks
Provides detailed, graph-based annotations of social situations depicted in movie clips.
13 papers · 0 benchmarks
Mutagenicity is a chemical compound dataset of drugs, which can be categorized into two classes: mutagen and non-mutagen.
13 papers · 1 benchmark
The SARDet-100K dataset encompasses a total of 116,598 images, and 245,653 instances distributed across six categories: Aircraft, Ship, Car, Bridge, Tank, and Harbor.
13 papers · 1 benchmark
UPFD (User Preference-aware Fake News Detection)
For benchmarking, please refer to its variant UPFD-POL and UPFD-GOS.
13 papers · 0 benchmarks
Yelp-Fraud (Multi-relational Graph Dataset for Yelp Spam Review Detection)
Yelp-Fraud is a multi-relational graph dataset built upon the Yelp spam review dataset, which can be used in evaluating graph-based node classification, fraud detection, and anomaly detection models.
13 papers · 3 benchmarks
The BotNet dataset is a set of topological botnet detection datasets forgraph neural networks.
12 papers · 0 benchmarks
The Dataset is part of the KELM corpus This is the Wikipedia text--Wikidata KG aligned corpus used to train the data-to-text generation model.
12 papers · 1 benchmark
4D-OR includes a total of 6734 scenes, recorded by six calibrated RGB-D Kinect sensors 1 mounted to the ceiling of the OR, with one frame-per-second, providing synchronized RGB and depth images.
11 papers · 3 benchmarks
GenWiki is a large-scale dataset for knowledge graph-to-text (G2T) and text-to-knowledge graph (T2G) conversion.
10 papers · 3 benchmarks
This relational database consists of 24 unique names in two families (they have equivalent structures).
10 papers · 0 benchmarks
Arxiv ASTRO-PH (Astro Physics) collaboration network is from the e-print arXiv and covers scientific collaborations between authors papers submitted to Astro Physics category.
10 papers · 0 benchmarks
The Ecoli dataset is a dataset for protein localization.
9 papers · 0 benchmarks
LDC2020T02 (Abstract Meaning Representation (AMR) Annotation Release 3.0)
Abstract Meaning Representation (AMR) Annotation Release 3.0 was developed by the Linguistic Data Consortium (LDC), SDL/Language Weaver, Inc., the University of Colorado's Computational Language and Educational Research group and the…
9 papers · 1 benchmark
Mindboggle is a large publicly available dataset of manually labeled brain MRI.
9 papers · 0 benchmarks
Leonardo Filipe Rodrigues Ribeiro, Pedro H.
9 papers · 1 benchmark
Amazon-Fraud (Multi-relational Graph Dataset for Amazon Fraudulent Account Detection)
Amazon-Fraud is a multi-relational graph dataset built upon the Amazon review dataset, which can be used in evaluating graph-based node classification, fraud detection, and anomaly detection models.
8 papers · 3 benchmarks
ProteinKG25 is a large-scale KG dataset with aligned descriptions and protein sequences respectively to GO terms and proteins entities.
8 papers · 0 benchmarks
Random sampled instances of the Capacitated Vehicle Routing Problem with Time Windows (CVRPTW) for 20, 50 and 100 customer nodes.
7 papers · 0 benchmarks
Language-molecule models have emerged as an exciting direction for molecular discovery and understanding.
7 papers · 1 benchmark
New3, a set of 527 instances from AMR 3.0, whose original source was the LORELEI DARPA project – not included in the AMR 2.0 training set – consisting of excerpts from newswires and online forum.
7 papers · 1 benchmark
GRB (Graph Robustness Benchmark)
Graph Robustness Benchmark (GRB) provides scalable, unified, modular, and reproducible evaluation on the adversarial robustness of graph machine learning models.
6 papers · 0 benchmarks
1.0 Version of OpenEA benchmark datasets.
6 papers · 4 benchmarks
SemOpenAlex is an extensive RDF knowledge graph that contains over 26 billion triples about scientific publications and their associated entities, such as authors, institutions, journals, and concepts.
6 papers · 0 benchmarks
We present a further analysis of visual modality incompleteness, benchmarking latest MMEA models on our proposed dataset MMEA-UMVM.
6 papers · 3 benchmarks
The ZS-F-VQA dataset is a new split of the F-VQA dataset for zero-shot problem.
6 papers · 1 benchmark
KG20C (A scholarly knowledge graph benchmark dataset)
KG20C is a Knowledge Graph about high quality papers from 20 top computer science Conferences.
5 papers · 1 benchmark
MarKG (Multimodal analogical reasoning Knowledge Graph)
The MarKG dataset has 11,292 entities, 192 relations and 76,424 images, including 2,063 analogy entities and 27 analogy relations.
5 papers · 0 benchmarks
MuMiN is a misinformation graph dataset containing rich social media data (tweets, replies, users, images, articles, hashtags), spanning 21 million tweets belonging to 26 thousand Twitter threads, each of which have been semantically…
5 papers · 0 benchmarks
This is a catalogue and repository of network datasets with the aim of aiding scientific research.
5 papers · 0 benchmarks
RARE (Randomized AMRs with Rewired Edges)
RARE consists of English AMR pairs with similarity scores that reflect the structural differences between them.
5 papers · 1 benchmark
SSN (Semantic Scholar Network)
SSN (short for Semantic Scholar Network) is a scientific papers summarization dataset which contains 141K research papers in different domains and 661K citation relationships.
5 papers · 0 benchmarks
This is a benchmark set for Traveling salesman problem (TSP) with characteristics that are different from the existing benchmark sets.
5 papers · 1 benchmark
The Vent dataset is a large annotated dataset of text, emotions, and social connections.
5 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
4 papers · 0 benchmarks
ACM (Association for Computing Machinery
Active Contour Model
algebraic collective model
and-Compare Module
Active Contour Models)
The ACM dataset contains papers published in KDD, SIGMOD, SIGCOMM, MobiCOMM, and VLDB and are divided into three classes (Database, Wireless Communication, Data Mining).
4 papers · 1 benchmark
Amazon Fine Foods is a dataset that consists of reviews of fine foods from amazon.
4 papers · 0 benchmarks
Chickenpox Cases in Hungary is a spatio-temporal dataset of weekly chickenpox (childhood disease) cases from Hungary.
4 papers · 0 benchmarks
DPB-5L is a Multilingual KG dataset containing 5 KGs in English, French, Japanese, Greek, and Spanish.
4 papers · 1 benchmark
GDSC (Genomics of Drug Sensitivity in Cancer)
We have characterized 1000 human cancer cell lines and screened them with 100s of compounds.
4 papers · 1 benchmark
Money laundering is a multi-billion dollar issue.
4 papers · 0 benchmarks
The IS-A dataset is a dataset of relations extracted from a medical ontology.
4 papers · 0 benchmarks
InferWiki is a Knowledge Graph Completion (KGC) dataset that improves upon existing benchmarks in inferential ability, assumptions, and patterns.
4 papers · 0 benchmarks
This is the set of instances use in the PACE 2018 competition, of optimal Steiner Tree computation.
4 papers · 0 benchmarks
VirtualHome2KG is a system for constructing and augmenting knowledge graphs (KGs) of daily living activities using virtual space.
4 papers · 0 benchmarks
WikiGraphs is a dataset of Wikipedia articles each paired with a knowledge graph, to facilitate the research in conditional text generation, graph generation and graph representation learning.
4 papers · 1 benchmark
Arxiv GR-QC (General Relativity and Quantum Cosmology collaboration network)
Arxiv GR-QC (General Relativity and Quantum Cosmology) collaboration network is from the e-print arXiv and covers scientific collaborations between authors papers submitted to General Relativity and Quantum Cosmology category.
3 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.