Home › Datasets › modality › Graphs

Graphs datasets

archive 2025-07-28

282 datasets carry the modality tag "Graphs", ordered by the archive's paper count. Page 4 of 6: 48 shown of 282. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Graphs datasets 145–192 of 282

BGP (Border Gateway Protocol (BGP) Network)
Border Gateway Protocol (BGP) Network describes the Internet's inter-domain structure, where nodes represent the autonomous systems and edges are the business relationships between nodes.
3 papers · 1 benchmark
CTU Relational (The CTU Prague Relational Learning Repository)
The CTU Relational Learning Repository offers relational database datasets to the machine learning community.
3 papers · 0 benchmarks
DPB-5L is a Multilingual KG dataset containing 5 KGs in English, French, Japanese, Greek, and Spanish.
3 papers · 1 benchmark
DIPS-Plus (The Enhanced Database of Interacting Protein Structures for Interface Prediction)
How and where proteins interface with one another can ultimately impact the proteins' functions along with a range of other biological processes.
3 papers · 0 benchmarks
DPB-5L is a Multilingual KG dataset containing 5 KGs in English, French, Japanese, Greek, and Spanish.
3 papers · 1 benchmark
The DeepNets-1M dataset is composed of neural network architectures represented as graphs where nodes are operations (convolution, pooling, etc.) and edges correspond to the forward pass flow of data through the network.
3 papers · 0 benchmarks
The data was collected from the music streaming service Deezer (November 2017).
3 papers · 0 benchmarks
The FB15k-237-low dataset is a variation of the FB15k-237 dataset where relations with a low number of triplets are kept.
3 papers · 0 benchmarks
FireRisk (FireRisk: A Remote Sensing Dataset for Fire Risk Assessment)
In this work, we propose a novel remote sensing dataset, FireRisk, consisting of 7 fire risk classes with a total of 91 872 labelled images for fire risk assessment.
3 papers · 1 benchmark
GlassTemp (Glass Transition Temperature)
The GlassTemp dataset is collected from Polyinfo.
3 papers · 1 benchmark
Question Answering (QA) is a widely-used framework for developing and evaluating an intelligent machine.
3 papers · 0 benchmarks
Operating rooms (ORs) are complex, high-stakes environments requiring precise understanding of interactions among medical staff, tools, and equipment for enhancing surgical assistance, situational awareness, and patient safety.
3 papers · 3 benchmarks
OQMD v1.2 (The Open Quantum Materials Database)
The OQMD is a database of DFT calculated thermodynamic and structural properties of one million materials, created in Chris Wolverton's group at Northwestern University.
3 papers · 1 benchmark
Placenta is a benchmark dataset for node classification in an underexplored domain: predicting microanatomical tissue structures from cell graphs in placenta histology whole slide images.
3 papers · 1 benchmark
SLNET (SLNET: A Redistributable Corpus of 3rd-party Simulink Models)
SLNET is collection of third party Simulink models.
3 papers · 0 benchmarks
Software Heritage is the largest existing public archive of software source code and accompanying development history.
3 papers · 0 benchmarks
SpaGBOL (Spatial-Graph-Based Orientated Cross-View Geo-Localisation)
Cross-View Geo-Localisation within urban regions is challenging in part due to the lack of geo-spatial structuring within current datasets and techniques.
3 papers · 1 benchmark
SupplyGraph (SupplyGraph: A Benchmark Dataset for Supply Chain Planning using Graph Neural Networks)
Graph Neural Networks (GNNs) have gained traction across different domains such as transportation, bio-informatics, language processing, and computer vision.
3 papers · 0 benchmarks
The Little Prince (The Little Prince Corpus)
This corpus is an annotation of the novel The Little Prince by Antoine de Saint-Exupéry, published in 1943.
3 papers · 1 benchmark
UPFD-GOS (User Preference-aware Fake News Detection)
The Gossipcop variant of the UPFD dataset for benchmarking.
3 papers · 1 benchmark
This package provides utilities for generation, filtering, solving, visualizing, and processing of mazes for training ML systems.
3 papers · 0 benchmarks
ApisTox contains molecules in SMILES format for predicting pesticides toxicity to honey bees.
2 papers · 0 benchmarks
Biographical (Biographical: A Semi-Supervised Relation Extraction Dataset)
Biographical is a semi-supervised dataset for RE.
2 papers · 0 benchmarks
The CHILI-100K dataset is a large-scale graph dataset (with overall >183M nodes, >1.2B edges) of nanomaterials generated from experimentally determined crystal structures.
2 papers · 8 benchmarks
The CHILI-3K dataset is a medium-scale graph dataset (with overall >6M nodes, >49M edges) of mono-metallic oxide nanomaterials generated from 12 selected crystal types.
2 papers · 8 benchmarks
ChEMBL is a manually curated database of bioactive molecules with drug-like properties.
2 papers · 0 benchmarks
DPPIN is a collection of dynamic networks, which consists of twelve generated dynamic protein-protein interaction networks of yeast cells, stored in twelve folders.
2 papers · 0 benchmarks
DTGB (Dynamic Text-attributed Graph Benchmark)
We introduce Dynamic Text-attributed Graph Benchmark (DTGB), a collection of large-scale, time-evolving graphs from diverse domains, with nodes and edges enriched by dynamically changing text attributes and categories.
2 papers · 0 benchmarks
This data was collected by performing a breadth-first search on the user-product-review graph until termination, meaning that it is a fairly comprehensive collection of English-language product data.
2 papers · 1 benchmark
GVLQA (Graph Vision-Language Question-Answering)
GVLQA is the first vision-language QA dataset for general graph reasoning.
2 papers · 0 benchmarks
This is a Twitter dataset of 100,386 users along with up to 200 tweets from their timelines with a random-walk-based crawler on the retweet graph, with a subsample of 4,972 which is manually annotated as hateful or not through…
2 papers · 0 benchmarks
HeriGraph (Multimodal Machine Learning Datasets on Graphs of Heritage Values and Attributes)
The dataset contains constructed multi-modal features (visual and textual), pseudo-labels (on heritage values and attributes), and graph structures (with temporal, social, and spatial links) constructed using User-Generated Content data…
2 papers · 0 benchmarks
HiAML Computational Graph (CG) family introduced in "GENNAPE: Towards Generalized Neural Architecture Performance Estimators", accepted to AAAI-23.
2 papers · 0 benchmarks
Inception Computational Graph (CG) family introduced in "GENNAPE: Towards Generalized Neural Architecture Performance Estimators", accepted to AAAI-23.
2 papers · 0 benchmarks
KGRC-RDF-star is an RDF-star dataset converted from KGRC-RDF, which is a Knowledge graph dataset of novel stories.
2 papers · 0 benchmarks
This dataset presents a set of large-scale ridesharing Dial-a-Ride Problem (DARP) instances.
2 papers · 0 benchmarks
An RDF knowledge graph that provides comprehensive, current information about almost 400,000 machine learning publications.
2 papers · 0 benchmarks
MetaVD (Meta Video Dataset)
MetaVD is a Meta Video Dataset for enhancing human action recognition datasets.
2 papers · 0 benchmarks
NBA: This is extended from a Kaggle dataset containing around 400 NBA basketball players.
2 papers · 1 benchmark
The Nations dataset is a small knowledge graph with 14 entities, 55 relations, and 1992 triples describing countries and their political relationships.
2 papers · 0 benchmarks
News SEO Dataset (Detection and Discovery of Misinformation Sources using Attributed Webgraphs)
Search Engine Optimization (SEO) attributes provide strong signals for predicting news site reliability.
2 papers · 0 benchmarks
Ocean Drifters (Madagascar Ocean Drifters)
From Schaub, Michael T., et al.
2 papers · 0 benchmarks
PACE 2016 Feedback Vertex Set (PACE 2016 Track B, Feedback Vertex Set)
This is the dataset used in the PACE 2016 challenge, Track B, which was computing minimal Feedback Vertex Set.
2 papers · 0 benchmarks
PACE 2022 Exact (PACE 2022 Directed Feedback Vertex Set, Exact Track)
This is the set of graphs used in the PACE 2022 challenge for computing the Directed Feedback Vertex Set, from the Exact track.
2 papers · 0 benchmarks
PointPattern is a graph classification dataset constructed by simple point patterns from statistical mechanics.
2 papers · 0 benchmarks
Rent3D++ is an extension of the Rent3D floorplans + photos dataset.
2 papers · 1 benchmark
TOMG-Bench (Text-based Open Molecule Generation Benchmark)
In this paper, we propose Text-based Open Molecule Generation Benchmark (TOMG-Bench), the first benchmark to evaluate the open-domain molecule generation capability of LLMs.
2 papers · 1 benchmark
TRN (Toulouse Road Network)
The Toulouse Road Network dataset describes patches of road maps from the city of Toulouse, represented both as spatial graphs G = (A, X) and as grayscale segmentation images.
2 papers · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.