Home › Datasets › modality › Graphs
Graphs datasets
archive 2025-07-28
282 datasets carry the modality tag "Graphs", ordered by the archive's paper count. Page 4 of 6: 48 shown of 282. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Graphs datasets 145–192 of 282
BGP (Border Gateway Protocol (BGP) Network)
Border Gateway Protocol (BGP) Network describes the Internet's inter-domain structure, where nodes represent the autonomous systems and edges are the business relationships between nodes.
3 papers · 1 benchmark
The CTU Relational Learning Repository offers relational database datasets to the machine learning community.
3 papers · 0 benchmarks
DPB-5L is a Multilingual KG dataset containing 5 KGs in English, French, Japanese, Greek, and Spanish.
3 papers · 1 benchmark
DIPS-Plus (The Enhanced Database of Interacting Protein Structures for Interface Prediction)
How and where proteins interface with one another can ultimately impact the proteins' functions along with a range of other biological processes.
3 papers · 0 benchmarks
DPB-5L is a Multilingual KG dataset containing 5 KGs in English, French, Japanese, Greek, and Spanish.
3 papers · 1 benchmark
The DeepNets-1M dataset is composed of neural network architectures represented as graphs where nodes are operations (convolution, pooling, etc.) and edges correspond to the forward pass flow of data through the network.
3 papers · 0 benchmarks
The data was collected from the music streaming service Deezer (November 2017).
3 papers · 0 benchmarks
The FB15k-237-low dataset is a variation of the FB15k-237 dataset where relations with a low number of triplets are kept.
3 papers · 0 benchmarks
FireRisk (FireRisk: A Remote Sensing Dataset for Fire Risk Assessment)
In this work, we propose a novel remote sensing dataset, FireRisk, consisting of 7 fire risk classes with a total of 91 872 labelled images for fire risk assessment.
3 papers · 1 benchmark
The GlassTemp dataset is collected from Polyinfo.
3 papers · 1 benchmark
Question Answering (QA) is a widely-used framework for developing and evaluating an intelligent machine.
3 papers · 0 benchmarks
Operating rooms (ORs) are complex, high-stakes environments requiring precise understanding of interactions among medical staff, tools, and equipment for enhancing surgical assistance, situational awareness, and patient safety.
3 papers · 3 benchmarks
OQMD v1.2 (The Open Quantum Materials Database)
The OQMD is a database of DFT calculated thermodynamic and structural properties of one million materials, created in Chris Wolverton's group at Northwestern University.
3 papers · 1 benchmark
Placenta is a benchmark dataset for node classification in an underexplored domain: predicting microanatomical tissue structures from cell graphs in placenta histology whole slide images.
3 papers · 1 benchmark
SLNET (SLNET: A Redistributable Corpus of 3rd-party Simulink Models)
SLNET is collection of third party Simulink models.
3 papers · 0 benchmarks
Software Heritage is the largest existing public archive of software source code and accompanying development history.
3 papers · 0 benchmarks
SpaGBOL (Spatial-Graph-Based Orientated Cross-View Geo-Localisation)
Cross-View Geo-Localisation within urban regions is challenging in part due to the lack of geo-spatial structuring within current datasets and techniques.
3 papers · 1 benchmark
SupplyGraph (SupplyGraph: A Benchmark Dataset for Supply Chain Planning using Graph Neural Networks)
Graph Neural Networks (GNNs) have gained traction across different domains such as transportation, bio-informatics, language processing, and computer vision.
3 papers · 0 benchmarks
This corpus is an annotation of the novel The Little Prince by Antoine de Saint-Exupéry, published in 1943.
3 papers · 1 benchmark
UPFD-GOS (User Preference-aware Fake News Detection)
The Gossipcop variant of the UPFD dataset for benchmarking.
3 papers · 1 benchmark
This package provides utilities for generation, filtering, solving, visualizing, and processing of mazes for training ML systems.
3 papers · 0 benchmarks
ApisTox contains molecules in SMILES format for predicting pesticides toxicity to honey bees.
2 papers · 0 benchmarks
Biographical (Biographical: A Semi-Supervised Relation Extraction Dataset)
Biographical is a semi-supervised dataset for RE.
2 papers · 0 benchmarks
The CHILI-100K dataset is a large-scale graph dataset (with overall >183M nodes, >1.2B edges) of nanomaterials generated from experimentally determined crystal structures.
2 papers · 8 benchmarks
The CHILI-3K dataset is a medium-scale graph dataset (with overall >6M nodes, >49M edges) of mono-metallic oxide nanomaterials generated from 12 selected crystal types.
2 papers · 8 benchmarks
ChEMBL is a manually curated database of bioactive molecules with drug-like properties.
2 papers · 0 benchmarks
DPPIN is a collection of dynamic networks, which consists of twelve generated dynamic protein-protein interaction networks of yeast cells, stored in twelve folders.
2 papers · 0 benchmarks
DTGB (Dynamic Text-attributed Graph Benchmark)
We introduce Dynamic Text-attributed Graph Benchmark (DTGB), a collection of large-scale, time-evolving graphs from diverse domains, with nodes and edges enriched by dynamically changing text attributes and categories.
2 papers · 0 benchmarks
This data was collected by performing a breadth-first search on the user-product-review graph until termination, meaning that it is a fairly comprehensive collection of English-language product data.
2 papers · 1 benchmark
GVLQA (Graph Vision-Language Question-Answering)
GVLQA is the first vision-language QA dataset for general graph reasoning.
2 papers · 0 benchmarks
This is a Twitter dataset of 100,386 users along with up to 200 tweets from their timelines with a random-walk-based crawler on the retweet graph, with a subsample of 4,972 which is manually annotated as hateful or not through…
2 papers · 0 benchmarks
HeriGraph (Multimodal Machine Learning Datasets on Graphs of Heritage Values and Attributes)
The dataset contains constructed multi-modal features (visual and textual), pseudo-labels (on heritage values and attributes), and graph structures (with temporal, social, and spatial links) constructed using User-Generated Content data…
2 papers · 0 benchmarks
HiAML Computational Graph (CG) family introduced in "GENNAPE: Towards Generalized Neural Architecture Performance Estimators", accepted to AAAI-23.
2 papers · 0 benchmarks
Inception Computational Graph (CG) family introduced in "GENNAPE: Towards Generalized Neural Architecture Performance Estimators", accepted to AAAI-23.
2 papers · 0 benchmarks
KGRC-RDF-star is an RDF-star dataset converted from KGRC-RDF, which is a Knowledge graph dataset of novel stories.
2 papers · 0 benchmarks
This dataset presents a set of large-scale ridesharing Dial-a-Ride Problem (DARP) instances.
2 papers · 0 benchmarks
An RDF knowledge graph that provides comprehensive, current information about almost 400,000 machine learning publications.
2 papers · 0 benchmarks
MetaVD is a Meta Video Dataset for enhancing human action recognition datasets.
2 papers · 0 benchmarks
NBA: This is extended from a Kaggle dataset containing around 400 NBA basketball players.
2 papers · 1 benchmark
The Nations dataset is a small knowledge graph with 14 entities, 55 relations, and 1992 triples describing countries and their political relationships.
2 papers · 0 benchmarks
News SEO Dataset (Detection and Discovery of Misinformation Sources using Attributed Webgraphs)
Search Engine Optimization (SEO) attributes provide strong signals for predicting news site reliability.
2 papers · 0 benchmarks
From Schaub, Michael T., et al.
2 papers · 0 benchmarks
This is the dataset used in the PACE 2016 challenge, Track B, which was computing minimal Feedback Vertex Set.
2 papers · 0 benchmarks
This is the set of graphs used in the PACE 2022 challenge for computing the Directed Feedback Vertex Set, from the Exact track.
2 papers · 0 benchmarks
PointPattern is a graph classification dataset constructed by simple point patterns from statistical mechanics.
2 papers · 0 benchmarks
Rent3D++ is an extension of the Rent3D floorplans + photos dataset.
2 papers · 1 benchmark
TOMG-Bench (Text-based Open Molecule Generation Benchmark)
In this paper, we propose Text-based Open Molecule Generation Benchmark (TOMG-Bench), the first benchmark to evaluate the open-domain molecule generation capability of LLMs.
2 papers · 1 benchmark
TRN (Toulouse Road Network)
The Toulouse Road Network dataset describes patches of road maps from the city of Toulouse, represented both as spatial graphs G = (A, X) and as grayscale segmentation images.
2 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.