Home › Datasets › modality › Graphs

Graphs datasets

archive 2025-07-28

282 datasets carry the modality tag "Graphs", ordered by the archive's paper count. Page 5 of 6: 48 shown of 282. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Graphs datasets 193–240 of 282

TTE-A&O (Travel Time Estimation: Abakan and Omsk)
The dataset includes two parts corresponding to the cities of Abakan (65524 nodes, 340012 edges) and Omsk (231688 nodes, 1149492 edges).
2 papers · 1 benchmark
Two-Path Computational Graph (CG) family introduced in "GENNAPE: Towards Generalized Neural Architecture Performance Estimators", accepted to AAAI-23.
2 papers · 0 benchmarks
UPFD-POL (User Preference-aware Fake News Detection)
The PolitiFact variant of the UPFD dataset for benchmarking.
2 papers · 1 benchmark
The raw data are obtained from an industrial plant for ultra-processed food production.
2 papers · 0 benchmarks
Wyze Rule Recommendation Dataset.
2 papers · 0 benchmarks
Source
2 papers · 0 benchmarks
Dataset of low fidelity resolutions of the RANS equations over airfoils.
1 paper · 0 benchmarks
This paper constructs 7-digit product Supply-Use Tables (SUTs) and symmetric Input-Output Tables (IOTs) for the Indian economy using microdata from the Annual Survey of Industries (ASI) for the period 2016-2021.
1 paper · 0 benchmarks
The AIDS Antiviral Screen dataset is a dataset of screens checking tens of thousands of compounds for evidence of anti-HIV activity.
1 paper · 0 benchmarks
ATC-GRAPH is the most extensive ATC benchmark dataset.
1 paper · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
AutoFR Dataset is broken down by each site that we crawl within a zip file.
1 paper · 0 benchmarks
B-XAIC consists of 50K small molecules represented as graphs and includes 7 graph classification tasks, each with ground truth labels and corresponding explanations.
1 paper · 0 benchmarks
BTS (Building Timeseries Dataset: Empowering Large-Scale Building Analytics)
The Building TimeSeries (BTS) dataset covers three buildings over a three-year period, comprising more than ten thousand timeseries data points with hundreds of unique ontologies.
1 paper · 0 benchmarks
BeGin provides 23 benchmark scenarios for graph from 14 real-world datasets, which cover 12 combinations of the incremental settings and the levels of problem.
1 paper · 0 benchmarks
The original paper contains a high-level explanation of the dataset characteristics, and potential use cases of the dataset.
1 paper · 0 benchmarks
Description This repository includes the experiment results, source code, and test data for Three Cs risk inference, using the CIRO (COVID-19 Infection Risk Ontology) and HermiT.
1 paper · 0 benchmarks
CTFW is a large annotated procedural text dataset in the cybersecurity domain (3154 documents).
1 paper · 0 benchmarks
Classifying all cells in an organ is a relevant and difficult problem from plant developmental biology.
1 paper · 1 benchmark
Clickable heat-map visualizations of the experiments run to quantify the Classic ECN AQM problem and to evaluate the success of the Classic AQM Detection and Fall-back algorithm.
1 paper · 0 benchmarks
Synthetic graph classification datasets with the task of recognizing the connectivity of same-colored nodes in 4 graphs of varying topology.
1 paper · 0 benchmarks
CompMix-IR Dataset Overview: Characteristics: CompMix-IR is a heterogeneous knowledge retrieval benchmark dataset, featuring four knowledge types (text, knowledge graphs, tables, and infoboxes), 9,400+ QA pairs, and a corpus of 10 million…
1 paper · 0 benchmarks
DPB-5L is a Multilingual KG dataset containing 5 KGs in English, French, Japanese, Greek, and Spanish.
1 paper · 0 benchmarks
[comment]:<> (Data for the paper "Deciphering Environmental Air Pollution with Large Scale City Data") Main Dataset citypollutiondata.csv Relevant Columns: Date: Date of the sample City: City of the sample Xmedian: Median value of the…
1 paper · 0 benchmarks
Source: Linking Datasets on Organizations Using Half-a-Billion Open-Collaborated Records (Description (Markdown and LATEX enabled)) High-Level Explanation of the Dataset - Scale and Composition: This repository provides millions of…
1 paper · 0 benchmarks
This repository contains three graph datasets for the UE traffic assignment problem on Sioux-Falls, Eastern-Massachusetts and Anaheim networks in both dgl and pyg formats.
1 paper · 0 benchmarks
The FB1.5M dataset is a benchmark for Knowledge Graph Completion.
1 paper · 0 benchmarks
FinDKG: The Global Financial Dynamic Knowledge Graph Dataset FinDKG is an open-source dataset focused on creating a temporally-resolved Financial Dynamic Knowledge Graph.
1 paper · 0 benchmarks
This repository is an extension of GEval.
1 paper · 0 benchmarks
GO21 is a biomedical knowledge graph that models genes, proteins, drugs, and the hierarchy of the biological processes they participate in.
1 paper · 1 benchmark
Genre2Movies (Compositional queries for Movie recommendation)
Genre annotations for movies The file genre2movies.csv contains genre-movie tuples based on Wikidata annotations (https://www.wikidata.org/).
1 paper · 0 benchmarks
GeoJEPAD (GeoJEPA Dataset)
GeoJEPAD is a multimodal dataset combining OpenStreetMap (OSM) data (attributes and geometries) with high-resolution aerial imagery from diverse urban areas.
1 paper · 0 benchmarks
Dataset introduced by Xifeng Yan et al.
1 paper · 0 benchmarks
Dataset introduced by Xifeng Yan et al.
1 paper · 0 benchmarks
HALvest-Geometric is a subset of HALvest: an academic citation network with 238,397 disambiguated authors and 18,662,037 scholarly papers.
1 paper · 0 benchmarks
HAM (Human-annotated Mappings)
HAM is a dataset for molecular graph partitioning.
1 paper · 0 benchmarks
HTDM (Hypertention Disease Medication)
Hypertention Disease Medication dataset.
1 paper · 0 benchmarks
Multi-Modal Hate Speech Detection with Graph Context.
1 paper · 0 benchmarks
Inpatient claims, Outpatient claims and Beneficiary details of each provider.
1 paper · 1 benchmark
HoaxItaly consists of over 1 million tweets shared during 2019 and containing links to thousands of news articles published on two classes of Italian outlets: (1) disinformation websites, i.e.
1 paper · 0 benchmarks
A large dataset from the Inductive Link Prediction Challenge 2022.
1 paper · 1 benchmark
A small dataset from the Inductive Link Prediction Challenge 2022.
1 paper · 1 benchmark
IMCPT-SparseGM dataset is a new visual graph matching benchmark addressing partial matching and graphs with larger sizes, based on the novel stereo benchmark Image Matching Challenge PhotoTourism (IMC-PT) 2020.
1 paper · 1 benchmark
IMCPT-SparseGM dataset is a new visual graph matching benchmark addressing partial matching and graphs with larger sizes, based on the novel stereo benchmark Image Matching Challenge PhotoTourism (IMC-PT) 2020.
1 paper · 1 benchmark
We release 280 synthetic IAM graphs generated using IAM graphs of commercial companies.
1 paper · 0 benchmarks
JoCAD is a dataset for anomaly detection in citation networks.
1 paper · 0 benchmarks
The KACC benchmark consists of three subtasks that can be applied to knowledge graphs: knowledge abstraction, knowledge concretization and knowledge completion.
1 paper · 0 benchmarks
LSEC (Live Stream E-Commerce)
The LSEC (Live Stream E-Commerce) dataset has two subsets: LSEC-Small and LSEC-Large.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.