Home › Datasets › task › Graph Regression
Graph Regression datasets
archive 2025-07-28
21 datasets carry the task tag "Graph Regression" (the task itself: Graph Regression), ordered by the archive's paper count. Page 1 of 1: 21 shown of 21. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Graph Regression datasets 1–21 of 21
ZINC is a free database of commercially-available compounds for virtual screening.
251 papers · 5 benchmarks
MoleculeNet is a large scale benchmark for molecular machine learning.
240 papers · 1 benchmark
The Long Range Graph Benchmark (LRGB) is a collection of 5 graph learning datasets that arguably require long-range reasoning to achieve strong performance in a given task.
76 papers · 4 benchmarks
QM9 provides quantum chemical properties (at DFT level) for a relevant, consistent, and comprehensive chemical space of small organic molecules.
76 papers · 9 benchmarks
OGB-LSC (OGB Large-Scale Challenge)
OGB Large-Scale Challenge (OGB-LSC) is a collection of three real-world datasets for advancing the state-of-the-art in large-scale graph ML.
34 papers · 3 benchmarks
The Tox21 data set comprises 12,060 training samples and 647 test samples that represent chemical compounds.
30 papers · 4 benchmarks
ESOL is a water solubility prediction dataset consisting of 1128 samples.
25 papers · 4 benchmarks
Using conservation of energy -- a fundamental property of closed classical and quantum mechanical systems -- we develop an efficient gradient-domain machine learning (GDML) approach to construct accurate molecular force fields using a…
20 papers · 0 benchmarks
Regression dataset for molecular docking scores (predicted molecule-protein binding affinity).
18 papers · 4 benchmarks
PCQM4Mv2 is a quantum chemistry dataset originally curated under the PubChemQC project.
17 papers · 1 benchmark
The $O2$Perm dataset is created from the Membrane Society of Australasia portal.
4 papers · 0 benchmarks
DrivAerNet (A Parametric Car Dataset for Data-driven Aerodynamic Design and Graph-Based Drag Prediction)
DrivAerNet is a large-scale, high-fidelity CFD dataset of 3D industry-standard car shapes designed for data-driven aerodynamic design.
4 papers · 0 benchmarks
The GlassTemp dataset is collected from Polyinfo.
3 papers · 1 benchmark
The MeltingTemp dataset is collected from Polyinfo.
3 papers · 0 benchmarks
SupplyGraph (SupplyGraph: A Benchmark Dataset for Supply Chain Planning using Graph Neural Networks)
Graph Neural Networks (GNNs) have gained traction across different domains such as transportation, bio-informatics, language processing, and computer vision.
3 papers · 0 benchmarks
Cylinder in Crossflow is a synthetic dataset that involves unsteady laminar flow past a cylinder that generates vortex shedding pattern known as a von Kármán vortex street.
2 papers · 0 benchmarks
GVLQA (Graph Vision-Language Question-Answering)
GVLQA is the first vision-language QA dataset for general graph reasoning.
2 papers · 0 benchmarks
The PolyDensity is collected from Polyinfo.
2 papers · 0 benchmarks
OCB (Open Circuit Benchmark)
OCB contains two graph datasets, Ckt-Bench-101 and Ckt-Bench-301, for representation learning over analog circuits.
1 paper · 0 benchmarks
This dataset are about Nafion 112 membrane standard tests and MEA activation tests of PEM fuel cell in various operation condition.
1 paper · 0 benchmarks
hERG is a large-scale biophysics federated molecular dataset related to cardiac toxicity.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.