Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 133 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 6337–6384 of 12,172
GVLQA (Graph Vision-Language Question-Answering)
GVLQA is the first vision-language QA dataset for general graph reasoning.
2 papers · 0 benchmarks
Hierarchical-multilabel classification dataset for functional genomics
2 papers · 1 benchmark
Hierarchical-multilabel classification dataset for functional genomics
2 papers · 1 benchmark
Geoclidean-Elements dataset is derived from definitions in the first book of Euclid’s Elements, which focuses on plane geometry.
2 papers · 0 benchmarks
The GermEval dataset is a valuable resource for natural language processing (NLP) tasks, specifically named entity recognition (NER), conducted in the German language.
2 papers · 0 benchmarks
GermanDPR is a dataset for passage retrieval in German.
2 papers · 0 benchmarks
Global WHEAT Dataset 2021 is the extentions of the Global Wheat Dataset 2020.
2 papers · 0 benchmarks
A collection of 1000 public domain volumes that were scanned as part of the Google Book Search project.
2 papers · 0 benchmarks
The Graphine dataset contains 2,010,648 terminology definition pairs organized in 227 directed acyclic graphs.
2 papers · 0 benchmarks
A small and simple dataset featuring RGB-D images and heightmaps of various objects in a bin with manually annotated suctionable regions
2 papers · 0 benchmarks
The Gun Violence Corpus (GVC) consists of 241 unique incidents for which we have structured data on a) location, b) time c) the name, gender and age of the victims and d) the status of the victims after the incident: killed or injured.
2 papers · 0 benchmarks
Instrument playing technique (IPT) is a key element of musical presentation.
2 papers · 0 benchmarks
The H01 dataset is a 1.4 petabyte rendering of a small sample of human brain tissue, released by a collaboration between the Lichtman Laboratory at Harvard University and Google.
2 papers · 0 benchmarks
HAMMER dataset contains 13 Scenes.
2 papers · 0 benchmarks
Automated measurement of fetal head circumference using 2D ultrasound images
2 papers · 0 benchmarks
In order to fill the gap of HC3 under semanticinvariant tasks, we extend HC3 and propose a larger ChatGPT-generated text dataset covering translation, summarization, and paraphrasing tasks, called HC3 Plus.
2 papers · 0 benchmarks
HCP Aging (Lifespan Human Connectome Project Aging)
Lifespan HCP Release 2.0 includes cross-sectional visit 1 (V1) preprocessed structural and functional imaging data, unprocessed V1 imaging data for all included modalities (structural, high-res hippocampal T2, resting state fMRI, task…
2 papers · 1 benchmark
The Hochschule Darmstadt (HDA) facial tattoo and paintings database contains 500 pairs of facial images of individuals with and without facial tattoos or paintings.
2 papers · 0 benchmarks
HEMEW^S-3D (HEterogeneous Materials and Elastic Waves with Source variability in 3D)
The HEterogeneous Materials and Elastic Waves with Source variability in 3D (HEMEWS-3D) database comprises 30,000 simulations of elastic wave propagation in 3D geological domains.
2 papers · 0 benchmarks
This dataset contains simulated and expert-labelled spectrograms from two radio telescopes: the Hydrogen Epoch of Reionization Array (HERA) in South Africa and the Low-Frequency Array (LOFAR) in the Netherlands.
2 papers · 1 benchmark
HERDPhobia is an annotated hate speech detection dataset on Fulani herders in Nigeria -- in three languages: English, Nigerian-Pidgin, and Hausa.
2 papers · 0 benchmarks
HLGD (Headline Grouping Dataset)
The Headline Grouping dataset is a binary classification dataset on pairs of news headline.
2 papers · 0 benchmarks
The HOI-Synth benchmark extends three egocentric datasets designed to study hand-object interaction detection, EPIC-KITCHENS VISOR, EgoHOS, and ENIGMA-51, with automatically labeled synthetic data obtained through a novel HOI generation…
2 papers · 0 benchmarks
HRA (Human Rights Archive Database)
A verified-by-experts repository of 3050 human rights violations photographs, labelled with human rights semantic categories, comprising a list of the types of human rights abuses encountered at present.
2 papers · 0 benchmarks
An eyeblink detection in the wild dataset.
2 papers · 1 benchmark
HalluEditBench is a comprehensive benchmark for evaluating knowledge editing methods' effectiveness in correcting real-world hallucinations.
2 papers · 0 benchmarks
Hansel is a human-annotated Chinese entity linking (EL) dataset, focusing on tail entities and emerging entities: - The test set contains Few-shot (FS) and zero-shot (ZS) slices, has 10K examples and uses Wikidata as the corresponding…
2 papers · 0 benchmarks
This is a Twitter dataset of 100,386 users along with up to 200 tweets from their timelines with a random-walk-based crawler on the retweet graph, with a subsample of 4,972 which is manually annotated as hateful or not through…
2 papers · 0 benchmarks
This dataset contains panoramic video captured from a helmet-mounted camera while riding a bike through suburban Northern Virginia.
2 papers · 0 benchmarks
The Herbarium Half-Earth dataset is a large and diverse dataset of herbarium specimens to date for automatic taxon recognition.
2 papers · 1 benchmark
HeriGraph (Multimodal Machine Learning Datasets on Graphs of Heritage Values and Attributes)
The dataset contains constructed multi-modal features (visual and textual), pseudo-labels (on heritage values and attributes), and graph structures (with temporal, social, and spatial links) constructed using User-Generated Content data…
2 papers · 0 benchmarks
Heritage Provider Network is providing Competition Entrants with deidentified member data collected during a forty-eight month period that is allocated among three data sets (the "Data Sets").
2 papers · 0 benchmarks
HiAML Computational Graph (CG) family introduced in "GENNAPE: Towards Generalized Neural Architecture Performance Estimators", accepted to AAAI-23.
2 papers · 0 benchmarks
High-gamma dataset discribed in Schirrmeister et al.
2 papers · 0 benchmarks
This dataset is the Hindi version of standard English MSR-VTT dataset.
2 papers · 1 benchmark
Whereas the action recognition community has focused mostly on detecting simple actions like clapping, walking or jogging, the detection of fights or in general aggressive behaviors has been comparatively less studied.
2 papers · 1 benchmark
The Horne 2017 Fake News Data contains two independed news datasets: 1.
2 papers · 0 benchmarks
Hotel (Hospitality > Tourism > Hotel Demand/Sales)
The dataset contains the hotel demand and revenue of 8 major tourist destinations in the US (e.g., Los Angeles, Orlando ...).
2 papers · 0 benchmarks
HuTics (Human Deictic Gestures Dataset)
HuTics contains 2040 images showing how humans use deictic gestures to interact with various daily-life objects.
2 papers · 0 benchmarks
Timely and effective response to humanitarian crises requires quick and accurate analysis of large amounts of text data, a process that can highly benefit from expert-assisted NLP systems trained on validated and annotated data in the…
2 papers · 0 benchmarks
Human Simulacra is a virtual character dataset that contains 129k texts across 11 virtual characters, with each character having unique attributes, biographies, and stories.
2 papers · 0 benchmarks
The Human-Parts dataset is a dataset for human body, face and hand detection with ~15k images.
2 papers · 0 benchmarks
I.PHI processes the Packard Humanities Institute (PHI) database of ancient Greek inscriptions including the geographical and chronological metadata into a machine actionable format.
2 papers · 1 benchmark
I2-2000FPS is the first high-speed video dataset offering an unprecedented temporal resolution of 2000 frames per second (fps).
2 papers · 0 benchmarks
IACC.3 (Internet Archive videos (IACC.3) under Creative Commons licenses.)
The IACC.3 dataset is approximately 4600 Internet Archive videos (144 GB, 600 h) with Creative Commons licenses in MPEG-4/H.264 format with duration ranging from 6.5 min to 9.5 min and a mean duration of almost 7.8 min.
2 papers · 0 benchmarks
A Simulated Benchmark for multi-modal SLAM Systems Evaluation in Large-scale Dynamic Environments.
2 papers · 0 benchmarks
Table is a compact and efficient form for summarizing and presenting correlative information in handwritten and printed archival documents, scientific journals, reports, financial statements and so on.
2 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.