Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 68 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 3217–3264 of 12,172
FinRED is a relation extraction dataset curated from financial news and earning call transcripts containing relations from the finance domain.
9 papers · 0 benchmarks
A popular dataset for node classification on heterogeneous graphs.
9 papers · 1 benchmark
A large publicly available retinal fundus image dataset for glaucoma classification called G1020.
9 papers · 0 benchmarks
GLOBEM is a multi-year passive sensing datasets, containing over 700 user-years and 497 unique users' data collected from mobile and wearable sensors, together with a wide range of well-being metrics.
9 papers · 0 benchmarks
GeoWebNews provides test/train examples and enable fine-grained Geotagging and Toponym Resolution (Geocoding).
9 papers · 0 benchmarks
GiantMIDI-Piano contains 10,854 unique piano solo pieces composed by 2,786 composers.
9 papers · 0 benchmarks
Global Voices is a multilingual dataset for evaluating cross-lingual summarization methods.
9 papers · 0 benchmarks
Global WHEAT Dataset is the first large-scale dataset for wheat head detection from field optical images.
9 papers · 0 benchmarks
This is a gun detection dataset with 51K annotated gun images for gun detection and other 51K cropped gun chip images for gun classification collected from a few different sources.
9 papers · 6 benchmarks
A new dataset of annotated pages from books of hours, a type of handwritten prayer books owned and used by rich lay people in the late middle ages.
9 papers · 0 benchmarks
HQ-WMCA (High-Quality Wide Multi-Channel Attack database)
The High-Quality Wide Multi-Channel Attack database (HQ-WMCA) database consists of 2904 short multi-modal video recordings of both bona-fide and presentation attacks.
9 papers · 0 benchmarks
HUMAN4D is a large and multimodal 4D dataset that contains a variety of human activities simultaneously captured by a professional marker-based MoCap, a volumetric capture and an audio recording system.
9 papers · 0 benchmarks
HappyDB is a corpus of 100,000 crowdsourced happy moments.
9 papers · 0 benchmarks
Paper | Github | Dataset| Model As a part of our research efforts toward making LLMs more safe for public use, we create HarmfulQA i.e.
9 papers · 1 benchmark
Home Action Genome is a large-scale multi-view video database of indoor daily activities.
9 papers · 2 benchmarks
Amazon Mechanical Turk (AMT) is used to collect annotations on HowTo100M videos.
9 papers · 0 benchmarks
HuGaDB is human gait data collection for analysis and activity recognition consisting of continues recordings of combined activities, such as walking, running, taking stairs up and down, sitting down, and so on; and the data recorded are…
9 papers · 0 benchmarks
We provide six different datasets with diverse range of activities
9 papers · 0 benchmarks
Colorization validation set for unconditional/conditional colorization tasks.
9 papers · 2 benchmarks
Imagewoof is a subset of 10 dog breed classes from Imagenet.
9 papers · 0 benchmarks
The IndoSum dataset is a benchmark dataset for Indonesian text summarization.
9 papers · 0 benchmarks
A multimodal dataset with comprehensive annotations of continuous emotions during naturalistic conversations.
9 papers · 0 benchmarks
KSoF (The Kassel State of Fluency Dataset – A Therapy Centered Dataset of Stuttering)
Stuttering is a complex speech disorder that negatively affects an individual’s ability to communicate effectively.
9 papers · 0 benchmarks
KaMed is a knowledge-aware medical dialogue dataset, which contains over 60,000 medical dialogue sessions with 5,682 entities (such as Asthma and Atropine).
9 papers · 0 benchmarks
We describe the 2020 edition of the DeepMind Kinetics human action dataset, which replenishes and extends the Kinetics-700 dataset.
9 papers · 1 benchmark
Dataset for document shadow removal
9 papers · 0 benchmarks
KnowIT VQA is a video dataset with 24,282 human-generated question-answer pairs about The Big Bang Theory.
9 papers · 0 benchmarks
LDC2020T02 (Abstract Meaning Representation (AMR) Annotation Release 3.0)
Abstract Meaning Representation (AMR) Annotation Release 3.0 was developed by the Linguistic Data Consortium (LDC), SDL/Language Weaver, Inc., the University of Colorado's Computational Language and Educational Research group and the…
9 papers · 1 benchmark
LEAF-QA, a comprehensive dataset of 250,000 densely annotated figures/charts, constructed from real-world open data sources, along with ~2 million question-answer (QA) pairs querying the structure and semantics of these charts.
9 papers · 0 benchmarks
For LIAR-RAW, we extended the public dataset LIAR-PLUS (Alhindi et al., 2018) with relevant raw reports, containing fine-grained claims from Politifact.
9 papers · 0 benchmarks
To make synthetic images match the property of real dark photography, we analyze the illumination distribution of low-light images.
9 papers · 1 benchmark
LoDoPaB-CT is a dataset of computed tomography images and simulated low-dose measurements.
9 papers · 1 benchmark
A logo detection dataset with full annotation, which has 3,000 logo categories, about 200,000 manually annotated logo objects and 158,652 images.
9 papers · 0 benchmarks
MEF (Multi-exposure image fusion)
Multi-exposure image fusion (MEF) is considered an effective quality enhancement technique widely adopted in consumer electronics, but little work has been dedicated to the perceptual quality assessment of multi-exposure fused images.
9 papers · 1 benchmark
MEIR (Multimodal Entity Image Repurposing)
MEIR is a substantially challenging dataset over that which has been previously available to support research into image repurposing detection.
9 papers · 0 benchmarks
MFRC (Moral Foundations Reddit Corpus)
Moral Foundations Reddit Corpus (MFRC) is a collection of 16,123 Reddit comments that have been curated from 12 distinct subreddits, hand-annotated by at least three trained annotators for 8 categories of moral sentiment (i.e., Care,…
9 papers · 0 benchmarks
Context This is the Original data provided by MIT .
9 papers · 1 benchmark
MMCU (Measuring Massive Multitask Chinese Understanding)
We propose a test to measure the multitask accuracy of large Chinese language models.
9 papers · 0 benchmarks
MOD (Meme incorporated Open-domain Dialogue)
MOD is a large-scale open-domain multimodal dialogue dataset incorporating abundant Internet memes into utterances.
9 papers · 0 benchmarks
MSASL is a real-life large-scale sign language data set comprising over 25,000 annotated videos.
9 papers · 1 benchmark
The datasets are machine learning data, in which queries and urls are represented by IDs.
9 papers · 0 benchmarks
MSP-Podcast (A large naturalistic speech emotional dataset)
The MSP-Podcast corpus contains speech segments from podcast recordings which are perceptually annotated using crowdsourcing.
9 papers · 4 benchmarks
This is a dataset for a video inverse-tone-mapping task.
9 papers · 1 benchmark
The dataset contains single-shot videos taken from moving cameras in underwater environments.
9 papers · 1 benchmark
MedNLI (Medical Natural Language Inference)
The MedNLI dataset consists of the sentence pairs developed by Physicians from the Past Medical History section of MIMIC-III clinical notes annotated for Definitely True, Maybe True and Definitely False.
9 papers · 2 benchmarks
Middlebury 2005 is a stereo dataset of indoor scenes.
9 papers · 0 benchmarks
Mindboggle is a large publicly available dataset of manually labeled brain MRI.
9 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.