Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 42 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 1969–2016 of 12,172
XLCoST (Cross-Lingual Code Snippet)
XLCoST is a benchmark dataset for cross-lingual code intelligence.
22 papers · 0 benchmarks
iPinYou (iPinYou Global RTB Bidding Algorithm Competition Dataset)
The iPinYou Global RTB(Real-Time Bidding) Bidding Algorithm Competition is organized by iPinYou from April 1st, 2013 to December 31st, 2013.The competition has been divided into three seasons.
22 papers · 1 benchmark
iSarcasmEval is the first shared task to target intended sarcasm detection: the data for this task was provided and labelled by the authors of the texts themselves.
22 papers · 0 benchmarks
iVQA (Instructional Video Question Answering)
An open-ended VideoQA benchmark that aims to: i) provide a well-defined evaluation by including five correct answer annotations per question and ii) avoid questions which can be answered without the video.
22 papers · 2 benchmarks
The examined group comprised kernels belonging to three different varieties of wheat: Kama, Rosa and Canadian, 70 elements each, randomly selected for the experiment.
22 papers · 1 benchmark
AIST++ is a 3D dance dataset which contains 3D motion reconstructed from real dancers paired with music.
21 papers · 2 benchmarks
Animal-Pose Dataset is an animal pose dataset to facilitate training and evaluation.
21 papers · 1 benchmark
CAL500 (Computer Audition Lab 500)
CAL500 (Computer Audition Lab 500) is a dataset aimed for evaluation of music information retrieval systems.
21 papers · 0 benchmarks
CDD-11 (Composite Degradation Dataset 11)
An image restoration dataset
21 papers · 1 benchmark
COCO-CN is a bilingual image description dataset enriching MS-COCO with manually written Chinese sentences and tags.
21 papers · 1 benchmark
Consists of 20k English biomedical entity mentions from Reddit expert-annotated with links to SNOMED CT, a widely-used medical knowledge graph.
21 papers · 0 benchmarks
COVID-Fact is a FEVER-like dataset of claims concerning the COVID-19 pandemic.
21 papers · 0 benchmarks
CliCR is a new dataset for domain specific reading comprehension used to construct around 100,000 cloze queries from clinical case reports.
21 papers · 1 benchmark
ContactDB is a dataset of contact maps for household objects that captures the rich hand-object contact that occurs during grasping, enabled by use of a thermal camera.
21 papers · 1 benchmark
A team of researchers from Qatar University, Doha, Qatar, and the University of Dhaka, Bangladesh along with their collaborators from Pakistan and Malaysia in collaboration with medical doctors have created a database of chest X-ray images…
21 papers · 0 benchmarks
DigiFace-1M is a synthetic dataset for face recognition, obtained by rendering digital faces using a computer graphics pipeline.
21 papers · 0 benchmarks
Digits (Optical Recognition of Handwritten Digits)
The DIGITS dataset consists of 1797 8×8 grayscale images (1439 for training and 360 for testing) of handwritten digits.
21 papers · 3 benchmarks
EPHOIE (phtnsantader@gmail.com)
EPHOIE is a fully-annotated dataset which is the first Chinese benchmark for both text spotting and visual information extraction.
21 papers · 2 benchmarks
Caenorhabditis elegans is a roundworm commonly used as a model organism in the study of genetics.
21 papers · 1 benchmark
Event2Mind is a corpus of 25,000 event phrases covering a diverse range of everyday events and situations.
21 papers · 2 benchmarks
Encourages machine learning research in this area and to help facilitate further work in understanding and mitigating the effects of climate change.
21 papers · 0 benchmarks
FIW (Families In The Wild)
FIW is a large and comprehensive database available for kinship recognition.
21 papers · 0 benchmarks
FQuAD (French Question Answering Dataset)
A French Native Reading Comprehension dataset of questions and answers on a set of Wikipedia articles that consists of 25,000+ samples for the 1.0 version and 60,000+ samples for the 1.1 version.
21 papers · 1 benchmark
Two datasets are provided.
21 papers · 0 benchmarks
The Respiratory Sound database was originally compiled to support the scientific challenge organized at Int.
21 papers · 2 benchmarks
The INRIA Aerial Image Labeling dataset is comprised of 360 RGB tiles of 5000×5000px with a spatial resolution of 30cm/px on 10 cities across the globe.
21 papers · 1 benchmark
A benchmark dataset for out-of-distribution detection.
21 papers · 1 benchmark
JAAD (Joint Attention in Autonomous Driving)
JAAD is a dataset for studying joint attention in the context of autonomous driving.
21 papers · 1 benchmark
JARVIS-DFT is a repository of density functional theory based calculation data for materials.
21 papers · 1 benchmark
KLUE (Korean Language Understanding Evaluation)
Korean Language Understanding Evaluation (KLUE) benchmark is a series of datasets to evaluate natural language understanding capability of Korean language models.
21 papers · 1 benchmark
LAG (Large-scale Attention based Glaucoma)
Includes 5,824 fundus images labeled with either positive glaucoma (2,392) or negative glaucoma (3,432).
21 papers · 1 benchmark
A social network of LastFM users which was collected from the public API in March 2020.
21 papers · 0 benchmarks
LinCE (Linguistic Code-switching Evaluation Dataset)
A centralized benchmark for Linguistic Code-switching Evaluation (LinCE) that combines ten corpora covering four different code-switched language pairs (i.e., Spanish-English, Nepali-English, Hindi-English, and Modern Standard…
21 papers · 0 benchmarks
MIR-1K (Multimedia Information Retrieval lab, 1000 song clips) is a dataset designed for singing voice separation.
21 papers · 0 benchmarks
MMT-Bench is a comprehensive benchmark designed to evaluate Large Vision-Language Models (LVLMs) across a wide array of multimodal tasks that require expert knowledge as well as deliberate visual recognition, localization, reasoning, and…
21 papers · 0 benchmarks
A composite dataset that unifies semantic segmentation datasets from different domains.
21 papers · 0 benchmarks
Publicly available dataset of naturally occurring factual claims for the purpose of automatic claim verification.
21 papers · 0 benchmarks
PLOS (Scientific Lay Summarization)
This dataset contains 27,525 full biomedical articles paired with non-technical lay summaries derived from various journals published by the Public Library of Science (PLOS).
21 papers · 2 benchmarks
SEP-28k (Stuttering Events in Podcasts)
Stuttering Events in Podcasts (SEP-28k) is a dataset containing over 28k clips labeled with five event types including blocks, prolongations, sound repetitions, word repetitions, and interjections.
21 papers · 0 benchmarks
SICAPv2 is a database containing prostate histology whole slide images with both annotations of global Gleason scores and path-level Gleason grades.
21 papers · 0 benchmarks
SMID (Seeing motion in the dark)
This is the low-light image enhancement dataset collected by the CVPR 2018 paper "Seeing Motion in the Dark".
21 papers · 1 benchmark
STARSS23 (STARSS23: An Audio-Visual Dataset of Spatial Recordings of Real Scenes with Spatiotemporal Annotations of Sound Events)
The Sony-TAu Realistic Spatial Soundscapes 2023 (STARSS23) dataset contains multichannel recordings of sound scenes in various rooms and environments, together with temporal and spatial annotations of prominent events belonging to a set of…
21 papers · 0 benchmarks
STREUSLE stands for Supersense-Tagged Repository of English with a Unified Semantics for Lexical Expressions.
21 papers · 1 benchmark
The Terms of Service dataset is a law dataset corresponding to the task of identifying whether contractual terms are potentially unfair.
21 papers · 1 benchmark
The UT-Interaction dataset contains videos of continuous executions of 6 classes of human-human interactions: shake-hands, point, hug, push, kick and punch.
21 papers · 1 benchmark
WIQA (What-If Question Answering)
The WIQA dataset V1 has 39705 questions containing a perturbation and a possible effect in the context of a paragraph.
21 papers · 0 benchmarks
Contains one million naturally occurring sentence rewrites, providing sixty times more distinct split examples and a ninety times larger vocabulary than the WebSplit corpus introduced by Narayan et al.
21 papers · 0 benchmarks
The 20 Newsgroups data set is a collection of approximately 20,000 newsgroup documents, partitioned (nearly) evenly across 20 different newsgroups.
20 papers · 3 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.