Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 64 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 3025–3072 of 12,172
DAWN emphasizes a diverse traffic environment (urban, highway and freeway) as well as a rich variety of traffic flow.
10 papers · 0 benchmarks
DCASE2018 Task 4 is a dataset for large-scale weakly labeled semi-supervised sound event detection in domestic environments.
10 papers · 0 benchmarks
DDPM (Deception Detection and Physiological Monitoring)
The Deception Detection and Physiological Monitoring (DDPM) dataset captures an interview scenario in which the interviewee attempts to deceive the interviewer on selected responses.
10 papers · 0 benchmarks
DOTA 2.0 (Dataset of Object deTection in Aerial images)
—In the past decade, object detection has achieved significant progress in natural images but not in aerial images, due to the massive variations in the scale and orientation of objects caused by the bird’s-eye view of aerial images.
10 papers · 0 benchmarks
DeepScores contains high quality images of musical scores, partitioned into 300,000 sheets of written music that contain symbols of different shapes and sizes.
10 papers · 0 benchmarks
A new dataset of handwritten text with fine-grained annotations at the character level and report results from an initial user evaluation.
10 papers · 0 benchmarks
Although deep face recognition has achieved impressive results in recent years, there is increasing controversy regarding racial and gender bias of the models, questioning their trustworthiness and deployment into sensitive scenarios.
10 papers · 0 benchmarks
Description Detection Dataset (D³, /dikju:b/) is an attempt at creating a next-generation object detection dataset.
10 papers · 1 benchmark
DiDi (Distractor Distilled Dataset)
DiDi is a distractor-distilled tracking dataset created to address the limitation of low distractor presence in current visual object tracking benchmarks.
10 papers · 1 benchmark
DiscoFuse was created by applying a rule-based splitting method on two corpora - sports articles crawled from the Web, and Wikipedia.
10 papers · 0 benchmarks
The Discovery datasets consists of adjacent sentence pairs (s1,s2) with a discourse marker (y) that occurred at the beginning of s2.
10 papers · 1 benchmark
This is the dataset for the 2020 Duolingo shared task on Simultaneous Translation And Paraphrase for Language Education (STAPLE).
10 papers · 0 benchmarks
Dynamic Replica is a synthetic dataset of stereo videos featuring humans and animals in virtual environments.
10 papers · 0 benchmarks
This dataset is composed of two collections of heartbeat signals derived from two famous PhysioNet datasets in heartbeat classification, the MIT-BIH Arrhythmia Dataset and the PTB Diagnostic ECG Database.
10 papers · 0 benchmarks
EEEyeNet is a dataset and benchmark with the goal of advancing research in the intersection of brain activities and eye movements.
10 papers · 0 benchmarks
Earnings-22 is a practical benchmark designed to evaluate automatic speech recognition (ASR) systems' performance on real-world, accented audio.
10 papers · 0 benchmarks
Satellite images are snapshots of the Earth surface.
10 papers · 4 benchmarks
EmoWOZ is the first large-scale open-source dataset for emotion recognition in task-oriented dialogues.
10 papers · 2 benchmarks
FERET-Morphs is a dataset of morphed faces selected from the publicly available FERET dataset [1].
10 papers · 0 benchmarks
FM-IQA (Freestyle Multilingual Image Question Answering)
FM-IQA is a question-answering dataset containing over 150,000 images and 310,000 freestyle Chinese question-answer pairs and their English translations.
10 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
10 papers · 0 benchmarks
Fig-QA consists of 10256 examples of human-written creative metaphors that are paired as a Winograd schema.
10 papers · 0 benchmarks
The dataset was created using high-resolution (8 m) satellite imagery from the Gaofen series (Gaofen-2 and Gaofen-6), captured in 2019 over Maduo County, China, located in the Yellow River source area.
10 papers · 1 benchmark
FunnyBirds is a synthetic vision dataset that is developed to automatically and quantitatively analyze XAI methods.
10 papers · 0 benchmarks
GDA (Gene-Disease Associations Corpus)
The gene-disease associations corpus contains 30,192 titles and abstracts from PubMed articles that have been automatically labelled for genes, diseases and gene-disease associations via distant supervision.
10 papers · 2 benchmarks
GRAZPEDWRI-DX is a public dataset of 20,327 pediatric wrist trauma X-ray images released by the University of Medicine of Graz.
10 papers · 4 benchmarks
The QMUL underGround Re-IDentification (GRID) dataset contains 250 pedestrian image pairs.
10 papers · 5 benchmarks
GUG (Grammatical” versus “UnGrammatical)
See article for detail
10 papers · 0 benchmarks
GenWiki is a large-scale dataset for knowledge graph-to-text (G2T) and text-to-knowledge graph (T2G) conversion.
10 papers · 3 benchmarks
The GraphInstruct dataset is part of a benchmark proposed in the paper titled "GraphInstruct: Empowering Large Language Models with Graph Understanding and Reasoning Capability." This benchmark is designed to evaluate and enhance the graph…
10 papers · 0 benchmarks
Griddly is an environment for grid-world based research.
10 papers · 0 benchmarks
The HO-3D v3 is the version 3 of the HO-3D dataset with more accurate hand-object poses.
10 papers · 1 benchmark
A new challenge domain with novel problems that arise from its combination of purely cooperative gameplay with two to five players and imperfect information.
10 papers · 0 benchmarks
HolStep is a dataset based on higher-order logic (HOL) proofs, for the purpose of developing new machine learning-based theorem-proving strategies.
10 papers · 2 benchmarks
ImageNet-VidVRD dataset contains 1,000 videos selected from ILVSRC2016-VID dataset based on whether the video contains clear visual relations.
10 papers · 2 benchmarks
ImgEdit is a large-scale, high-quality image-editing dataset comprising 1.2 million carefully curated edit pairs, which contain both novel and complex single-turn edits, as well as challenging multi-turn tasks.
10 papers · 1 benchmark
Dataset Introduction In this work, we introduce the In-Diagram Logic (InDL) dataset, an innovative resource crafted to rigorously evaluate the logic interpretation abilities of deep learning models.
10 papers · 1 benchmark
This relational database consists of 24 unique names in two families (they have equivalent structures).
10 papers · 0 benchmarks
Kuzushiji-49 is an MNIST-like dataset that has 49 classes (28x28 grayscale, 270,912 images) from 48 Hiragana characters and one Hiragana iteration mark.
10 papers · 0 benchmarks
LAD (Large-scale Attribute Dataset)
LAD (Large-scale Attribute Dataset) has 78,017 images of 5 super-classes and 230 classes.
10 papers · 0 benchmarks
A challenging new benchmark for language-agnostic answer retrieval from a multilingual candidate pool.
10 papers · 0 benchmarks
LC-QuAD (Largescale Complex Question Answering Dataset)
LC-QuAD is a Large Question Answering dataset with 30,000 pairs of questions and its corresponding SPARQL query.
10 papers · 1 benchmark
These are the official datasets for the LHC Olympics 2020 Anomaly Detection Challenge.
10 papers · 0 benchmarks
The Lakh Pianoroll Dataset (LPD) is a collection of 174,154 multitrack pianorolls derived from the Lakh MIDI Dataset (LMD).
10 papers · 0 benchmarks
LectureBank Dataset is a manually collected dataset of lecture slides.
10 papers · 0 benchmarks
MACS (Multi-Annotator Captioned Soundscapes)
This is a dataset containing audio captions and corresponding audio tags for a number of 3930 audio files of the TAU Urban Acoustic Scenes 2019 development dataset (airport, public square, and park).
10 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.