Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 240 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 11473–11520 of 12,172
DR HAGIS (Diabetic Retinopathy, Hypertension, Age-related macular degeneration and Glacuoma ImageS)
The DR HAGIS database has been created to aid the development of vessel extraction algorithms suitable for retinal screening programmes.
0 papers · 0 benchmarks
Dataset for the DREAMING - Diminished Reality for Emerging Applications in Medicine through Inpainting Challenge!
0 papers · 0 benchmarks
Dialog System Technology Challenges 8 (DSTC) Track 2 builds on the success of DSTC 7 Track 1 (NOESIS: Noetic End-to-End Response Selection Challenge).
0 papers · 0 benchmarks
The DUC 2005 data set is a dataset for summarization which consists of 50 document collections of 25 documents each; each document collection includes a human-written query.
0 papers · 0 benchmarks
DUS (Daimler Urban Segmentation)
The Daimler Urban Segmentation Dataset is a dataset for semantic segmentation.
0 papers · 0 benchmarks
This is a dataset of 18624 fonts labeled as 100% Free and Public domain / GPL / OFL on https://www.dafont.com/ with .ttf and .otf extensions.
0 papers · 0 benchmarks
The DanishPoliticalComments dataset is a collection of sentences that are labeled with fine-grained polarity in the range from -2 to 2 (negative to positive).
0 papers · 0 benchmarks
This dataset is part of the Data Wrangling Dataset Repository created by the DMiP Team (UPV).
0 papers · 0 benchmarks
Cancer genomics and precision oncology: The TCGA Research Network started in 2005 has profiled and analyzed a large number of human tumors to discover molecular aberrations at the DNA, RNA, protein, and epigenetic levels and thereby…
0 papers · 0 benchmarks
The landmark Cancer Genomics Program launched in 2006 has contributed immensely to the awareness of the importance of cancer genomics in our understanding of cancer over the past decade and has begun to change the way the disease is…
0 papers · 0 benchmarks
The dataset for the Deenz Emotional Abuse Scale (DEAS-18) study comprises detailed participant responses collected to validate the psychometric properties of the scale.
0 papers · 0 benchmarks
Dataset contains 16.000 electric power distribution transformers from Cauca Department (Colombia).
0 papers · 0 benchmarks
The DeepSpeak dataset contains over 43 hours of real and deepfake footage of people talking and gesturing in front of their webcams.
0 papers · 0 benchmarks
Deeply Korean read speech corpus contains pairs of Korean speakers reading a script with 3 distinct text sentiments (negative, neutral, positive), with 3 distinct voice sentiments (negative, neutral, positive), are recorded.
0 papers · 0 benchmarks
Deeply Parent-Child Vocal Interaction contains the interaction of 24 pairs of parent and child(total 48 speakers), such as reading fairy tales, singing children’s songs, conversing, and others, is recorded.
0 papers · 0 benchmarks
Deeply vocal characterizer is a human nonverbal vocalization dataset.
0 papers · 0 benchmarks
DementiaBank is a shared database of multimedia interactions for the study of communication in dementia.
0 papers · 0 benchmarks
DemonsP300 MOABB (Visual P300 dataset recorded in Virtual Reality (VR) game Raccoons versus Demons.)
0 papers · 0 benchmarks
Benchmark dataset for low-resource multiclass classification, with 4,015 training, 500 testing, and 500 validation examples, each labeled as part of five classes.
0 papers · 0 benchmarks
DetReIDX (A Stress-Test Dataset for Real-World UAV-Based Person Recognition)
We are proud to introduce DetReIDX - is a new benchmark dataset built for real-world, long-range human recognition.
0 papers · 0 benchmarks
The DevOps-Eval is an industrial-first evaluation benchmark specifically designed for Large Language Models (LLMs) in the DevOps/AIOps domain¹.
0 papers · 0 benchmarks
RGB-D images of 60 western dishes, home made.
0 papers · 0 benchmarks
This Dataset is the official dataset accompanying the paper "DiffCkt: A Diffusion Model-Based Hybrid Neural Network Framework for Automatic Transistor-Level Generation of Analog Circuits".
0 papers · 0 benchmarks
DigiLeTs (Digit- and Letter Trajectories)
A dataset with 23 870 digital trajectories (i.e.
0 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
0 papers · 0 benchmarks
The deliberate manipulation of public opinion, especially through altered images, poses a significant danger to society.
0 papers · 0 benchmarks
This dataset consists of images of various hand and power tools, specifically designed to aid in training and improving AI-based object recognition systems.
0 papers · 0 benchmarks
The DogCentric Activity dataset is composed of dog activity videos taken from a first-person animal viewpoint.
0 papers · 0 benchmarks
A dog face dataset for dog face verification and recognition/identification.
0 papers · 0 benchmarks
This dataset is collected by Datacluster Labs.
0 papers · 0 benchmarks
a large video dataset captured with UAVs in different complex real-world scenes, with multiple representations, suitable for multi-task learning.
0 papers · 0 benchmarks
The Dutch Social Dataset is a collection of tweets primarily written in Dutch.
0 papers · 0 benchmarks
This dataset is based on FB15k237 and a pre-trained language-model-based KGE.
0 papers · 0 benchmarks
This dataset contains electroencephalogram (EEG) signals, event-related potentials (ERP), and demographic attributes aimed at the early identification of schizophrenia.
0 papers · 0 benchmarks
Context As mentioned in the reference paper: Dust storms are considered a severe meteorological disaster, especially in arid and semi-arid regions, which is characterized by dust aerosol-filled air and strong winds across an extensive area.
0 papers · 0 benchmarks
The MICCAI 2020 EMIDEC dataset is a dataset for classifying normal and pathological cases from the clinical information with or without DE-MRI, and secondly to automatically detect the different relevant areas (the myocardial contours, the…
0 papers · 0 benchmarks
ENEM (Brazilian High School National Exam)
The ENEM dataset refers to data collected from the Brazilian High School National Exam (ENEM).
0 papers · 0 benchmarks
ESP dataset (Evaluation for Styled Prompt dataset) is a new benchmark for zero-shot domain-conditional caption generation.
0 papers · 0 benchmarks
ESPADA (Extended Synthetic and Photogrammetric Aerial-Image Dataset)
We present a new aerial image dataset, named ESPADA, intended for the training of deep neural networks for depth image estimation from a single aerial image.
0 papers · 0 benchmarks
The EU-ADR corpus is a biomedical relation extraction dataset that contains 100 abstracts, with relations between drug, disorder, and targets.
0 papers · 0 benchmarks
The Edge Milling Heads data set comprises 144 images of an edge profile cutting head of a milling machine.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.