Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 111 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 5281–5328 of 12,172
DFDM (Deepfake videos generated from different models)
We created a new dataset, named DFDM, with 6,450 Deepfake videos generated by different Autoencoder models.
3 papers · 0 benchmarks
The dataset contains the main components of the news articles published online by the newspaper named Gazzetta di Modena: url of the web page, title, sub-title, text, date of publication, crime category assigned to each news article by the…
3 papers · 0 benchmarks
DIPS-Plus (The Enhanced Database of Interacting Protein Structures for Interface Prediction)
How and where proteins interface with one another can ultimately impact the proteins' functions along with a range of other biological processes.
3 papers · 0 benchmarks
DISRPT2021 (DISRPT2021 shared task on Discourse Unit Segmentation, Connective Detection and Discourse Relation Classification)
The DISRPT 2021 shared task, co-located with CODI 2021 at EMNLP, introduces the second iteration of a cross-formalism shared task on discourse unit segmentation and connective detection, as well as the first iteration of a cross-formalism…
3 papers · 0 benchmarks
DIVOTrack is a cross-view multi-object tracking dataset for DIVerse Open scenes with dense tracking pedestrians in realistic and non-experimental environments.
3 papers · 0 benchmarks
The DeepMind Q&A Dataset consists of two datasets for Question Answering, CNN and DailyMail.
3 papers · 0 benchmarks
Intended to provide freely available data sets in various formats together with basic annotation to be useful for applications in computational linguistics, translation studies and cross-linguistic corpus studies.
3 papers · 0 benchmarks
DOORS (Dataset fOr bOuldeRs Segmentation)
DOORS is a dataset designed for boulders recognition, centroid regression, segmentation, and navigation applications.
3 papers · 0 benchmarks
DPB-5L is a Multilingual KG dataset containing 5 KGs in English, French, Japanese, Greek, and Spanish.
3 papers · 1 benchmark
DR.BENCH (Diagnostic Reasoning Benchmark for clinical natural language processing)
DR.BENCH is a dataset for developing and evaluating cNLP models with clinical diagnostic reasoning ability.
3 papers · 0 benchmarks
This dataset is designed to enhance the progress of event-based optical flow algorithms.
3 papers · 1 benchmark
We present a dataset, DANFEVER, intended for claim verification in Danish.
3 papers · 1 benchmark
DangerousQA refers to a set of harmful questions used to evaluate the safety and behavior of large language models (LLMs) in generating responses.
3 papers · 0 benchmarks
This dataset is the outcome of a data challenge conducted as part of the Dark Machines Initiative and the Les Houches 2019 workshop on Physics at TeV colliders.
3 papers · 0 benchmarks
Two single cell datsets for 3D shape reconstruction from 2D microscopy images used for our three previous publication’s, together with the respective model predictions.
3 papers · 0 benchmarks
The Deep Fakes Dataset is a collection of "in the wild" portrait videos for deepfake detection.
3 papers · 0 benchmarks
DeepFake MNIST+ is a deepfake facial animation dataset.
3 papers · 0 benchmarks
The DeepNets-1M dataset is composed of neural network architectures represented as graphs where nodes are operations (convolution, pooling, etc.) and edges correspond to the forward pass flow of data through the network.
3 papers · 0 benchmarks
DeepSportradar is a benchmark suite of computer vision tasks, datasets and benchmarks for automated sport understanding.
3 papers · 0 benchmarks
https://github.com/anhaidgroup/deepmatcher/blob/master/Datasets.md
3 papers · 0 benchmarks
The data was collected from the music streaming service Deezer (November 2017).
3 papers · 0 benchmarks
DeliData is the first publicly available dataset containing collaborative conversations on solving a cognitive task, consisting of 500 group dialogues and 14k utterances.
3 papers · 1 benchmark
The dermatology differential diagnoses (ddx) dataset for skin condition classification includes expert annotations and model predictions for 1947 cases.
3 papers · 0 benchmarks
DevBench is a comprehensive benchmark designed to evaluate Large Language Models (LLMs) across various stages of the software development lifecycle.
3 papers · 0 benchmarks
The dataset contains historical technical data of Dhaka Stock Exchange (DSE).
3 papers · 0 benchmarks
DiFair serves as a meticulous endeavor to address the oversight in evaluating the impact of bias mitigation on useful gender knowledge while assessing gender neutrality in pretrained language models.
3 papers · 0 benchmarks
The DialSummEval is a multi-faceted dataset of human judgments.
3 papers · 0 benchmarks
Human face Deepfake dataset sampled from large datasets - High Quality Dataset - Diverse Dataset - Challenging Dataset - Large Dataset - Text prompts
3 papers · 0 benchmarks
Diffusion4D is a large-scale, high-quality dynamic 3D(4D) dataset sourced from the vast 3D data corpus of Objaverse-1.0 and Objaverse-XL.
3 papers · 0 benchmarks
Digital Peter is a dataset of Peter the Great's manuscripts annotated for segmentation and text recognition.
3 papers · 1 benchmark
DirtyMNIST is a concatenation of MNIST + AmbiguousMNIST, with 60k samples each in the training set.
3 papers · 0 benchmarks
Disfl-QA is a targeted dataset for contextual disfluencies in an information seeking setting, namely question answering over Wikipedia passages.
3 papers · 0 benchmarks
Distributional MIPLIB is a dataset of Mixed Integer Linear Programming (MILP) instances designed to advance research on learning to optimize \url{https://www.arxiv.org/abs/2406.06954}.
3 papers · 0 benchmarks
DivEMT (Post-Editing Effort Across Typologically-diverse Languages)
DivEMT, the first publicly available post-editing study of Neural Machine Translation (NMT) over a typologically diverse set of target languages.
3 papers · 0 benchmarks
This dataset consists of 1634 biomedical abstracts, expert-annotated for the purpose of extracting information about the efficacy of drug combinations from the scientific literature.
3 papers · 1 benchmark
DrugBank (DrugBank 6.0: the DrugBank Knowledgebase for 2024)
Abstract: First released in 2006, DrugBank (https://go.drugbank.com) has grown to become the 'gold standard' knowledge resource for drug, drug-target and related pharmaceutical information.
3 papers · 1 benchmark
Contains over 70,000 question-answer pairs from both structured tables and unstructured notes from a publicly available Electronic Health Record (EHR).
3 papers · 0 benchmarks
This repository contains gzipped files containing more than 2 million tokens (words) from answers submitted by more than 6,000 students over the course of their first 30 days of using Duolingo.
3 papers · 0 benchmarks
This is a gzipped CSV file containing the 13 million Duolingo student learning traces used in experiments by Settles & Meeder (2016).
3 papers · 0 benchmarks
E-GMD (Expanded Groove MIDI Dataset)
Expanded Groove MIDI dataset (E-GMD) is an automatic drum transcription (ADT) dataset that contains 444 hours of audio from 43 drum kits, making it an order of magnitude larger than similar datasets, and the first with human-performed…
3 papers · 0 benchmarks
E-NER is a publicly available legal Named Entity Recognition (NER) data set.
3 papers · 0 benchmarks
The automated recognition of different vehicle classes and their orientation on aerial images is an important task in the field of traffic research and also finds applications in disaster management, among other things.
3 papers · 0 benchmarks
EC-FUNSD is introduced in [[arXiv:2402.02379]](https://arxiv.org/abs/2402.02379) as a benchmark of semantic entity recognition (SER) and entity linking (EL), designed for the entity-centric robustness evaluation of pre-trained…
3 papers · 2 benchmarks
The ECUST Food Dataset is a food recognition dataset that contains 2978 images Source: https://github.com/Liang-yc/ECUSTFD-resized- Image Source: https://github.com/Liang-yc/ECUSTFD-resized-
3 papers · 0 benchmarks
Bi-temporal images in the EGY-BCD dataset are taken from 4 different regions located in Egypt, including New Mansoura, El Galala City, New Cairo, and New Thebes.
3 papers · 1 benchmark
EMOTyDA (Emotion aware Dialogue Act)
EMOTyDA is a multimodal Emotion aware Dialogue Act dataset collected from open-sourced dialogue datasets.
3 papers · 1 benchmark
EMOVIE is a Mandarin emotion speech dataset including 9,724 samples with audio files and its emotion human-labeled annotation.
3 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.