Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 171 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 8161–8208 of 12,172
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Disaster is a dataset that contains images collected from various sources for three different disasters: fire, water and land.
1 paper · 0 benchmarks
Mapping of detailed discipline tags to one of three broader disciplines (Arts, Science, Business)
1 paper · 0 benchmarks
Dataset Summary The DiscoEval is an English-language Benchmark that contains a test suite of 7 tasks to evaluate whether sentence representations include semantic information relevant to discourse processing.
1 paper · 0 benchmarks
DiscoSense is a benchmark sourced from datasets that contain two sentences connected through a discourse connective.
1 paper · 0 benchmarks
Project: Discrete-Time Modeling of Interturn Short Circuits in Interior PMSMs Authors: Lukas Zezula, Matus Kozovsky, Ludek Buchta and Petr Blaha (corresponding author: Lukas Zezula, e-mail: lukas.zezula@ceitec.vutbr.cz) Affiliation: CEITEC…
1 paper · 0 benchmarks
Extracts diseases and syndromes (DsSs) from more than 65,000 neurology case reports from 66 journals in PubMed over the last six decades from 1955 to 2017.
1 paper · 0 benchmarks
Dissonance Twitter Dataset is a dataset collected from annotating tweets for dissonance.
1 paper · 0 benchmarks
This dataset is named as the DistNLI dataset, which is a synthesized benchmark aiming to probe neural network models from the aspect of conjunctions on distributivity in NLI task in American English.
1 paper · 0 benchmarks
DivShift North American West Coast DivShift Paper | Extended Version | Code https://doi.org/10.1609/aaai.v39i27.35060 Highlighting biases through partitions of volunteer-collected biodiversity data and ecologically relevant context for…
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Dizi is a dataset of music style of the Northern school and the Southern School.
1 paper · 0 benchmarks
DnR-nonverbal is a dataset for cinematic audio source separation (CASS) based on Divide and Remaster (DnR) dataset.
1 paper · 0 benchmarks
DoPose (Dortmund 6D Pose dataset)
DoPose (Dortmund 6D Pose dataset) is a dataset of highly cluttered and closely stacked objects.
1 paper · 0 benchmarks
Doc3DShade extends Doc3D with realistic lighting and shading.
1 paper · 0 benchmarks
This dataset consisting 500 set of caption, table and coresponding paper page, processed from DocBank.
1 paper · 0 benchmarks
We manually annotate 800 sentences from 80 documents in two domains (Healthcare and Transportation) to form a DocOIE dataset for evaluation.
1 paper · 2 benchmarks
DocRED-FE (DocRED with Fine-Grained Entity Type)
DocRED-FE is the DocRED with Fine-Grained Entity Type
1 paper · 0 benchmarks
The DocRED Information Extraction (DocRED-IE) dataset extends the DocRED dataset for the Document-level Closed Information Extraction (DocIE) task.
1 paper · 6 benchmarks
These are the test and training data used for experiments presented in BioNLP 2017.
1 paper · 0 benchmarks
The dataset is composed of 95 unique document texts spanning the period 2005-2022.
1 paper · 0 benchmarks
The MLCommons Dollar Street Dataset is a collection of images of everyday household items from homes around the world that visually captures socioeconomic diversity of traditionally underrepresented populations.
1 paper · 0 benchmarks
An adaption of the MVTec Anomaly Detection dataset, presented in the paper "Domain-independent detection of known anomalies".
1 paper · 1 benchmark
Domains Project is a public dataset contains freely available sorted list of Internet domains.
1 paper · 0 benchmarks
DotPrompts is a set of testcases derived from PragmaticCode, such that each testcase consists of a prompt to a dereference location (a code location having the "." operator in Java).
1 paper · 1 benchmark
DpgMedia2019 is a Dutch news dataset for partisanship detection.
1 paper · 0 benchmarks
Abstract Forecasting methods from averaging regression analysis lines, reversed and direct lines.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
The Drag100 dataset is introduced in the paper "GoodDrag: Towards Good Practices for Drag Editing with Diffusion Models"¹.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
About the Dataset: 4 classes of drinking waste: Aluminium Cans, Glass bottles, PET (plastic) bottles and HDPE (plastic) Milk bottles.
1 paper · 1 benchmark
Driver Micro Hand Gestures (DriverMHG) is a dataset for dynamic recognition of driver micro hand gestures, which consists of RGB, depth and infrared modalities.
1 paper · 0 benchmarks
This dataset consists of a number of sequences that were recorded with a VGA (640x480) event camera (Samsung DVS Gen3) and a conventional RGB camera (Huawei P20 Pro) placed on the windshield of a car driving through Zurich.
1 paper · 0 benchmarks
A synthetic dataset including driving under adverse weather conditions | Autonomous Driving
1 paper · 0 benchmarks
DroidBugs is a benchmark for Automated Program Repair (APR) of Android applications.
1 paper · 0 benchmarks
This dataset contains videos where a flying drone (hexacopter) is captured with multiple consumer-grade cameras (smartphones, compact cameras, gopro,...) with highly accurate 3D drone trajectory ground truth recorderd by a precise…
1 paper · 0 benchmarks
For the Drone-vs-Bird Detection Challenge 2021, 77 different video sequences have been made available as training data.
1 paper · 1 benchmark
Drone-Anomaly rovides 37 training video sequences and 22 testing video sequences from 7 different realistic scenes with various anomalous events.
1 paper · 0 benchmarks
From DroneDeploy: We’ve collected a dataset of aerial orthomosaics and elevation images.
1 paper · 1 benchmark
人群计数旨在识别物体的数量,在智能交通、城市管理和安全监控中发挥着重要作用。由于比例变化、照明变化、遮挡和较差的成像条件,尤其是在夜间和雾霾条件下,人群计数的任务非常具有挑战性。 在本文中,我们提出了一个基于无人机的 RGB-Thermal 人群计数数据集 (DroneRGBT),该数据集由 3600…
1 paper · 1 benchmark
The data used for all results in this paper can be found here.
1 paper · 0 benchmarks
DrugComb is an open-access, community-driven data portal where the results of drug combination screening studies for a large variety of cancer cell lines are accumulated, standardized and harmonized.
1 paper · 0 benchmarks
DrugProt corpus, where domain experts have exhaustively labeled:(a) all chemical and gene mentions, and (b) all binary relationships between them corresponding to a specific set of biologically relevant relation types (DrugProt relation…
1 paper · 1 benchmark
Seven different types of dry beans were used in this research, taking into account the features such as form, shape, type, and structure by the market situation.
1 paper · 0 benchmarks
Dubbing Test Set consists of two subsets extracted from the En→De test set of COVOST-2, a large-scale multilingual speech translation corpus based on Common Voice.
1 paper · 0 benchmarks
DukeMTMC-VideoReID-DL processed with our re-Detect and Link (DL) module.
1 paper · 0 benchmarks
Replication datasets (200 million rows) used in experiments by Yancey & Settles (2020).
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.