Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 168 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 8017–8064 of 12,172
DAVIS-Edit is a curated testing benchmark for video editing.
1 paper · 0 benchmarks
This dataset includes Direct Borohydride Fuel Cell (DBFC) impedance and polarization test in anode with Pd/C, Pt/C and Pd decorated Ni–Co/rGO catalysts.
1 paper · 0 benchmarks
DPB-5L is a Multilingual KG dataset containing 5 KGs in English, French, Japanese, Greek, and Spanish.
1 paper · 0 benchmarks
The DBP2.0 dataset can be downloaded from the figshare repository.
1 paper · 1 benchmark
The dataset provides the content of all articles for 128 Wikipedia languages.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
[comment]:<> (Data for the paper "Deciphering Environmental Air Pollution with Large Scale City Data") Main Dataset citypollutiondata.csv Relevant Columns: Date: Date of the sample City: City of the sample Xmedian: Median value of the…
1 paper · 0 benchmarks
DEEP-VOICE: Real-time Detection of AI-Generated Speech for DeepFake Voice Conversion This dataset contains examples of real human speech, and DeepFake versions of those speeches by using Retrieval-based Voice Conversion.
1 paper · 1 benchmark
This dataset contains synthetic text data generated to train models for text generation.
1 paper · 0 benchmarks
DET is a lane detection dataset that consists of the raw event data, accumulated images over 30ms and corresponding lane labels.
1 paper · 1 benchmark
DEplain-APA-doc: A German Parallel Corpus for Document Simplification on News Texts DEplain is a new dataset of parallel, professionally written and manually aligned simplifications in plain German “plain DE” (or in German: “Einfache…
1 paper · 1 benchmark
DEplain-web-doc: A German Parallel Corpus for Document Simplification on Web Texts DEplain is a new dataset of parallel, professionally written and manually aligned simplifications in plain German “plain DE” (or in German: “Einfache…
1 paper · 1 benchmark
We provide DGL compatible graphs in lmdb format for the OpenCatalyst IS2RE task based on the OC20 dataset.
1 paper · 0 benchmarks
Object Detection data set created from the engine DeepGTAV, which is based on the video game GTAV.
1 paper · 0 benchmarks
HeLa cells on a flat glass Dr.
1 paper · 1 benchmark
DIGITal (Digitally Generated Numerals)
Digitally Generated Numerals (DIGITal) Description The Digitally Generated Numerals (DIGITal) dataset consists of 100,000 image pairs representing digits from 0 to 9.
1 paper · 0 benchmarks
DINOS (Diverse INdustrial Operation Sounds)
DINOS (Diverse INdustrial Operation Sounds) is a large-scale, open-access dataset consisting of over 74,000 audio samples totaling more than 1,093 hours, collected from a wide range of industrial acoustic scenarios.
1 paper · 0 benchmarks
DIO (Discovering Interacted Objects)
Discovering Interacted Objects (DIO) is a benchmark containing 51 interactions and 1,000+ objects designed for Spatio-temporal Human-Object Interaction (ST-HOI) detection.
1 paper · 0 benchmarks
Datasets are built upon three other datasets: DISEC 2013, RVL-CDIP, RDCL 2017.
1 paper · 1 benchmark
DLBCL-Morph is a dataset containing 42 digitally scanned high-resolution tissue microarray (TMA) slides accompanied by clinical, cytogenetic, and geometric features from 209 DLBCL cases.
1 paper · 0 benchmarks
DLCN (Dynamic-lighting Conditions at Night)
DLCN (Dynamic-lighting Conditions at Night) is a dataset collected specifically for rPPG signal evaluation under complex lighting environments.
1 paper · 0 benchmarks
Contains ~60000 HD images of Deformable Linear Objects (DLOs) generated using blender.
1 paper · 0 benchmarks
DMAD (Deepfake Massively Annotated Databases)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Medical VQA dataset built from the IDRiD and eOphta datasets.
1 paper · 0 benchmarks
A large scale dataset to pre-train optical flow prediction network.
1 paper · 0 benchmarks
DODa (Darija Open Dataset)
Darija Open Dataset (DODa) is an open-source project for the Moroccan dialect.
1 paper · 0 benchmarks
The dataset was collected from DOTA 2 using OpenDota API via Python.
1 paper · 0 benchmarks
DP0E is a public dataset of anti-counterfeiting printable graphical codes (PGC) based on DataMatrix modulation.
1 paper · 0 benchmarks
dataset in WWW 2019 "DPLink: User Identity Linkage via Deep Neural Network From Heterogeneous Mobility Data".
1 paper · 0 benchmarks
We provide the code to generate base and query vector datasets for similarity search benchmarking and evaluation on high-dimensional vectors stemming from large language models.
1 paper · 0 benchmarks
This dataset is a new variant of the voice cloning toolkit (VCTK) dataset: device-recorded VCTK (DR-VCTK), where the high-quality speech signals recorded in a semi-anechoic chamber using professional audio devices are played back and…
1 paper · 0 benchmarks
DRAM dataset is the first dataset to introduce a fully labeled test set for the task of semantic segmentation of art paintings.
1 paper · 0 benchmarks
DRIFT (Domain-Adaptive Regression for Forest Monitoring)
The DRIFT dataset includes 25k image patches collected in five European countries sourced from aerial and nanosatellite image archives.
1 paper · 0 benchmarks
DRKG (Drug Repurposing Knowledge graph)
Drug Repurposing Knowledge Graph (DRKG) is a comprehensive biological knowledge graph relating genes, compounds, diseases, biological processes, side effects and symptoms.
1 paper · 0 benchmarks
DSBEC (Dark solitons in BECs dataset)
The data set consists of 6257 labeled images of Bose-Einstein condensates (BECs) with and without solitonic excitations, including kink solitons and solitonic vortices.
1 paper · 0 benchmarks
DSSN (DAIICT Spatio-Temporal Network)
DSSN is a spatiotemporal dataset of 0.7 million data points of continuous location data logged at an interval of every 2 minutes by mobile phones of 46 subjects.
1 paper · 0 benchmarks
DSurVD (Distorted Surveillance Video Database)
A large-scale dataset, namely Distorted Surveillance Video Database (DSurVD), which can be downloaded from the link: https://sites.google.com/site/sorsyuanyuan/home/dsurvd Image source: https://sites.google.com/site/sorsyuanyuan/home/dsurvd
1 paper · 0 benchmarks
DTBM (Digital Twin Benchmark Model)
DTBM is a benchmark dataset for Digital Twins that reflects these characteristics and look into the scaling challenges of different knowledge graph technologies.
1 paper · 0 benchmarks
DUC 2006 (Document Understanding Conferences)
There is currently much interest and activity aimed at building powerful multi-purpose information systems.
1 paper · 0 benchmarks
DUSK (Do not Unlearn Shared Knowledge)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
This is a dataset of 22.5 hours of synthesized audio using the open-source learnfm clone of the DX7 FM synthesizer, based upon 31K presets from Bobby Blue.
1 paper · 0 benchmarks
A line drawing restoration dataset which consists of 71 line drawing sketches by Leonardo Da Vinci.
1 paper · 0 benchmarks
The dataset comprises motion sensor data of 19 daily and sports activities each performed by 8 subjects in their own style for 5 minutes.
1 paper · 0 benchmarks
Daily load patterns (Daily load patterns of six global exchanges as recorded by vwd/Infront Financial Technology)
This data set provides fine-granular statistics on trading traffic generated by six global exchanges over the course of two days in February 2019 for a set of representative feeds and recorded by the systems of vwd Vereinigte…
1 paper · 0 benchmarks
DailyMoth-70h is a fully self-contained ASL-to-English sign language dataset containing over 70h of video (48K clips) with aligned English captions of a single native ASL signer (white, male, and early middle-aged) from the ASL news…
1 paper · 0 benchmarks
The Daimler Monocular Pedestrian Detection dataset is a dataset for pedestrian detection in urban environments.
1 paper · 0 benchmarks
A dataset of images obtained from DALL-E 3 for 67 countries and 10 concept classes, similar to DollarStreet images.
1 paper · 0 benchmarks
DanbooRegion is a dataset consists of 5377 in-the-wild illustration downloaded from the Danbooru2018 and region segment map annotation pairs samples are provided as at 1024px 8-bit RGB images, and region segment maps as int-32 index images.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.