Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 130 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 6193–6240 of 12,172
The dataset contains 578,731 structures for methane combustion and their energies and forces under MN15/6-31G level.
2 papers · 0 benchmarks
Dataset of Legal Documents consists of court decisions from 2017 and 2018 were selected for the dataset, published online by the Federal Ministry of Justice and Consumer Protection.
2 papers · 0 benchmarks
This is a part of dataset of the paper published in UAI 2021 (37th Conference on Uncertainty in Artificial Intelligence).
2 papers · 0 benchmarks
DeCOCO is a bilingual (English-German) corpus of image descriptions, where the English part is extracted from the COCO dataset, and the German part are translations by a native German speaker.
2 papers · 0 benchmarks
The dataset enables the mapping from text space to numerical space and vice versa.
2 papers · 0 benchmarks
DeePhy is a novel DeepFake Phylogeny dataset consisting of 5040 DeepFake videos generated using three different generation techniques.
2 papers · 0 benchmarks
The Deep Blending Dataset comprises 19 diverse scenes, offering comprehensive resources for free-viewpoint image-based rendering (IBR).
2 papers · 1 benchmark
DeepPCB Dataset Link : A dataset contains 1,500 image pairs, each of which consists of a defect-free template image and an aligned tested image with annotations including positions of 6 most common types of PCB defects: open, short,…
2 papers · 1 benchmark
The Java dataset introduced in DeepCom (Deep Code Comment Generation), commonly used to evaluate automated code summarization.
2 papers · 1 benchmark
This collection contains data and code associated with the IPCAI/IJCARS 2020 paper “Automatic Annotation of Hip Anatomy in Fluoroscopy for Robust and Efficient 2D/3D Registration.” The data hosted here consists of annotated datasets of…
2 papers · 0 benchmarks
Delicious : This data set contains tagged web pages retrieved from the website delicious.com.
2 papers · 1 benchmark
DenseUAV is a dataset of drone and satellite perspectives collected from 14 universities in low-altitude urban scenes.
2 papers · 0 benchmarks
Hierarchical-multilabel classification dataset for functional genomics
2 papers · 1 benchmark
Includes gold-standard labels for identifying statements of desire, textual evidence for desire fulfillment, and annotations for whether the stated desire is fulfilled given the evidence in the narrative context.
2 papers · 0 benchmarks
DialogUSR dataset covers 23 domains with a multi-step crowd-sourcing procedure.
2 papers · 0 benchmarks
The Dialogue Fairness dataset is used to evaluate and understand fairness in dialogue models, focusing on gender and racial biases.
2 papers · 0 benchmarks
435 vocal presets retrieval from MedleyDB and a private collection of multi-track mixes.
2 papers · 0 benchmarks
DifferSketching is a dataset of freehand sketches to understand how differently professional and novice users sketch 3D objects.
2 papers · 0 benchmarks
This is a dataset used to test deep learning-supported deep learning for fault diagnosis: - A digital twin model for a robot.
2 papers · 1 benchmark
DisKnE (Disease Knowledge Evaluation)
DisKnE is a benchmark for Disease Knowledge Evaluation built from MedNLI and MEDIQA-NLI.
2 papers · 0 benchmarks
The DispScenes dataset was created to address the specific problem of disparate image matching.
2 papers · 0 benchmarks
Doodle to UI Dataset contains 11 thousand drawings from 16 categories.
2 papers · 0 benchmarks
Drone-Action (Drone-Action: An Outdoor Recorded Drone Video Dataset for Action Recognition)
Website: https://asankagp.github.io/droneaction/
2 papers · 1 benchmark
DropletVideo is a project exploring high-order spatio-temporal consistency in image-to-video generation.
2 papers · 0 benchmarks
Estimating camera motion in deformable scenes poses a complex and open research challenge.
2 papers · 1 benchmark
Background: Lung cancer risk classification is an increasingly important area of research as low-dose thoracic CT screening programs have become standard of care for patients at high risk for lung cancer.
2 papers · 1 benchmark
Dusha (Dusha Crowd, Dusha Podcast)
Dusha is a dataset for speech emotion recognition (SER) tasks.
2 papers · 2 benchmarks
DyML-Animal is based on animal images selected from ImageNet-5K [1].
2 papers · 1 benchmark
DyML-Product is derived from iMaterialist-2019, a hierarchical online product dataset.
2 papers · 1 benchmark
DyML-Vehicle merges two vehicle re-ID datasets PKU VehicleID [1], VERI-Wild [1].
2 papers · 1 benchmark
To provide ground truth supervision for video consistency modeling, we build up a high-quality dynamic OLAT dataset.
2 papers · 0 benchmarks
E-ReDial (Explainable Recommendation Dialogues)
E-ReDial is a conversational recommender system dataset with high-quality explanations.
2 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 1 benchmark
The EARS-Reverb dataset uses real recorded room impulse responses (RIRs) from multiple public datasets (ACE-Challenge, AIR, ARNI, BRUDEX, dEchorate, DetmoldSRIR, and Palimpsest).
2 papers · 1 benchmark
This dataset contains around 5K pairs of aligned images captured using Canon 70D DSLR with low and high apertures, modeling normal photos and photos with bokeh (blur) effect.
2 papers · 0 benchmarks
A large-scale benchmark with 1605 high-resolution, well-annotated images, featuring more complex scenes and a wider range of DOF settings.
2 papers · 1 benchmark
EDFace-Celeb-1M is a public Ethnically Diverse Face dataset which is used to benchmark the task of face hallucination.
2 papers · 0 benchmarks
EDGAR-CORPUS is a novel corpus comprising annual reports from all the publicly traded companies in the US spanning a period of more than 25 years.
2 papers · 0 benchmarks
EDNA-Covid is a multilingual, large-scale dataset of coronavirus-related tweets collected since January 25, 2020.
2 papers · 0 benchmarks
EFO-1-QA is a new dataset to benchmark the combinatorial generalizability of Complex Query Answering (CQA) models by including 301 different queries types, which is 20 times larger than existing datasets.
2 papers · 0 benchmarks
Contains annotations of human activity with different sub-actions, e.g., activity Ping-Pong with four sub-actions which are pickup-ball, hit, bounce-ball and serve.
2 papers · 0 benchmarks
EHE (Elderly Home Exercise)
Human Action Evaluation (HAE) has rarely been applied to real-world disease monitoring, the EHE dataset aims to gather sample data to validate effective HAE methods that could then be expanded on a larger validation scale.
2 papers · 1 benchmark
ELFW (Extended Labeled Faces in-the-Wild)
Extended Labeled Faces in-the-Wild (ELFW) is a dataset supplementing with additional face-related categories —and also additional faces— the originally released semantic labels in the vastly used Labeled Faces in-the-Wild (LFW) dataset.
2 papers · 0 benchmarks
In EMDS-6, there are 21 classes of environmental microorganisms (EMs).
2 papers · 0 benchmarks
ENRICH (Multi-purposE dataset for beNchmaRking In Computer vision and pHotogrammetry)
A new synthetic, multi-purpose dataset - called ENRICH - for testing photogrammetric and computer vision algorithms.
2 papers · 0 benchmarks
ENTIGEN (Ethical NaTural Language Interventions in Text-to-Image GENeration)
ENTIGEN is a benchmark dataset to evaluate the change in image generations conditional on ethical interventions across three social axes -- gender, skin color, and culture.
2 papers · 0 benchmarks
EPISURG (EPISURG: a dataset of postoperative MRI for quantitative analysis of resection neurosurgery for refractory epilepsy)
EPISURG is a clinical dataset of T1-weighted magnetic resonance images (MRI) from 430 epileptic patients who underwent resective brain surgery at the National Hospital of Neurology and Neurosurgery (Queen Square, London, United Kingdom)…
2 papers · 0 benchmarks
ERATO is a large-scale multi-modal dataset for Pairwise Emotional Relationship Recognition (PERR).
2 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.