Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 99 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 4705–4752 of 12,172
The dataset consists of over 350,000 public domain patent drawings collected from the United States Patent and Trademark Office (USPTO).
4 papers · 1 benchmark
During the MILAN research project (MachIne Learning for AstroNomy), we have compiled a large collection of deep sky images during Electronically Assisted Astronomy sessions in Luxembourg, France, Belgium.
4 papers · 0 benchmarks
DEVAI is a benchmark of 55 realistic AI development tasks.
4 papers · 0 benchmarks
DiS-ReX is a multilingual dataset for distantly supervised (DS) relation extraction (RE).
4 papers · 0 benchmarks
DiSCQ (Discharge Summary Clinical Questions)
DiSCQ is a newly curated question dataset composed of 2,000+ questions paired with the snippets of text (triggers) that prompted each question.
4 papers · 0 benchmarks
DiaASQ (Conversational Aspect-based Sentiment Quadruple Extraction)
DiaASQ is a fine-grained Aspect-based Sentiment Analysis (ABSA) benchmark under the conversation scenario.
4 papers · 2 benchmarks
A new English-French test set for the evaluation of Machine Translation (MT) for informal, written bilingual dialogue.
4 papers · 1 benchmark
DiaMOS Plant (A Dataset for Diagnosis and Monitoring Plant Disease)
Abstract The classification and recognition of foliar diseases is an increasingly developing field of research, where the concepts of machine and deep learning are used to support agricultural stakeholders.
4 papers · 0 benchmarks
The DiaTrend dataset is composed of intensive longitudinal data from wearable medical devices, including a total of 27,561 days of continuous glucose monitor data and 8,220 days of insulin pump data from 54 patients with diabetes.
4 papers · 0 benchmarks
DiagSet is a histopathological dataset for prostate cancer detection.
4 papers · 0 benchmarks
Diamante is a novel and efficient framework consisting of a data collection strategy and a learning method to boost the performance of pre-trained dialogue models.
4 papers · 0 benchmarks
DoMSEV (Dataset of Multimodal Semantic Egocentric Video)
The Dataset of Multimodal Semantic Egocentric Video (DoMSEV) contains 80-hours of multimodal (RGB-D, IMU, and GPS) data related to First-Person Videos with annotations for recorder profile, frame scene, activities, interaction, and…
4 papers · 0 benchmarks
Draper VDISC Dataset - Vulnerability Detection in Source Code The dataset consists of the source code of 1.27 million functions mined from open source software, labeled by static analysis for potential vulnerabilities.
4 papers · 0 benchmarks
DrivAerNet (A Parametric Car Dataset for Data-driven Aerodynamic Design and Graph-Based Drag Prediction)
DrivAerNet is a large-scale, high-fidelity CFD dataset of 3D industry-standard car shapes designed for data-driven aerodynamic design.
4 papers · 0 benchmarks
The images in DukeMTMC-attribute dataset comes from Duke University.
4 papers · 1 benchmark
This dataset provides a large number of training and testing example which is sufficient for a deep learning approach to address Dunhuang Grotto Painting restoration.
4 papers · 0 benchmarks
ECG-Image-Database (Digitization and Classification of ECG Images: The George B. Moody PhysioNet Challenge 2024)
The George B.
4 papers · 1 benchmark
EDEN (Enclosed garDEN) is a multimodal synthetic dataset, a dataset for nature-oriented applications.
4 papers · 0 benchmarks
EDGE-IIOTSET (A NEW COMPREHENSIVE REALISTIC CYBER SECURITY DATASET OF IOT AND IIOT APPLICATIONS: CENTRALIZED AND FEDERATED LEARNING)
ABSTRACT In this project, we propose a new comprehensive realistic cyber security dataset of IoT and IIoT applications, called Edge-IIoTset, which can be used by machine learning-based intrusion detection systems in two different modes,…
4 papers · 0 benchmarks
EDUB-Seg (Egocentric Dataset of the University of Barcelona – Segmentation)
Egocentric Dataset of the University of Barcelona – Segmentation (EDUB-Seg) is a dataset for egocentric event segmentation acquired by the Narrative Clip, which takes a picture every 30 seconds.
4 papers · 0 benchmarks
This data set consists of over 1500 one- and two-minute EEG recordings, obtained from 109 volunteers.
4 papers · 1 benchmark
EGSet12 (Twelve real & original solo electric guitar performances with diverse playing styles to evaluate guitar tablature transcription)
EGSet12 is a small dataset with twelve original and real solo electric guitar performances (31.65 seconds avg.
4 papers · 0 benchmarks
EMU (Edited Media Understanding)
48k question-answer pairs written in rich natural language.
4 papers · 0 benchmarks
ESAD (SARAS Endoscopic Surgeon Action Detection)
ESAD is a large-scale dataset designed to tackle the problem of surgeon action detection in endoscopic minimally invasive surgery.
4 papers · 0 benchmarks
ETHEC (ETH Entomological Collection (ETHEC) Dataset)
It includes 47,978 butterfly images with a 4-level label-hierarchy.
4 papers · 0 benchmarks
The dataset contains 7000 videos: native, altered and exchanged through social platforms.
4 papers · 0 benchmarks
EVJVQA (English-Japanese-Vietnamese Visual Question Answering)
EVJVQA, the first multilingual Visual Question Answering dataset with three languages: English, Vietnamese, and Japanese, is released in this task.
4 papers · 0 benchmarks
EdAcc (Edinburgh International Accents of English Corpus)
The Edinburgh International Accents of English Corpus (EdAcc) is a new automatic speech recognition (ASR) dataset composed of 40 hours of English dyadic conversations between speakers with a diverse set of accents.
4 papers · 0 benchmarks
Election2020 is a Twitter dataset on the 2020 US presidential elections.
4 papers · 0 benchmarks
The endoscopic SLAM dataset (EndoSLAM) is a dataset for depth estimation approach for endoscopic videos.
4 papers · 0 benchmarks
Cholecystectomy is a very common abdominal surgical procedure almost ubiquitously performed with a laparoscopic approach, hence guided by an endoscopic video.
4 papers · 1 benchmark
This dataset contains around 5000 scholarly articles and their corresponding easy summary from eureka alert blog, the dataset can be used for the combined task of summarization and simplification.
4 papers · 2 benchmarks
A SAR version of the EuroSAT dataset.
4 papers · 1 benchmark
Builds upon the event-centric EventKG knowledge graph and language-specific information on user interactions with events, entities, and their relations derived from the Wikipedia clickstream.
4 papers · 0 benchmarks
ExVo2022 (ICML ExVo 2022 Workshop & Competition Data)
Baseline code for the three tracks of ExVo 2022 competition.
4 papers · 0 benchmarks
Explorall font image dataset https://drive.google.com/file/d/1P2DbNbVw4QWcV1YdzE7zsDKilmd3pO/view
4 papers · 1 benchmark
F-SIOL-310 is a robotic dataset and benchmark for Few-Shot Incremental Object Learning, which is used to test incremental learning capabilities for robotic vision from a few examples.
4 papers · 0 benchmarks
FACTIFY (a dataset on multi-modal fact verification)
FACTIFY is a dataset on multi-modal fact verification.
4 papers · 0 benchmarks
FES (Fisheye Evaluation Suite)
FES is an indoor dataset that can be used for evaluation of deep learning approaches.
4 papers · 0 benchmarks
FEWS (FEWS: Large-Scale, Low-Shot Word Sense Disambiguation with the Dictionary)
FEWS (Few-shot Examples of Word Senses) is a few-shot dataset for English Word Sense Disambiguation (WSD) gathered from Wiktionary, an online, crowd-sourced dictionary.
4 papers · 1 benchmark
FG-OVD (Fine-Grained Open-Vocabulary object Detection benchmarks)
Benchmark Suite Description for PapersWithCode Fine-Grained Open-Vocabulary Detection (FG-OVD) Benchmark Suite The FG-OVD benchmark suite evaluates the ability of open-vocabulary object detectors to discern fine-grained object properties…
4 papers · 0 benchmarks
The FIGR-8 database is a dataset containing 17,375 classes of 1,548,256 images representing pictograms, ideograms, icons, emoticons or object or conception depictions.
4 papers · 0 benchmarks
FINO-Net is a multimodal (RGB, depth and audio) dataset, containing 229 real-world manipulation data of 5 different manipulation types recorded with a Baxter robot.
4 papers · 0 benchmarks
FIRE (Fundus Image Registration Dataset)
Fundus Image Registration Dataset (FIRE) is a dataset consisting of 129 retinal images forming 134 image pairs.
4 papers · 1 benchmark
FLAG3D is a large-scale 3D fitness activity dataset with language instruction containing 180K sequences of 60 categories.
4 papers · 0 benchmarks
The French National Institute of Geographical and Forest Information (IGN) has the mission to document and measure land-cover on French territory and provides referential geographical datasets, including high-resolution aerial images and…
4 papers · 1 benchmark
FR-FS (Fall Recognition in Figure Skating)
The FR-FS dataset contains 417 videos collected from FIV dataset and Pingchang 2018 Winter Olympic Games.
4 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.