Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 82 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 3889–3936 of 12,172
This is a synthetic dataset for defect detection on textured surfaces.
6 papers · 1 benchmark
An extension of DAIR-V2X with addition temporal information
6 papers · 0 benchmarks
A dataset for interactive segmentation with simulated initial masks.
6 papers · 1 benchmark
Dataset Description: The interaction of 72 kinase inhibitors with 442 kinases covering >80% of the human catalytic protein kinome.
6 papers · 3 benchmarks
DIBCO 2017 is the international Competition on Document Image Binarization organized in conjunction with the ICDAR 2017 conference.
6 papers · 0 benchmarks
The DLR-ACD dataset is a collection of aerial images for crowd counting and density estimation, as well as for person localization at mass events.
6 papers · 1 benchmark
DRTiD is a benchmark dataset for DR grading, consisting of 3,100 two-field fundus images.
6 papers · 0 benchmarks
A set of 10 DSC datasets (reviews of 10 products) to produce sequences of tasks.
6 papers · 1 benchmark
DaLAJ 1.0, a dataset for Linguistic Acceptability Judgments for Swedish, comprising 9,596 sentences in its first version; and the initial experiment using it for the binary classification task.
6 papers · 1 benchmark
DaNE (Danish Dependency Treebank)
Danish Dependency Treebank (DaNE) is a named entity annotation for the Danish Universal Dependencies treebank using the CoNLL-2003 annotation scheme.
6 papers · 5 benchmarks
DebateSum consists of 187328 debate documents, arguments (also can be thought of as abstractive summaries, or queries), word-level extractive summaries, citations, and associated metadata organized by topic-year.
6 papers · 1 benchmark
Integrals and Differential Equations Dataset using the generators from the paper
6 papers · 0 benchmarks
DiFF (Diffusion Facial Forgery Detection)
Detecting diffusion-generated images has recently grown into an emerging research area.
6 papers · 0 benchmarks
DialogStudio, a meticulously curated collection of dialogue datasets.
6 papers · 0 benchmarks
DocCVQA (Document Collection Visual Question Answering)
DocCVQA is a Document Visual Question Answering dataset, where the questions are posed over a whole collection of 14,362 scanned documents.
6 papers · 0 benchmarks
Cant (also known as doublespeak, cryptolect, argot, anti-language or secret language) is important for understanding advertising, comedies and dog-whistle politics.
6 papers · 0 benchmarks
DreamBench++ is a human-aligned benchmark for personalized image generation¹.
6 papers · 0 benchmarks
DroneCrowd is a benchmark for object detection, tracking and counting algorithms in drone-captured videos.
6 papers · 0 benchmarks
DroneSURF (DroneSURF: Benchmark Dataset for Drone-based Face Recognition)
Drone Surveillance of Faces, is a large-scale drone dataset intended to facilitate research for face recognition using drones.
6 papers · 1 benchmark
The original dataset for "ECG5000" is a 20-hour long ECG downloaded from Physionet.
6 papers · 3 benchmarks
EHR-RelB is a benchmark dataset for biomedical concept relatedness, consisting of 3630 concept pairs sampled from electronic health records (EHRs).
6 papers · 0 benchmarks
EarthVQA (A multi-modal multi-task VQA dataset for remote sensing)
Earth vision research typically focuses on extracting geospatial object locations and categories but neglects the exploration of relations between objects and comprehensive reasoning.
6 papers · 1 benchmark
EmbSpatial-Bench (EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models)
The recent rapid development of Large Vision-Language Models (LVLMs) has indicated their potential for embodied tasks.
6 papers · 1 benchmark
The EMODB database is the freely available German emotional database.
6 papers · 1 benchmark
The EuroCity Persons dataset provides a large number of highly diverse, accurate and detailed annotations of pedestrians, cyclists and other riders in urban traffic scenes.
6 papers · 0 benchmarks
This dataset contains around 10000 videos generated by various methods using the Prompt list.
6 papers · 1 benchmark
F-CelebA - This dataset is adapted from federated learning.
6 papers · 1 benchmark
Falling Things (FAT) is a dataset for advancing the state-of-the-art in object detection and 3D pose estimation in the context of robotics.
6 papers · 0 benchmarks
FDF (Flickr Diverse Faces)
A diverse dataset of human faces, including unconventional poses, occluded faces, and a vast variability in backgrounds.
6 papers · 0 benchmarks
FFHQ-Aging is a Dataset of human faces designed for benchmarking age transformation algorithms as well as many other possible vision tasks.
6 papers · 0 benchmarks
FMFCC-A is a large publicly-available Mandarin dataset for synthetic speech detection, which contains 40,000 synthesized Mandarin utterances that generated by 11 Mandarin TTS systems and two Mandarin VC systems, and 10,000 genuine Mandarin…
6 papers · 0 benchmarks
Flightmare is composed of two main components: a configurable rendering engine built on Unity and a flexible physics engine for dynamics simulation.
6 papers · 0 benchmarks
The Food-101N dataset is introduced in "CleanNet: Transfer Learning for Scalable Image Training with Label Noise (CVPR'18).
6 papers · 1 benchmark
The Freiburg Forest dataset was collected using a Viona autonomous mobile robot platform equipped with cameras for capturing multi-spectral and multi-modal images.
6 papers · 2 benchmarks
French TimeBank, a corpus for French annotated in ISO-TimeML.
6 papers · 1 benchmark
French Wikipedia is a dataset used for pretraining the CamemBERT French language model.
6 papers · 0 benchmarks
FrenchMedMCQA (FrenchMedMCQA: A French Multiple-Choice Question Answering Dataset for Medical domain)
This paper introduces FrenchMedMCQA, the first publicly available Multiple-Choice Question Answering (MCQA) dataset in French for medical domain.
6 papers · 1 benchmark
A public data set of walking full-body kinematics and kinetics in individuals with Parkinson’s disease
6 papers · 1 benchmark
FusedChat is an inter-mode dialogue dataset.
6 papers · 1 benchmark
GBCU (Gallbladder Cancer Ultrasound Dataset)
GBCU is the first public dataset for Gallbladder Cancer identification from Ultrasound images.
6 papers · 1 benchmark
Geo-Diverse Visual Commonsense Reasoning (GD-VCR) is a new dataset to test vision-and-language models' ability to understand cultural and geo-location-specific commonsense.
6 papers · 1 benchmark
This dataset is a new benchmark, grounded in real-world usages is developed to support more authentic and comprehensive evaluation of image editing models.
6 papers · 2 benchmarks
GFF (Global Flood Forecasting)
Floods are among the most common and devastating natural hazards, imposing immense costs on our society and economy due to their disastrous consequences.
6 papers · 1 benchmark
GICoref (Gender Inclusive Coreference)
GICoref is a fully annotated coreference resolution dataset written by and about trans people.
6 papers · 0 benchmarks
The German Lipreading dataset consists of 250,000 publicly available videos of the faces of speakers of the Hessian Parliament, which was processed for word-level lip reading using an automatic pipeline.
6 papers · 0 benchmarks
GRB (Graph Robustness Benchmark)
Graph Robustness Benchmark (GRB) provides scalable, unified, modular, and reproducible evaluation on the adversarial robustness of graph machine learning models.
6 papers · 0 benchmarks
GTSinger (GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks)
The scarcity of high-quality and multi-task singing datasets significantly hinders the development of diverse controllable and personalized singing tasks, as existing singing datasets suffer from low quality, limited diversity of languages…
6 papers · 0 benchmarks
GWA (Geometric-Wave Acoustic)
GWA is a large-scale audio dataset of over 2 million synthetic room impulse responses (IRs) and their corresponding detailed geometric and simulation configurations.
6 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.