Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 89 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 4225–4272 of 12,172
ComplexCodeEval ComplexCodeEval is an evaluation benchmark designed to accommodate multiple downstream tasks, accurately reflect different programming environments, and deliberately avoid data leakage issues.
5 papers · 0 benchmarks
CovidET (Emotions and their Triggers during Covid-19)
Crises such as the COVID-19 pandemic continuously threaten our world and emotionally affect billions of people worldwide in distinct ways.
5 papers · 0 benchmarks
A richer dataset based on real items on Craigslist.
5 papers · 0 benchmarks
CrossRE is a cross-domain benchmark for Relation Extraction (RE), which comprises six distinct text domains and includes multi-label annotations.
5 papers · 0 benchmarks
Cube++ is a novel dataset for the color constancy problem that continues on the Cube+ dataset.
5 papers · 1 benchmark
D3D-HOI is a dataset of monocular videos with ground truth annotations of 3D object pose, shape and part motion during human-object interactions.
5 papers · 0 benchmarks
DBE-KT22 contains student exercise answering activities collected through an online practicing platform for the database systems course taught at the Australian National University within the period 2018-2021.
5 papers · 0 benchmarks
TAU Urban Acoustic Scenes 2019 Mobile development dataset consists of 10-seconds audio segments from 10 acoustic scenes: Airport Indoor shopping mall Metro station Pedestrian street Public square Street with medium level of traffic…
5 papers · 1 benchmark
A large-scale and diverse duet interactive dance dataset.
5 papers · 0 benchmarks
DIBCO 2013 is the international Document Image Binarization Contest organized in the context of ICDAR 2013 conference.
5 papers · 0 benchmarks
DIMO (Dataset of Industrial Metal Objects)
The Industrial Metal Objects dataset is a diverse dataset of industrial metal objects.
5 papers · 0 benchmarks
A corpus of Offensive Language and Hate Speech Detection for Danish This DKhate dataset contains 3600 comments from the web annotated for offensive language, following the Zampieri et al.
5 papers · 1 benchmark
DREAM-dataset (Deep Robot-to-camera Extrinsics for Articulated Manipulators)
The DREAM dataset is introduce by the paper "Camera-to-Robot Pose Estimation from a Single Image" (ICRA 2020).
5 papers · 1 benchmark
DaN+ is a new multi-domain corpus and annotation guidelines for Danish nested named entities (NEs) and lexical normalization to support research on cross-lingual cross-domain learning for a less-resourced language.
5 papers · 0 benchmarks
DeToxy (DeToxy: A Large-Scale Multimodal Dataset for Toxicity Classification in Spoken Utterances)
DeToxy is a publicly available toxicity annotated dataset for the English language.
5 papers · 0 benchmarks
It contains 19 HDF5 files that represent a data collection campaign run on the NI mmWave Transceiver System with four SiBeam 60 GHz radio heads and on two Pi-Radio digital 60 GHz radios.
5 papers · 0 benchmarks
The rise of deepfake images, especially of well-known personalities, poses a serious threat to the dissemination of authentic information.
5 papers · 0 benchmarks
This basketball dataset was acquired under the Walloon region project DeepSport, using the Keemotion system installed in multiple arenas.
5 papers · 0 benchmarks
Dem@Care is providing the following datasets, which are collected during lab and home experiments.
5 papers · 0 benchmarks
Diabetes (Diabetes 130-US Hospitals for Years 1999-2008)
What do the instances in this dataset represent?
5 papers · 3 benchmarks
The Distress Analysis Interview Corpus/Wizard-of-Oz set (DAIC-WOZ) dataset [24, 25] comprises voice and text samples from 189 interviewed healthy and control persons and their PHQ-8 depression detection questionnaire.
5 papers · 0 benchmarks
DurLAR (A High-Fidelity 128-Channel LiDAR Dataset with Panoramic Ambient and Reflectivity Imagery)
DurLAR is a high-fidelity 128-channel 3D LiDAR dataset with panoramic ambient (near infrared) and reflectivity imagery for multi-modal autonomous driving applications.
5 papers · 0 benchmarks
The EDT dataset is designed for corporate event detection and text-based stock prediction (trading strategy) benchmark.
5 papers · 0 benchmarks
ELAS is a dataset for lane detection.
5 papers · 0 benchmarks
ESC50 (ESC: Dataset for Environmental Sound Classification)
The ESC-50 dataset is a labeled collection of 2000 environmental audio recordings suitable for benchmarking methods of environmental sound classification.
5 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
5 papers · 0 benchmarks
The Epinions dataset is trust network dataset.
5 papers · 1 benchmark
Europeana Newspapers consists of four datasets with 100 pages each for the languages Dutch, French, German (including Austrian) as part of the Europeana Newspapers project is expected to contribute to the further development and…
5 papers · 0 benchmarks
FAUST-partial is a 3D registration benchmark dataset created to provide a more informative evaluation of 3D registration methods.
5 papers · 9 benchmarks
FFHQ-UV is a large-scale facial UV-texture dataset that contains over 50,000 high-quality texture UV-maps with even illuminations, neutral expressions, and cleaned facial regions, which are desired characteristics for rendering realistic…
5 papers · 0 benchmarks
FLD (Formal Logic Deduction)
A deductive reasoning benchmark based on formal logic theory.
5 papers · 0 benchmarks
FS2K is a high-quality Facial Sketch Synthesis (FSS).
5 papers · 0 benchmarks
FSDKaggle2019 is an audio dataset containing 29,266 audio files annotated with 80 labels of the AudioSet Ontology.
5 papers · 0 benchmarks
FinRL-Meta is universe of market environments for data-driven financial reinforcement learning.
5 papers · 0 benchmarks
The FineWeb dataset consists of more than 15T tokens of cleaned and deduplicated English web data from CommonCrawl.
5 papers · 0 benchmarks
FishEye8K (FishEye8K: A Benchmark and Dataset for Fisheye Camera Object Detection)
With the advance of AI, road object detection has been a prominent topic in computer vision, mostly using perspective cameras.
5 papers · 1 benchmark
A synthetically generated QA dataset for text-based reasoning.
5 papers · 0 benchmarks
The Flick Cropping Dataset consists of high quality cropping and pairwise ranking annotations used to evaluate the performance of automatic image cropping approaches.
5 papers · 0 benchmarks
Contains 8k flickr Images with captions.
5 papers · 2 benchmarks
FluidLab is a simulation environment with a diverse set of manipulation tasks involving complex fluid dynamics.
5 papers · 0 benchmarks
1000 query triples on 120 tables.
5 papers · 0 benchmarks
We construct the ForgeryNet dataset, an extremely large face forgery dataset with unified annotations in image- and video-level data across four tasks: 1) Image Forgery Classification, including two-way (real / fake), three-way (real /…
5 papers · 1 benchmark
FormNLU (Form-NLU: Dataset for the Form Language Understanding)
We introduce a new dataset for form structure understanding and key information extraction.
5 papers · 0 benchmarks
GAD (Gene Associations Database)
GAD, or Gene Associations Database, is a corpus of gene-disease associations curated from genetic association studies.
5 papers · 1 benchmark
GBSG2 (German Breast Cancer Study Group 2)
The German Breast Cancer Study Group (GBSG2) dataset studies the effects of hormone treatment on recurrence-free survival time.
5 papers · 0 benchmarks
GMVD (Generalized Multi-View Detection Dataset)
The GMVD dataset consists of synthetic scenes captured using the GTA-V and Unity graphics engines.
5 papers · 1 benchmark
GOF (Gyroscope Optical Flow)
Optical Flow in challenging scenes with gyroscope readings!
5 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.