Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 153 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 7297–7344 of 12,172
AOMIC (Amsterdam Open MRI Collection)
The Amsterdam Open MRI Collection (AOMIC) is a collection of three datasets with multimodal (3T) MRI data including structural (T1-weighted), diffusion-weighted, and (resting-state and task-based) functional BOLD MRI data, as well as…
1 paper · 0 benchmarks
AP (Adversarial Paraphrase)
This is a paraphrasing dataset created using the adversarial paradigm.
1 paper · 1 benchmark
APIBench is a benchmark dataset designed for evaluating the performance of API recommendation approaches.
1 paper · 0 benchmarks
APND (Arm Point Nav Dataset)
APND (Arm Point Nav Dataset) is a dataset for the generalizable object manipulation task called ARMPOINTNAV, which consists on moving an object in the scene from a source location to a target location.
1 paper · 0 benchmarks
A database of 56 high quality fabric material measurements, provided as carefully calibrated rectified HDR images, together with SVBRDF fits.
1 paper · 0 benchmarks
We present the APRICOT-Mask dataset, which augments the APRICOT dataset with pixel-level annotations of adversarial patches.
1 paper · 0 benchmarks
The Archive Query Log (AQL) is a previously unused, comprehensive query log collected at the Internet Archive over the last 25 years.
1 paper · 0 benchmarks
ARC Ukiyo-e Faces is a large-scale (>10k paintings, >20k faces) Ukiyo-e dataset with coherent semantic labels and geometric annotations through augmenting and organizing existing datasets with automatic detection.
1 paper · 0 benchmarks
The ARC-100 dataset was collected as part of a prototype retail checkout system titled ARC (Automatic Retail Checkout).
1 paper · 0 benchmarks
ARCT (Argument Reasoning Comprehension Task)
Freely licensed dataset with warrants for 2k authentic arguments from news comments.
1 paper · 0 benchmarks
ARIA (Automated Retinal Image Analysis (ARIA) Data Set)
This data set was collected in 2004 to 2006 in the United Kingdom.
1 paper · 0 benchmarks
This page contains ARINC 429 message data recorded from the hardware-in-a-loop simulator.
1 paper · 0 benchmarks
We complement ARKitScenes dataset with dense semantic annotations that are automatically generated at scale.
1 paper · 0 benchmarks
ARKitTrack is a new RGB-D tracking dataset for both static and dynamic scenes captured by consumer-grade LiDAR scanners equipped on Apple's iPhone and iPad.
1 paper · 0 benchmarks
Around 90k different RL runs: 256 hyperparameter configurations for 10 seeds each across a total of 3 algorithms (PPO, SAC, DQN) and 22 environments.
1 paper · 0 benchmarks
ARTE (Ambisonics Recordings of Typical Environments)
The ARTE database, so far, contains 13 acoustic environments that were recorded with a purpose-built 62-channel microphone array in various locations around Sydney (Australia), and was decoded into the higher-order Ambisonics (HOA) format.
1 paper · 0 benchmarks
ARVSU (Addressee Recognition in Visual Scenes with Utterances)
ARVSU contains a vast body of image variations in visual scenes with an annotated utterance and a corresponding addressee for each scenario.
1 paper · 0 benchmarks
ASCAD (ANSSI SCA Database) is a set of databases that aims at providing a benchmarking reference for the SCA community: the purpose is to have something similar to the MNIST database that the Machine Learning community has been using for…
1 paper · 0 benchmarks
ASCAD database version 2.
1 paper · 0 benchmarks
ASLG-PC12 (English-ASL Gloss Parallel Corpus 2012)
An artificial corpus built using grammatical dependencies rules due to the lack of resources for Sign Language.
1 paper · 1 benchmark
ASLLVD (American Sign Language Lexicon Video Dataset)
Extremely important: The ASLLVD video data are subject to Terms of Use: http://www.bu.edu/asllrp/signbank-terms.pdf.
1 paper · 0 benchmarks
ASOS Data (Automated Surface/Weather Observing Systems (ASOS/AWOS) Data)
The Automated Surface Observing Systems (ASOS) program is a joint effort of the National Weather Service (NWS), the Federal Aviation Administration (FAA), and the Department of Defense (DOD).
1 paper · 1 benchmark
A Rich Annotated Mandarin Conversational (RAMC) Speech Dataset, including 180 hours of Mandarin Chinese dialogue, 150, 10 and 20 hours for the training set, development set and test set respectively.
1 paper · 0 benchmarks
ASRD (Anime Style Recognition Dataset)
A well-labeled challenging dataset, to facilitate the research on style recognition on anime images by collecting images from 190 anime and cartoon works covering 93 years from 13 countries and regions, 2D and 3D work into consideration…
1 paper · 0 benchmarks
ASyMOB (ASyMOB: Algebraic Symbolic Mathematical Operations Benchmark)
ASyMOB (pronounced Asimov, in tribute to the renowned author), is a novel assessment framework focused exclusively on symbolic manipulation, featuring 17,092 unique math challenges, organized by similarity and complexity.
1 paper · 0 benchmarks
ATC-GRAPH is the most extensive ATC benchmark dataset.
1 paper · 1 benchmark
ATD-Dataset (Auto-Tune Detection Dataset (ATD-Dataset))
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Multimodal object recognition is still an emerging field.
1 paper · 0 benchmarks
AUR & UMB dataset (Anticancer Efficacy of Auraptene & Umbelliprenin: In Vitro Viability Dataset)
This dataset contains quantitative data on the anticancer effects of the natural coumarins Auraptene (AUR) and Umbelliprenin (UMB) across 27 studies.
1 paper · 1 benchmark
AUT-VI (Amirkabir campus dataset)
AUT-VI is a super-challenging visual inertial dataset with 126 diverse sequences in 17 locations.
1 paper · 0 benchmarks
Science Journal for Kids Data This repository contains a dataset of abstracts from the Science Journal for Kids website and the original academic papers.
1 paper · 0 benchmarks
Science Journal for Kids Data This repository contains a dataset of abstracts from the Science Journal for Kids website and the original academic papers.
1 paper · 0 benchmarks
This is a processed dataset comprising the interactions of automated vehicles and human-driven vehicles at unsignalized intersections, extracted from the Waymo Open Motion Dataset and Lyft Level 5 Dataset.
1 paper · 0 benchmarks
AVASpeech-SMAD (AVASpeech-SMAD: A Strongly Labelled Speech and Music Activity Detection Dataset with Label Co-Occurrence)
We propose a dataset, AVASpeech-SMAD, to assist speech and music activity detection research.
1 paper · 0 benchmarks
A dataset for audio-visual event classification and localization in the context of office environments.
1 paper · 0 benchmarks
Aircraft collision avoidance systems rely on sensor information to detect and track intruding aircraft so that they may issue proper collision avoidance advisories.
1 paper · 0 benchmarks
Provided in the linked paper.
1 paper · 0 benchmarks
AViMoS (Audio-Visual Mouse Saliency)
A novel audio-visual mouse saliency (AViMoS) dataset with the following key-features: Diverse content: movie, sports, live, vertical videos, etc.; Large scale: 1500 videos with mean 19s duration; High resolution: all streams are FullHD;…
1 paper · 0 benchmarks
The existing multi-modality image fusion dataset lacks comprehensive coverage of adverse weather scenarios.
1 paper · 0 benchmarks
The AbstRCT dataset consists of randomized controlled trials retrieved from the MEDLINE database via PubMed search.
1 paper · 2 benchmarks
A benchmark evaluating LLMs’ abstention: the skill of knowing when NOT to answer!
1 paper · 0 benchmarks
The dataset contains 7,601 Gab posts classified on three different aspects: abuse presence or not, abuse severity and abuse target.
1 paper · 0 benchmarks
Accidental Turntables contains a challenging set of 41,212 images of cars in cluttered backgrounds, motion blur and illumination changes that serves as a benchmark for 3D pose estimation.
1 paper · 0 benchmarks
Accompnaying Dataset for: Chemical Heredity as Group Selection at the Molecular Level.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Yavuz Selim TASPINAR, Murat KOKLU and Mustafa ALTIN Citation Request : 1: KOKLU M., TASPINAR Y.S., (2021).
1 paper · 0 benchmarks
Hahmann, Manuel; Verburg Riezu, Samuel Arturo (2021): Acoustic frequency responses in a conventional classroom.
1 paper · 0 benchmarks
AcousticRooms is a large-scale synthetic room impulse response (RIR) dataset designed for cross-room RIR prediction tasks.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.