Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 155 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 7393–7440 of 12,172
This dataset is suitable for the image classification model.
1 paper · 0 benchmarks
Official repository for the AnnoMI dataset: the first public collection of expert-annotated MI transcripts.
1 paper · 0 benchmarks
AnnoPage Dataset contains pages of mostly historical documents and annotations of non-textual objects, such as images, diagrams, symbols, initials, etc.
1 paper · 0 benchmarks
Includes two datasets for this task, one for English-French (En-Fr) and another for English-German (En-De).
1 paper · 0 benchmarks
AnswerSumm is a dataset of 4,631 CQA threads for answer summarization, curated by professional linguists.
1 paper · 0 benchmarks
Antibody Watch is a dataset of text snippets extracted from over 2000 PubMed articles with annotations denoting specificity of antibodies.
1 paper · 0 benchmarks
Contains a large number of online videos and subtitles.
1 paper · 0 benchmarks
The Apiza Corpus is a WoZ-like (Wizard of Oz) set of dialogues between 30 programmers and a simulated virtual assistant.
1 paper · 0 benchmarks
The Inpainting dataset consists of synchronized Labeled image and LiDAR scanned point clouds.
1 paper · 1 benchmark
Dataset used for the paper entitled "Towards a Fair Comparison and Realistic Evaluation Framework of Android Malware Detectors based on Static Analysis and Machine Learning".
1 paper · 0 benchmarks
The AppealCase dataset is the first large-scale resource specifically designed to support LegalAI research in appellate judgment scenarios.
1 paper · 0 benchmarks
The Apron Dataset focuses on training and evaluating classification and detection models for airport-apron logistics.
1 paper · 0 benchmarks
This dataset contains 369 images of Trash used for deep learning.
1 paper · 1 benchmark
ArEEGChars, the first EEG dataset for Arabic characters, consists of high-quality recordings for 31 unique characters from 30 participants (21 males and 9 females) using the Epoc X 14-channel device.
1 paper · 0 benchmarks
ArEEGWords dataset is a novel EEG dataset recorded from 22 participants with mean age of 22 years (5 female, 17 male) using a 14-channel Emotiv Epoc X device.
1 paper · 0 benchmarks
Sentiment analysis is pivotal in Natural Language Processing for understanding opinions and emotions in text.
1 paper · 0 benchmarks
ArVoice (ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis)
We introduce ArVoice, a multi-speaker Modern Standard Arabic (MSA) speech corpus with diacritized transcriptions, intended for multi-speaker speech synthesis, and can be useful for other tasks such as speech-based diacritic restoration,…
1 paper · 0 benchmarks
AraCovid19-SSD is a manually annotated Arabic COVID-19 sarcasm and sentiment detection dataset containing 5,162 tweets.
1 paper · 0 benchmarks
This dataset is designed to help train simple machine learning models that serve educational and research purposes in the speech recognition domain, mainly for keyword spotting tasks.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Arabic-ToD (Arabic-ToD: Arabic Task Oriented Dialogue dataset)
The Arabic-TOD dataset is based on the BiToD dataset.
1 paper · 0 benchmarks
Follow the instructions provided in the companion repo to automatically download and decompress the archive.
1 paper · 0 benchmarks
Digital Edition: Essays from Hannah Arendt We have created a NER dataset from the digital edition "Sechs Essays" by Hannah Arendt.
1 paper · 0 benchmarks
ArgSciChat is an argumentative dialogue dataset.
1 paper · 2 benchmarks
A small-scale, real-world Project Aria dataset with high quality static 3D oriented bounding boxs annotations.
1 paper · 1 benchmark
Aria Scenes is a benchmark dataset for future research on photorealistic reconstruction.
1 paper · 0 benchmarks
The Aristo Tuple KB contains a collection of high-precision, domain-targeted (subject,relation,object) tuples extracted from text using a high-precision extraction pipeline, and guided by domain vocabulary constraints.
1 paper · 1 benchmark
This dataset contains 2,360 paraphrases in Armenian that can be used for paraphrase detection.
1 paper · 0 benchmarks
ArtDL is a novel painting data set for iconography classification composed of images collected from online sources.
1 paper · 1 benchmark
The ArtFID dataset contains around 250k labeled artworks.
1 paper · 0 benchmarks
ArtImage is a synthetic dataset of articulated object models of 5 categories from PartNet-Mobility for articulated object tasks in category level.
1 paper · 0 benchmarks
The task of Visual Question Answering (VQA) has been studied extensively on general-domain real-world images.
1 paper · 1 benchmark
Article-Bias-Prediction Dataset The articles crawled from www.allsides.com are available in the ./data folder, along with the different evaluation splits.
1 paper · 0 benchmarks
Checkpoints, generated EMA representations, audio outputs, and annotations for paper titled "Articulation GAN: Unsupervised modeling of articulatory learning"
1 paper · 0 benchmarks
These images consist of a series of bacteria of the type Bacillus Subtilis that are suspended and captured by a digital microscope.
1 paper · 0 benchmarks
This is a set of signals-pairs, univariate and multivariate, that can be used to test alignment algorithms.
1 paper · 0 benchmarks
A dataset to enable automatic academic paper rating.
1 paper · 0 benchmarks
AsEP (Antibody-specific Epitope Prediction)
AsEP is a protein structure dataset that includes 1723 filtered antibody-antigen complexes from abYbank/AbDb.
1 paper · 0 benchmarks
A mapping of Ascent to the relations of ConceptNet.
1 paper · 0 benchmarks
AskParents is a dataset for advice classification extracted from Reddit.
1 paper · 0 benchmarks
AskUbuntu question dataset is a preprocessed collection of questions taken from the AskUbuntu.com 2014 corpus dump.
1 paper · 0 benchmarks
A radar-centric automotive dataset based on radar, lidar and camera data for the purpose of 3D object detection.
1 paper · 0 benchmarks
AtyPict is a dataset of atypical sketch content designed for atypical sketch content detection tasks.
1 paper · 0 benchmarks
Dataset Description: The dataset comprises audio recordings of the wing beats of Aedes aegypti mosquitoes and others, conducted in a semi-controlled environment.
1 paper · 0 benchmarks
Audio files that supplement "Treatise on Hearing: The Temporal Auditory Imaging Theory Inspired by Optics and Communication".
1 paper · 0 benchmarks
Test dataset for unsupervised anomaly detection in sound (ADS).
1 paper · 0 benchmarks
The dataset utilized for this study is the Wine Quality dataset, which comprises 1,599 rows and 11 features related to the chemical properties of wine samples.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.