Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 124 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 5905–5952 of 12,172
AIDA/testc is a new challenging test set for entity linking systems containing 131 Reuters news articles published between December 5th and 7th, 2020.
2 papers · 1 benchmark
AIH is created for hand deocclusion and removal.
2 papers · 0 benchmarks
AIME (AI Music Evaluation Dataset)
The AIME dataset contains 6,000 audio tracks generated by 12 music generation models in addition to 500 tracks from MTG-Jamendo.
2 papers · 0 benchmarks
AI Playground (AIP) is an open-source, Unreal Engine-based tool for generating and labeling virtual image data.
2 papers · 0 benchmarks
Adverbs in Recipes (AIR) is a dataset specifically collected for adverb recognition.
2 papers · 1 benchmark
ALGAD (Andy Lomas Generative Art Dataset)
Repository of a generative art dataset by computer artist Andy Lomas.
2 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
AM2iCo (Adversarial and Multilingual Meaning in Context)
AM2iCo is a wide-coverage and carefully designed cross-lingual and multilingual evaluation set.
2 papers · 0 benchmarks
ANETAC (Arabic Named Entity Transliteration and Classification)
An English-Arabic named entity transliteration and classification dataset built from freely available parallel translation corpora.
2 papers · 0 benchmarks
This is the Current Video sequence set from the AOM-CTC.
2 papers · 0 benchmarks
APE (Automatic Post-Editing)
APE is useful to evaluate Machine Translation automatic post-editing (APE), which is the task of improving the output of a blackbox MT system by automatically fixing its mistakes.
2 papers · 0 benchmarks
The APT Malware dataset is used to train classifiers to predict if a given malware belongs to the “Advanced Persistent Threat” (APT) type or not.
2 papers · 0 benchmarks
ARAS (Action with RAre Scene)
Action with RAre Scene is a small scale dataset collected from Youtube.
2 papers · 0 benchmarks
The Abstraction and Reasoning Corpus (ARC) is a dataset created by François Chollet in 2019.
2 papers · 0 benchmarks
ARC-AGI (Abstraction and Reasoning Corpus for Artificial General Intelligence)
The ARC-AGI benchmark is a significant measure in the field of artificial intelligence, focusing on an AI's general reasoning capabilities.
2 papers · 0 benchmarks
ARCA23K is a dataset of labelled sound events created to investigate real-world label noise.
2 papers · 0 benchmarks
The ARKitFace dataset is established by this work in order to train and evaluate both 3D face shape and 6DoF in the setting of perspective projection.
2 papers · 1 benchmark
The goal of ARQMath is to advance techniques for mathematical information retrieval, in particular, retrieving answers to mathematical questions (Task 1), and formula retrieval (Task 2).
2 papers · 1 benchmark
ASD (Annotated Semantic Dataset)
The Annotated Semantic Dataset is composed of $11$ videos, divided in $3$ activity categories: Biking; Driving and Walking, according to their amount of semantic information.
2 papers · 0 benchmarks
Assessing the value of energy efficiency improvements can be challenging as there's no way to truly know how much energy a building would have used without the improvements.
2 papers · 0 benchmarks
ASIRRA ((Animal Species Image Recognition for Restricting Access)
Web services are often protected with a challenge that's supposed to be easy for people to solve, but difficult for computers.
2 papers · 0 benchmarks
A novel dataset that can support the end-to-end design and running of Online Controlled Experiments (OCE) with adaptive stopping.
2 papers · 0 benchmarks
ASVspoof 5 is the fifth edition in a series of challenges which promote the study of speech spoofing and deepfake attacks as well as the design of detection solutions.
2 papers · 0 benchmarks
AStitchInLanguageModels is a dataset for the exploration of idiomaticity in pre-trained language models.
2 papers · 0 benchmarks
The benchmark ATC-SMILES is built for ATC classification.
2 papers · 1 benchmark
ATLANTIS is a benchmark for semantic segmentation of waterbody images.
2 papers · 1 benchmark
The collected dataset consists of multivariate time series (MTS) data belonging to several ATMs banking along with the annotations that the operators did when they performed a maintenance task on any of the machines.
2 papers · 0 benchmarks
ATPChecker (Automated Third-party library Privacy compliance Checker)
A novel dataset for identifying privacy policy compliance of Android third-party libraries.
2 papers · 0 benchmarks
ATUE is an antibody study benchmark with four real-world supervised tasks covering therapeutic antibody engineering, B cell analysis, and antibody discovery.
2 papers · 0 benchmarks
The AU-AIR is a multi-modal aerial dataset captured by a UAV.
2 papers · 0 benchmarks
AV Digits Database is an audiovisual database which contains normal, whispered and silent speech.
2 papers · 0 benchmarks
AVCAffe (A Large Scale Audio-Visual Dataset of Cognitive Load and Affect for Remote Work)
We introduce AVCAffe, the first Audio-Visual dataset consisting of Cognitive load and Affect attributes.
2 papers · 0 benchmarks
AVMIT (Audiovisual Moments in Time)
Audiovisual Moments in Time (AVMIT) is a large-scale dataset of audiovisual action events.
2 papers · 0 benchmarks
AVSync15 is a high-quality synchronized audio-video dataset curated from VGGSound.
2 papers · 0 benchmarks
AWARE (AWARE: Aspect-Based Sentiment Analysis Dataset of Apps Reviews for Requirements Elicitation)
The peer-reviewed paper of AWARE dataset is published in ASEW 2021, and can be accessed through: http://doi.org/10.1109/ASEW52652.2021.00049.
2 papers · 3 benchmarks
We present the AWS documentation corpus, an open-book QA dataset, which contains 25,175 documents along with 100 matched questions and answers.
2 papers · 0 benchmarks
Fanpage dataset, containing news articles taken from Fanpage.
2 papers · 1 benchmark
IlPost dataset, containing news articles taken from IlPost.
2 papers · 1 benchmark
ActionBench contains two carefully designed probing tasks: Action Antonym and Video Reversal, which targets multimodal alignment capabilities and temporal understanding skills of the model, respectively.
2 papers · 0 benchmarks
Measurement data related to the publication „Active TLS Stack Fingerprinting: Characterizing TLS Server Deployments at Scale“.
2 papers · 0 benchmarks
Contains ten synthetic time series with five days of high activity and two days of low activity.
2 papers · 0 benchmarks
AdvNet is a dataset of traffic signs images.
2 papers · 0 benchmarks
An exhaustive list of stop lemmas created from 12 corpora across multiple domains, consisting of over 13 million words, from which more than 200,000 lemmas were generated, and 11 publicly available stop word lists comprising over 1000…
2 papers · 0 benchmarks
A set of synthetic MNIST-style datasets for four orthographies used in Afro-Asiatic and Niger-Congo languages: Geez (Ethiopic), Vai, Osmanya, and N'Ko.
2 papers · 0 benchmarks
The dataset contains historical financial transactions, including time, category and cost fields.
2 papers · 1 benchmark
Almawave-SLU is the first Italian dataset for Spoken Language Understanding (SLU).
2 papers · 0 benchmarks
Alsat-2B is a remote sensing dataset of low and high spatial resolution images (10m and 2.5m respectively) for the single-image super-resolution task.
2 papers · 0 benchmarks
Amazon MTPP (Marked Temporal Point Processes on Amazon data)
The dataset includes time-stamped user product reviews behavior from January, 2008 to October, 2018.
2 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.