Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 81 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 3841–3888 of 12,172
The AQUAINT Corpus consists of newswire text data in English, drawn from three sources: the Xinhua News Service (People's Republic of China), the New York Times News Service, and the Associated Press Worldstream News Service.
6 papers · 1 benchmark
ARCADE (Automatic Region-based Coronary Artery Disease diagnostics using x-ray angiography imagEs Dataset)
ARCADE: Automatic Region-based Coronary Artery Disease diagnostics using x-ray angiography imagEs Dataset Phase 2 consist of two folders with 300 images in each of them as well as annotations.
6 papers · 0 benchmarks
ASCEND (A Spontaneous Chinese-English Dataset) introduces a high-quality resource of spontaneous multi-turn conversational dialogue Chinese code-switching corpus collected in Hong Kong.
6 papers · 0 benchmarks
AcinoSet is a dataset of free-running cheetahs in the wild that contains 119,490 frames of multi-view synchronized high-speed video footage, camera calibration files and 7,588 human-annotated frames.
6 papers · 0 benchmarks
ARID is a dataset for action recognition in dark videos.
6 papers · 0 benchmarks
This data was extracted from the 1994 Census bureau database by Ronny Kohavi and Barry Becker (Data Mining and Visualization, Silicon Graphics).
6 papers · 1 benchmark
AnoShift (AnoShift: A Distribution Shift Benchmark for Unsupervised Anomaly Detection)
AnoShift is a large-scale anomaly detection benchmark, which focuses on splitting the test data based on its temporal distance to the training set, introducing three testing splits: IID, NEAR, and FAR.
6 papers · 1 benchmark
ArCOV-19 is an Arabic COVID-19 Twitter dataset that covers the period from 27th of January till 30th of April 2020.
6 papers · 0 benchmarks
ArmanEmo is a human-labeled emotion dataset of more than 7000 Persian sentences labeled for seven categories.
6 papers · 1 benchmark
AtariARI (Atari Annotated RAM Interface)
The AtariARI (Atari Annotated RAM Interface) is an environment for representation learning.
6 papers · 0 benchmarks
BAAI-VANJEE is a dataset for benchmarking and training various computer vision tasks such as 2D/3D object detection and multi-sensor fusion.
6 papers · 0 benchmarks
A dataset with a single banking domain, includes both general Out-of-Scope (OOD-OOS) queries and In-Domain but Out-of-Scope (ID-OOS) queries, where ID-OOS queries are semantically similar intents/queries with in-scope intents.
6 papers · 1 benchmark
This dataset was created using a dataset used for data categorization that onsists of 2225 documents from the BBC news website corresponding to stories in five topical areas from 2004-2005 used in the paper of D.
6 papers · 0 benchmarks
BC4CHEMD (BioCreative IV Chemical compound and drug name recognition)
Introduced by Krallinger et al.
6 papers · 1 benchmark
BIPIA (Benchmark of Indirect Prompt Injection Attacks)
Recent advancements in large language models (LLMs) have led to their adoption across various applications, notably in combining LLMs with external content to generate responses.
6 papers · 0 benchmarks
BMELD is a bilingual (English-Chinese) dialogue corpus for Neural chat translation.
6 papers · 0 benchmarks
Dataset for multi-class cell classification in breast cancer H\&E images using dot annotations .
6 papers · 0 benchmarks
BSARD (Belgian Statutory Article Retrieval Dataset)
The Belgian Statutory Article Retrieval Dataset (BSARD) is a French native corpus for studying statutory article retrieval.
6 papers · 1 benchmark
Introduces three datasets of expressing hate, commonly used topics, and opinions for hate speech detection, document classification, and sentiment analysis, respectively.
6 papers · 0 benchmarks
BiPaR is a manually annotated bilingual parallel novel-style machine reading comprehension (MRC) dataset, developed to support monolingual, multilingual and cross-lingual reading comprehension on novels.
6 papers · 0 benchmarks
The Bimanual Actions Dataset is a collection of 540 RGB-D videos, showing subjects perform bimanual actions in a kitchen or workshop context.
6 papers · 0 benchmarks
Multimodal Brain Tumor Segmentation Challenge 2018
6 papers · 0 benchmarks
BreizhCrops is a satellite image time series dataset for crop type classification.
6 papers · 0 benchmarks
- An underwater dataset collected from several field trials within the EU FP7 project “Cognitive autonomous diving buddy (CADDY)”, where an Autonomous Underwater Vehicle (AUV) was used to interact with divers and monitor their activities.
6 papers · 0 benchmarks
CAIS (Chinese Artificial Intelligence Speakers)
We collect utterances from the Chinese Artificial Intelligence Speakers (CAIS), and annotate them with slot tags and intent labels.
6 papers · 2 benchmarks
CMD is a publicly available collection of hundreds of thousands 2D maps and 3D grids containing different properties of the gas, dark matter, and stars from more than 2,000 different universes.
6 papers · 0 benchmarks
CCMixter is a singing voice separation dataset consisting of 50 full-length stereo tracks from ccMixter featuring many different musical genres.
6 papers · 0 benchmarks
CELLS is a large (63k pairs) and broadest-ranging (12 journals) parallel corpus for lay language generation.
6 papers · 0 benchmarks
CGIQA-6K (Computer Graphics Image Quality Assessment)
CGIQA-6k database is a large-scale, in-the-wild CGIQA database consisting of 6,000 CGIs.
6 papers · 0 benchmarks
The CHB-MIT dataset is a dataset of EEG recordings from pediatric subjects with intractable seizures.
6 papers · 1 benchmark
CITE is a crowd-sourced resource for multimodal discourse: this resource characterises inferences in image-text contexts in the domain of cooking recipes in the form of coherence relations.
6 papers · 1 benchmark
CLAMS (Cross-linguistic Analysis of Models on Syntax)
Targeted syntactic evaluation datasets in 5 languages: English, French, German, Russian, and Hebrew.
6 papers · 0 benchmarks
CLUECorpus2020 is a large-scale corpus that can be used directly for self-supervised learning such as pre-training of a language model, or language generation.
6 papers · 0 benchmarks
Comprises about 40,000 images where the most suitable objects for 14 tasks have been annotated.
6 papers · 0 benchmarks
A dataset of 12-lead ECGs with annotations.
6 papers · 1 benchmark
A large challenging dataset, COUGH, for COVID-19 FAQ retrieval.
6 papers · 0 benchmarks
Median house prices for California districts derived from the 1990 census.
6 papers · 2 benchmarks
CeyMo is a novel benchmark dataset for road marking detection which covers a wide variety of challenging urban, sub-urban and rural road scenarios.
6 papers · 1 benchmark
ChiMed-VL Dataset ChiMed-VL-Alignment dataset ## ChiMed-VL-Alignment consists of 580,014 image-text couplings, each pair falling into one of two categories: context information of an image or descriptions of an image.
6 papers · 0 benchmarks
ChineseFoodNet aims to automatically recognizing pictured Chinese dishes.
6 papers · 0 benchmarks
ClipShots is a large-scale dataset for shot boundary detection collected from Youtube and Weibo covering more than 20 categories, including sports, TV shows, animals, etc.
6 papers · 1 benchmark
The ClonedPerson dataset is a large-scale synthetic person re-identification dataset introduced in the paper "Cloning Outfits from Real-World Images to 3D Characters for Generalizable Person Re-Identification" in CVPR 2022.
6 papers · 4 benchmarks
The Color Dataset (CoDa) is a probing dataset to evaluate the representation of visual properties in language models.
6 papers · 0 benchmarks
The CoarseWSD-20 dataset is a coarse-grained sense disambiguation dataset built from Wikipedia (nouns only) targeting 2 to 5 senses of 20 ambiguous words.
6 papers · 0 benchmarks
This is a dataset with spurious correlations which can be used to evaluate machine learning methods for out-of-distribution generalization, causal inference, and related field.
6 papers · 1 benchmark
Common Phone is a gender-balanced, multilingual corpus recorded from more than 76.000 contributors via Mozilla's Common Voice project.
6 papers · 0 benchmarks
Concepticon (Concepticon. A Resource for the Linking of Concept Lists)
This resource, our Concepticon, links concept labels from different conceptlists to concept sets.
6 papers · 0 benchmarks
ConvoSumm is a suite of four datasets to evaluate a model’s performance on a broad spectrum of conversation data.
6 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.