Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 97 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 4609–4656 of 12,172
BNaT (Blockchain Network Attack Traffic dataset)
Tran Viet Khoa, Do Hai Son, Dinh Thai Hoang, Nguyen Linh Trung, Tran Thi Thuy Quynh, Diep N.
4 papers · 0 benchmarks
The Bach Doodle Dataset is composed of 21.6 million harmonizations submitted from the Bach Doodle.
4 papers · 0 benchmarks
This dataset includes the beat and downbeat annotations for Beatles albums.
4 papers · 2 benchmarks
Bentham manuscripts refers to a large set of documents that were written by the renowned English philosopher and reformer Jeremy Bentham (1748-1832).
4 papers · 1 benchmark
Botswana is a hyperspectral image classification dataset.
4 papers · 1 benchmark
This brain tumor dataset contains 3064 T1-weighted contrast-enhanced images with three kinds of brain tumor.
4 papers · 0 benchmarks
This self-driving dataset collected in Brno, Czech Republic contains data from four WUXGA cameras, two 3D LiDARs, inertial measurement unit, infrared camera and especially differential RTK GNSS receiver with centimetre accuracy.
4 papers · 0 benchmarks
CADB (Composition Assessment DataBase)
To the best of our knowledge, there is no prior dataset specifically constructed for composition assessment.
4 papers · 1 benchmark
Chinese AI and Law 2019 Similar Case Matching dataset.
4 papers · 0 benchmarks
Large-scale human activity recognition dataset in free-living environment for 151 participants.
4 papers · 0 benchmarks
We create 64-beam LiDAR dataset with settings similar to Velodyne VLP-64 LiDAR on the CARLA simulator.
4 papers · 0 benchmarks
Unsupervised Domain Adaptation demonstrates great potential to mitigate domain shifts by transferring models from labeled source domains to unlabeled target domains.
4 papers · 3 benchmarks
This dataset is an OSN-transmitted (OSN = Online Social Network) version of the CASIA dataset.
4 papers · 1 benchmark
This dataset is an OSN-transmitted (OSN = Online Social Network) version of the CASIA dataset.
4 papers · 1 benchmark
This dataset is an OSN-transmitted (OSN = Online Social Network) version of the CASIA dataset.
4 papers · 1 benchmark
The COVID-19 pandemic raises the problem of adapting face recognition systems to the new reality, where people may wear surgical masks to cover their noses and mouths.
4 papers · 1 benchmark
CAsT-snippets is a high-quality dataset for conversational information seeking containing snippet-level annotations for all queries in the TREC CAsT 2020 and 2022 datasets.
4 papers · 0 benchmarks
CC-DBP is a dataset for knowledge base population research using Common Crawl and DBpedia.
4 papers · 0 benchmarks
CCPE-M (Coached Conversational Preference Elicitation dataset for Movies)
A dataset consisting of 502 English dialogs with 12,000 annotated utterances between a user and an assistant discussing movie preferences in natural language.
4 papers · 0 benchmarks
CCv2 (Casual Conversations v2)
Casual Conversations v2 (CCv2) is composed of over 5,567 participants (26,467 videos) and intended mainly to be used for assessing the performance of already trained models in computer vision and audio applications for the purposes…
4 papers · 0 benchmarks
CDS2K is a benchmark for Concealed scene understanding (CSU), which is a hot computer vision topic aiming to perceive objects with camouflaged properties.
4 papers · 0 benchmarks
CEDAR Signature is a database of off-line signatures for signature verification.
4 papers · 1 benchmark
CrackForest Dataset is an annotated road crack image database which can reflect urban road surface condition in general.
4 papers · 0 benchmarks
CHOCOLATE (Captions Have Often ChOsen Lies About The Evidence)
CHOCOLATE is a benchmark for detecting and correcting factual inconsistency in generated chart captions.
4 papers · 4 benchmarks
CI-MNIST (Correlated and Imbalanced MNIST)
CI-MNIST (Correlated and Imbalanced MNIST) is a variant of MNIST dataset with introduced different types of correlations between attributes, dataset features, and an artificial eligibility criterion.
4 papers · 0 benchmarks
CIDAR contains 10,000 instructions and their output.
4 papers · 0 benchmarks
The quality of AI-generated images has rapidly increased, leading to concerns of authenticity and trustworthiness.
4 papers · 2 benchmarks
CODA-19 is a human-annotated dataset that denotes the Background, Purpose, Method, Finding/Contribution, and Other for 10,966 English abstracts in the COVID-19 Open Research Dataset.
4 papers · 0 benchmarks
COMP6 (COmprehensive Machine-learning Potential)
COMP6 is a benchmark for evaluating the extensibility of machine-learning based molecular potentials.
4 papers · 0 benchmarks
COMPARE is a taxonomy and a dataset of comparison discussions in peer reviews of research papers in the domain of experimental deep learning.
4 papers · 0 benchmarks
MICCAI Challenge on Circuit Reconstruction from Electron Microscopy Images.
4 papers · 1 benchmark
CS1QA is a dataset for code-based question answering in the programming education domain.
4 papers · 0 benchmarks
CSAW-S is a dataset of mammography images which includes expert annotations of tumors and non-expert annotations of breast anatomy and artifacts in the image.
4 papers · 0 benchmarks
CSPubSum is a dataset for summarisation of computer science publications, created by exploiting a large resource of author provided summaries and show straightforward ways of extending it further.
4 papers · 0 benchmarks
CUB-GHA (CUB Gaze-based Human Attention)
CUB-GHA is a dataset for fine-grained classification with human attention annotations.
4 papers · 0 benchmarks
CUGE is a Chinese Language Understanding and Generation Evaluation benchmark with the following features: (1) Hierarchical benchmark framework, where datasets are principally selected and organized with a language capability-task-dataset…
4 papers · 0 benchmarks
A video dataset for recognising traffic signs hosted with the first IEEE Video and Image Processing (VIP) Cup within the IEEE Signal Processing Society.
4 papers · 0 benchmarks
We introduce the Cambridge Law Corpus (CLC), a corpus for legal AI research.
4 papers · 0 benchmarks
The nine (moving camera) videos in this benchmark exhibit camouflaged animals that are difficult to see in a single frame, but can be detected based upon their motion across frames.
4 papers · 1 benchmark
The Cat Facial Landmarks in the Wild (CatFLW) dataset contains 2079 images of cats' faces in various environments and conditions, annotated with 48 facial landmarks and a bounding box on the cat’s face.
4 papers · 1 benchmark
CausalBench is a comprehensive benchmark suite for evaluating network inference methods on large-scale perturbational single-cell gene expression data.
4 papers · 0 benchmarks
The COVID-19 pandemic raises the problem of adapting face recognition systems to the new reality, where people may wear surgical masks to cover their noses and mouths.
4 papers · 1 benchmark
ChemDisGene, a new dataset for training and evaluating multi-class multi-label biomedical relation extraction models.
4 papers · 1 benchmark
Chickenpox Cases in Hungary is a spatio-temporal dataset of weekly chickenpox (childhood disease) cases from Hungary.
4 papers · 0 benchmarks
Children's Song Dataset is open source dataset for singing voice research.
4 papers · 0 benchmarks
The Chilean Waiting List corpus comprises de-identified referrals from the waiting list in Chilean public hospitals.
4 papers · 1 benchmark
The ChineseLP dataset contains 411 vehicle images (mostly of passenger cars) with Chinese license plates (LPs).
4 papers · 1 benchmark
CiteSum is a large-scale scientific extreme summarization benchmark.
4 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.