Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 164 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 7825–7872 of 12,172
Card is a dataset of playing card images, which consists of 8,029 images with two clusterings, i.e., rank (Ace, King, Queen, etc.) and suits (clubs, diamonds, hearts, spades).
1 paper · 0 benchmarks
A dataset of games played in the card game "Cards Against Humanity" (CAH), by human players, derived from the online CAH labs.
1 paper · 0 benchmarks
The CareerCoach 2022 gold standard is available for download in the NIF and JSON format, and draws upon documents from a corpus of over 99,000 education courses which have been retrieved from 488 different education providers.
1 paper · 0 benchmarks
A large-scale benchmark dataset involving well-labelled datasets to employ the state-of-the-art machine intelligence technologies for map text annotation recognition, map scene classification, map super-resolution reconstruction, and map…
1 paper · 0 benchmarks
Caselaw4 is a dataset of 350k common law judicial decisions from the U.S.
1 paper · 0 benchmarks
Casino Reviews (Online reviews of North American Casinos from Google Reviews)
This dataset contain online reviews gathered from google reviews written by north american casino users.
1 paper · 0 benchmarks
CatalogBank (CatalogBank: A Structured and Interoperable Catalog Dataset for Engineering System Design)
In the realm of document engineering and Natural Language Processing (NLP), the integration of digitally born catalogs into product design processes presents a novel avenue for enhancing information extraction and interoperability.
1 paper · 0 benchmarks
Cattle data set, which was introduced in a paper.
1 paper · 0 benchmarks
SyntaxGym, adapted for interventional interpretability.
1 paper · 1 benchmark
CelebAGaze consists of 25283 high-resolution celebrity images that are collected from CelebA and the Internet.
1 paper · 0 benchmarks
Cell-200 is a a dataset of synthetic fluorescence microscopy images with cell populations generated by SIMCEP.
1 paper · 0 benchmarks
Classifying all cells in an organ is a relevant and difficult problem from plant developmental biology.
1 paper · 1 benchmark
Hyperquack v.2 response data which contains structured data records in JSON.
1 paper · 0 benchmarks
Dataset Overview The dataset used for training and testing consists of vibration signals for six pump conditions: | Fault Type | Training Samples | Testing Samples | Flow Range (L/min) |…
1 paper · 0 benchmarks
The dataset has 93 image stacks and their corresponding Extended Depth of Field (EDF) image acquired from cases with grades Nagative, LSIL or HSIL (The Bethesda System): - Negative: 16 - LSIL: 46 - HSIL: 31 The ground truth includes the…
1 paper · 0 benchmarks
The dataset covers Hindi and Tamil, collected without the use of translation.
1 paper · 1 benchmark
ChCatExt (Chinese Catalog Extraction Dataset)
ChCatExt is composed of BidAnn (bid announcement), FinAnn (financial announcement) and CreRat (credit rating report).
1 paper · 1 benchmark
ChMusic is a traditional Chinese music dataset for training model and performance evaluation of musical instrument recognition.
1 paper · 0 benchmarks
This meta-dataset is first used in the AutoML1 challenge organized by Chalearn in 2015.
1 paper · 1 benchmark
Scene change detection (SCD) dataset tailored for generalizable SCD algorithm.
1 paper · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
We propose ChaosBench, a large-scale, multi-channel, physics-based benchmark for subseasonal-to-seasonal (S2S) climate prediction.
1 paper · 0 benchmarks
This is the training datasets for Character-LLM, which contains nine characters experience data used to train Character-LLMs.
1 paper · 0 benchmarks
Charlotte-ThermalFace is a thermal face dataset.
1 paper · 0 benchmarks
Taking Advice from ChatGPT is a laboratory study of how student participants incorporate advice generated by ChatGPT.
1 paper · 0 benchmarks
Dataset Overview vanilla.csv: Represents the interactions without specific role-play instructions.
1 paper · 0 benchmarks
ChatGPT Software Testing Study Dataset contains questions from a well-known software testing book by Ammann and Offutt.
1 paper · 0 benchmarks
The data can be found in the Data folder, which contains two files: - tickertraindata.json: This file holds the data utilized for training and validation of our model.
1 paper · 0 benchmarks
ChatLog is a coarse-to-fine temporal dataset called ChatLog, consisting of two parts that update monthly and daily: 1.
1 paper · 0 benchmarks
The dataset contains two few-shot chemical fine-grained entity extraction datasets, based on human-annotated ChemNER+ and CHEMET.
1 paper · 0 benchmarks
A novel remote sensing dataset for evaluating a geospatial machine learning model's ability to learn long range dependencies and spatial context understanding.
1 paper · 1 benchmark
The Chess Recognition Dataset (ChessReD) comprises a diverse collection of images of chess formations captured using smartphone cameras; a sensor choice made to ensure real-world applicability.
1 paper · 0 benchmarks
The Chess Recognition Dataset 2K (ChessReD2K) comprises a diverse collection of images of chess formations captured using smartphone cameras; a sensor choice made to ensure real-world applicability.
1 paper · 0 benchmarks
This dataset contains 1125 X-ray images of the studied individuals’ chests, including 125 images labeled as COVID-19, 500 images labeled as pneumonia, and 500 images labeled as no findings.
1 paper · 0 benchmarks
ChiQA is a dataset designed for visual question answering tasks that not only measures the relatedness but also measures the answerability, which demands more fine-grained vision and language reasoning.
1 paper · 0 benchmarks
A large-scale, first-of-its-kind database aimed at generating a better understanding of the way children interact with mobile devices during their development process.
1 paper · 0 benchmarks
ChinaOpen is a new video dataset targeted at open-world multimodal learning, with raw data gathered from Bilibili, a popular Chinese video-sharing website.
1 paper · 1 benchmark
Large-scale Chinese legal dataset for judgment prediction.
1 paper · 0 benchmarks
Chinese Literature NER RE is a Discourse-Level Named Entity Recognition and Relation Extraction Dataset for Chinese Literature Text.
1 paper · 0 benchmarks
The Chinese Traditional Painting dataset for style transfer contains 1000 content images and 100 style images.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Description - Venue: NeurIPS 2024 D&B Spotlight - Repository: Code, Page, Data - Paper: arxiv.org/abs/2406.18522 - Point of Contact: Shenghai Yuan Citation If you find our paper and code useful in your research, please consider giving a…
1 paper · 0 benchmarks
CiNAT Birds 2021 (Cross-View iNaturalist-2021 Birds) dataset contains ground-level images of bird species along with satellite images associated with the geolocation of the ground-level images.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.