Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 165 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 7873–7920 of 12,172
The Cifar10Mnist dataset is created using CIFAR-10 and MNIST data sources.
1 paper · 0 benchmarks
CinePile is a question-answering-based, long-form video understanding dataset.
1 paper · 1 benchmark
Ciona17 is a semantic segmentation dataset with pixel-level annotations pertaining to invasive species in a marine environment.
1 paper · 0 benchmarks
📥 The CipherSpectrum dataset has 4 ZIP files, each containing network traffic data for the same set of 40 domains as follows: - 📦 aes-128-gcm.zip — Network traffic encrypted with TLSAES128GCMSHA256 - 📦 aes-256-gcm.zip — Network traffic…
1 paper · 0 benchmarks
This dataset contains a two-column CSV file, where the first column ("ValidcitingDOI") contains the DOI of a citing entity retrieved in Crossref, while the second column ("InvalidcitedDOI") contains the invalid DOI of a cited entity…
1 paper · 0 benchmarks
This is the full dataset for the paper Fourier neural operator for real-time simulation of 3D dynamic urban microclimate.
1 paper · 0 benchmarks
CityNet is a multi-modal urban dataset containing data from 7 cities, each of which coming from 3 data sources, which can be used for urban computing and smart city research.
1 paper · 0 benchmarks
CityTopia is the largest synthetic dataset for 3D cities with annotations, offering high-fidelity scenes generated using 3D assets from the Unreal Engine 5 CitySample project.
1 paper · 0 benchmarks
- An evaluation test bed for assessing the robustness of sentence embedding models against user-informed misinformation edits.
1 paper · 0 benchmarks
Clarkson Fingerprint Generator consists of a dataset of 50K synthetically generated fingerprints.
1 paper · 0 benchmarks
Clickable heat-map visualizations of the experiments run to quantify the Classic ECN AQM problem and to evaluate the success of the Classic AQM Detection and Fall-back algorithm.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Lang-8 Preprocessed Dataset (for GED): - Dataset: Lang-8, a publicly available dataset containing user-generated content, primarily from second-language learners, focused on writing errors.
1 paper · 0 benchmarks
The Cleft dataset is a collection of ultrasound tongue imaging and audio data, gathered from children with cleft lip and palate by a research speech and language therapist working in a hospital environment.
1 paper · 0 benchmarks
Clickbait PDFs (From Attachments to SEO: Click Here to Learn More about Clickbait PDFs!)
The paper presents a study of Clickbait PDFs, which are PDF documents leading to various attacks on the Web.
1 paper · 0 benchmarks
CloudCast (CloudCast: A Satellite-Based Dataset and Baseline for Forecasting Clouds)
A satellite-based dataset called "CloudCast".
1 paper · 0 benchmarks
The ClueWeb09 dataset was created to support research on information retrieval and related human language technologies.
1 paper · 0 benchmarks
This repository includes the experimental dataset acquired to evaluate our radar-leg odometry algorithm, Co-RaL, accepted by IEEE IROS 2024.
1 paper · 0 benchmarks
Co/FeMn bilayers measured.
1 paper · 0 benchmarks
CoCaHis (Colon Cancer Histology Dataset)
Highlights • Publicly available dataset with 82 H&E stained images of frozen sections.
1 paper · 0 benchmarks
One of the central tasks in software maintenance is being able to understand and develop code changes.
1 paper · 0 benchmarks
CoNECo (Complex Named Entity Corpus)
Complex Named Entity Corpus (CoNECo) is an annotated corpus for NER and NEN of protein-containing complexes.
1 paper · 0 benchmarks
Automatic segmentation, tokenization and morphological and syntactic annotations of raw texts in 45 languages, generated by UDPipe (http://ufal.mff.cuni.cz/udpipe), together with word embeddings of dimension 100 computed from lowercased…
1 paper · 0 benchmarks
CoRAL dataset (CoRAL: a Context-aware Croatian Abusive Language Dataset)
CoRAL is a language and culturally aware Croatian Abusive dataset covering phenomena of implicitness and reliance on local and global context.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
This dataset provides simulated flood inundation maps of Abu Dhabi's coast under 174 different shoreline protection scenarios.
1 paper · 1 benchmark
Dataset used in research submitted to ICPC ERA 2022
1 paper · 0 benchmarks
Code and Data for Replication of "Microsimulation Estimates of Decision Uncertainty and Value of Information Are Biased but Consistent" This is the full data set for replication of all results in the paper along with the R code for doing…
1 paper · 0 benchmarks
The dataset is specifically constructed for the library-oriented code generation task, which are constructed in the paper “CodeGen4Libs: A Two-Stage Approach for Library-Oriented Code Generation”.
1 paper · 0 benchmarks
InstructCoder is the first dataset designed to adapt LLMs for general code editing.
1 paper · 0 benchmarks
A diverse dataset of written code-switched productions, curated from topical threads of multiple bilingual communities on the Reddit discussion platform, and explore questions that were mainly addressed in the context of spoken language…
1 paper · 0 benchmarks
Dataset for machine learning based performance prediction in online coding competitions.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
The data set is based on roughly 6,000 coffee bean review spublished on the website Coffereviews going back to 1997.
1 paper · 0 benchmarks
With two listening files, four multimedia was made.
1 paper · 0 benchmarks
This dataset derives from Coil100.
1 paper · 0 benchmarks
The Collision Avoidance Challenge dataset is the official dataset used during the ESA's Kelvins competition for "Collision Avoidance Challenge".
1 paper · 0 benchmarks
Synthetic graph classification datasets with the task of recognizing the connectivity of same-colored nodes in 4 graphs of varying topology.
1 paper · 0 benchmarks
ColorSVG-100K contains: - 100K samples - 500 categories Project website
1 paper · 0 benchmarks
A large dataset of color names and their respective RGB values stores in CSV.
1 paper · 1 benchmark
ColosseumRL is a framework for research in reinforcement learning in n-player games.
1 paper · 0 benchmarks
Contains correlation data for 119,384 column pairs, taken from 3,952 data sets, including Pearson correlation, Spearman correlation, and Theil's U.
1 paper · 0 benchmarks
ComSum is a data set of 7 million commit messages for text summarization.
1 paper · 0 benchmarks
The ComTQA dataset is a visual table question answering benchmark.
1 paper · 0 benchmarks
The combinatorial 3D shape dataset is composed of 406 instances of 14 classes.
1 paper · 0 benchmarks
Comet is a dataset which contains 11.5k user-assistant dialogs (totalling 103k utterances), grounded in simulated personal memory graphs.
1 paper · 0 benchmarks
Probes to evaluate commonsense in language models.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.