Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 161 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 7681–7728 of 12,172
a new self-annotated CC-ReID dataset named Cloth-Changing Unreal Person.
1 paper · 0 benchmarks
CD-HARD comprises 102 images featuring vehicles with oblique license plates sourced from the Cars dataset.
1 paper · 0 benchmarks
CD18 (Cellphone Dataset with 18 Features)
1 paper · 1 benchmark
CDC fluview (National, Regional, and State Level Outpatient Illness and Viral Surveillance)
country- and state-level historical ILI data from 2010 to 2018 from the CDC (CDC).
1 paper · 0 benchmarks
Given the difficulty to handle planetary data we provide downloadable files in PNG format from the missions Chang'E-3 and Chang'E-4.
1 paper · 0 benchmarks
CECW (Colorful Extended Cleanup World)
The CECW dataset is a color-extended version of the Cleanup World (CW) borrowed from the mobile-manipulation robot domain.
1 paper · 0 benchmarks
The dataset includes annotations for burned area delineation and land cover segmentation, with a focus on European soil.
1 paper · 1 benchmark
CEREC (Corpus for Entity Resolution in Email Conversations)
CEREC is a large scale corpus for entity resolution in email conversations.
1 paper · 0 benchmarks
This repo contains open-source channel measurement data for research and development purposes.
1 paper · 0 benchmarks
CFC-DAOD (Caltech Fish Counting – Domain Adaptive Object Detection)
CFC-DAOD is a domain adaptation extension to the Caltech Fish Counting domain generalization benchmark.
1 paper · 1 benchmark
- CFEVER is a Chinese Fact Extraction and VERification dataset published at AAAI 2024.
1 paper · 0 benchmarks
CGHD1152 (Circuit Graph Hand Drawn 1152)
- 1152 Images - 144 Circuits - 12 Drafter - 48,563 Object (Symbol, Structural, Text) Annotations
1 paper · 0 benchmarks
An image sequence dataset of growing snowflakes in HDF5 format.
1 paper · 0 benchmarks
CHAMP (Concept and Hint-Annotated Math Problems)
The Concept and Hint-Annotated Math Problems (CHAMP) consists of high school math competition problems, annotated with concepts, or general math facts, and hints, or problem-specific tricks.
1 paper · 0 benchmarks
CHIP Clinical Trial Classification, a dataset aimed at classifying clinical trials eligibility criteria, which are fundamental guidelines of clinical trials defined to identify whether a subject meets a clinical trial or not, is used for…
1 paper · 1 benchmark
CHORD (CHOrus Recognition Dataset)
CHORD is the first chorus recognition dataset containing 627 songs for public use.
1 paper · 0 benchmarks
CI-ToD is a dataset for Consistency Identification in Task-oriented Dialog system.
1 paper · 0 benchmarks
Collected more than 10,854 samples (4,354 malware and 6,500 benign) from several sources.
1 paper · 0 benchmarks
The CICIoMT2024 dataset is a comprehensive dataset designed for cybersecurity research focused on the Internet of Medical Things (IoMT).
1 paper · 0 benchmarks
CIP (Complete Inertial Pose)
The CIP dataset is composed of 2 subsets, containing low-cost (MPU9250) and high-end (MTwAwinda) Magnetic, Angular Rate, and Gravity (MARG) sensor data respectively.
1 paper · 0 benchmarks
The main goal of the Continuous Integration of Performance Models (CIPM) is to enable an accurate architecture-based performance prediction at each point of the systems development life cycle.
1 paper · 0 benchmarks
Description This repository includes the experiment results, source code, and test data for Three Cs risk inference, using the CIRO (COVID-19 Infection Risk Ontology) and HermiT.
1 paper · 0 benchmarks
CISOL (Construction Industry Steel Ordering Lists Dataset)
The Construction Industry Steel Ordering Lists (CISOL) dataset comprises table-centric, real-world documents from the construction industry, annotated to facilitate the testing and training of deep learning models for table detection (TD)…
1 paper · 2 benchmarks
CITR Dataset consists of experimentally designed fundamental VCI scenarios (front, back, and lateral VCIs) and provides unique ID for each pedestrian, which is suitable for exploring a specific aspect of VCI.
1 paper · 0 benchmarks
CKBP v2 is a new CSKB Population benchmark, which addresses the two mentioned problems by using experts instead of crowd-sourced annotation and by adding diversified adversarial samples to make the evaluation set more representative.
1 paper · 0 benchmarks
CLAD (Complex and Long Activities Dataset)
CLAD (Compled and Long Activities Dataset) is an activity dataset which exhibits real-life and diverse scenarios of complex, temporally-extended human activities and actions.
1 paper · 0 benchmarks
CLCXray (Cutters and Liquid Containers X-ray Dataset)
The CLCXray dataset contains 9,565 X-ray images, in which 4,543 X-ray images (real data) are obtained from the real subway scene and 5,022 X-ray images (simulated data) are scanned from manually designed baggage.
1 paper · 1 benchmark
CLEVR Mental Rotation Tests (CLEVR-MRT) is a new version of the CLEVR dataset.
1 paper · 0 benchmarks
CLIPS (Corpora e Lessici dell'Italiano Parlato e Scritto)
CLIPS, ovvero Corpora e Lessici dell'Italiano Parlato e Scritto, è uno degli otto progetti (Progetto n.
1 paper · 0 benchmarks
The CLOUD dataset is a set of Optical Coherence Tomography of the Anterior Segment images (AS-OCT) used to the automatic identification and representation of the cornea-contact lens relationship.
1 paper · 0 benchmarks
CLPD (China License Plate Dataset)
The CLPD dataset comprises 1200 images that encompass various regions within mainland China.
1 paper · 0 benchmarks
CLUES (Constrained Language Understanding Evaluation Standard)
CLUES (Constrained Language Understanding Evaluation Standard) is a benchmark for evaluating the few-shot learning capabilities of NLU models.
1 paper · 0 benchmarks
CLUES is a benchmark for Classifier Learning Using natural language ExplanationS, consisting of a range of classification tasks over structured data along with natural language supervision in the form of explanations.
1 paper · 0 benchmarks
CLaRO is a new dataset of 234 Competency Questions that had been processed automatically into 106 patterns.
1 paper · 0 benchmarks
CMACD (Chinese Multi-label Affective Computing Dataset)
This study collected data from the major social media platform Weibo, screening 11,338 valid users from over 50,000 individuals with diverse MBTI personality labels and acquiring 566,900 posts along with the user MBTI personality tags.
1 paper · 0 benchmarks
CMCNC (Coherent Multiple Choice Narrative Cloze)
The Coherent Multiple Choice Narrative Cloze (CMCNC) dataset is an evaluation dataset for the multi-choice narrative cloze task, where the goal is to distinguish which event has been held out from a document from a small set of randomly…
1 paper · 0 benchmarks
The field of biomechanics is at a turning point, with marker-based motion capture set to be replaced by portable and inexpensive hardware, rapidly improving markerless tracking algorithms, and open datasets that will turn these new…
1 paper · 0 benchmarks
CMWD (Cloud Motion Wind Dataset)
CMWD (Cloud Motion Wind Dataset) is the first cloud motion wind dataset for deep learning research.
1 paper · 0 benchmarks
CMeIE (Chinese Medical Information Extraction Dataset)
Chinese Medical Information Extraction, a dataset that is also released in CHIP2020, is used for CMeIE task.
1 paper · 1 benchmark
Cmedia dataset: This dataset consists of 200 Youtube links of pop songs (most of them are Chinese songs), together with their groundtruth files of vocal transcription.
1 paper · 0 benchmarks
CN-Celeb-AV is a multi-genre AVPR dataset collected 'in the wild'.
1 paper · 0 benchmarks
Contains a dataset of 241 Chinese dishes with 191,811 images.
1 paper · 0 benchmarks
CNFOOD-241 Contains a dataset of 241 Chinese dishes with 191,811 images.
1 paper · 1 benchmark
Dataset for the Paper "Adversarial Robustness through the Lens of Convolutional Filters".
1 paper · 0 benchmarks
COAT (CommonSense Object Affordance Task)
Useful for checking the physical reasoning capabilities in household agents.
1 paper · 0 benchmarks
COCO Earthquake is a dataset similar to Common Objects in Context (COCO) used for cracking segmentation.
1 paper · 0 benchmarks
COCO-MEBOW (Monocular Estimation of Body Orientation In the Wild)
COCO-MEBOW (Monocular Estimation of Body Orientation in the Wild) is a new large-scale dataset for orientation estimation from a single in-the-wild image.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.