Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 127 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 6049–6096 of 12,172
This provides a benchmark for cyclist's orientation detection, "CIMAT-Cyclist" with bounding box based labels according to eight different classes depending on the orientation.
2 papers · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
CLEAR-Bias (Corpus for Linguistic Evaluation of Adversarial Robustness against Bias)
CLEAR-Bias is a benchmark dataset designed to evaluate the robustness of large language models (LLMs) against bias elicitation, particularly under adversarial conditions.
2 papers · 0 benchmarks
CLOTH3D (CLOTH3D: Clothed 3D Humans)
This work presents CLOTH3D, the first big scale synthetic dataset of 3D clothed human sequences.
2 papers · 0 benchmarks
CLSE (Corpus of Linguistically Significant Entities)
2 papers · 0 benchmarks
CLUENER2020 is a well-defined fine-grained dataset for named entity recognition in Chinese.
2 papers · 0 benchmarks
A benchmark dataset for the co-skeletonization task.
2 papers · 0 benchmarks
The training and validation data are subsets of the training split of the MS COCO dataset (2017 release, bounding boxes only).
2 papers · 0 benchmarks
A novel road corner case dataset for object detection in autonomous driving which contains ~10000 carefully selected road driving scenes with high-quality bounding box annotation for 43 representative road object categories.
2 papers · 0 benchmarks
COFAR (Commonsense and Factual Reasoning in Image Search)
The COFAR (COmmonsense and FActual Reasoning) dataset is a collection of images and text queries specifically designed to challenge and evaluate image search models that aim to go beyond simple visual matching.
2 papers · 1 benchmark
COMPASS-XP is a dataset of matched photographic and X-ray images of single objects, made available for use in Machine Learning & Computer Vision research, in particular in the context of transport security.
2 papers · 0 benchmarks
A dataset for position-constrained robot grasp planning.
2 papers · 0 benchmarks
CORE (Company Relation Extraction)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
A public database of COVID-19 CXR and CT images, which are automatically extracted from COVID-19-relevant articles from the PubMed Central Open Access (PMC-OA) Subset.
2 papers · 0 benchmarks
The researchers of Qatar University have compiled the COVID-QU-Ex dataset, which consists of 33,920 chest X-ray (CXR) images including: 11,956 COVID-19 11,263 Non-COVID infections (Viral or Bacterial Pneumonia) 10,701 Normal Ground-truth…
2 papers · 0 benchmarks
A coronavirus dataset with 98 countries constructed from different reliable sources, where each row represents a country, and the columns represent geographic, climate, healthcare, economic, and demographic factors that may contribute to…
2 papers · 0 benchmarks
A dataset of tweets that reference the COVID-19 pandemic with emotion labels.
2 papers · 0 benchmarks
CPAP (Kang et. al.'s model of adherence behavior for the treatment of obstructive sleep apnea)
Kang et al.'s Markovian model for treatment adherence in obstructive sleep apnea.
2 papers · 0 benchmarks
CPPE-5 (Medical Personal Protective Equipment Dataset)
CPPE - 5 (Medical Personal Protective Equipment) is a new challenging dataset with the goal to allow the study of subordinate categorization of medical personal protective equipments, which is not possible with other popular data sets that…
2 papers · 1 benchmark
The CPU dataset, first introduced by Rahimi and Recht (2007) and then used by Balog et al (2016).
2 papers · 0 benchmarks
The general multi-turn dialogue evaluation dataset with nine topics.
2 papers · 0 benchmarks
Non-contrast head/brain CT of patients with head trauma or stroke symptoms.
2 papers · 0 benchmarks
CQR (Contextual Query Rewrite)
CQR is an extension to the Stanford Dialogue Corpus.
2 papers · 0 benchmarks
CRC100K (100,000 histological images of human colorectal cancer and healthy tissue)
This is a set of 100,000 non-overlapping image patches from hematoxylin & eosin (H&E) stained histological images of human colorectal cancer (CRC) and normal tissue.
2 papers · 0 benchmarks
CREPE is QA dataset containing a natural distribution of presupposition failures from online information-seeking forums.
2 papers · 0 benchmarks
A fundamental characteristic common to both human vision and natural language is their compositional nature.
2 papers · 1 benchmark
CRIPP-VQA (Counterfactual Reasoning about Implicit Physical Properties via Video Question Answering)
CRIPP-VQA is a video question answering dataset for reasoning about the implicit physical properties of objects in a scene.
2 papers · 0 benchmarks
Chinese Spelling Correction Dataset for errors generated by pinyin IME (CSCD-IME), a dataset containing 40,000 annotated sentences from real posts of official media on Sina Weibo.
2 papers · 0 benchmarks
CUHK-SYSU-TBPS is a dataset for text-based person search task.
2 papers · 0 benchmarks
CUP (Context-sitUated Pun) is a dataset containing 4.5k tuples of context words and pun pairs, each labelled with whether they are compatible for composing a pun.
2 papers · 0 benchmarks
CVB (Video Dataset of Cattle Visual Behaviors)
Existing image/video datasets for cattle behavior recognition are mostly small, lack well-defined labels, or are collected in unrealistic controlled environments.
2 papers · 0 benchmarks
The CVL Database is a public database for writer retrieval, writer identification and word spotting.
2 papers · 0 benchmarks
CWD30 (Crop Weed Dataset 30 species)
CWD30 comprises over 219,770 high-resolution images of 20 weed species and 10 crop species, encompassing various growth stages, multiple viewing angles, and environmental conditions.
2 papers · 0 benchmarks
CaBuAr (CaBuAr: California Burned Areas dataset)
This dataset contains images from Sentinel-2 satellites taken before and after a wildfire.
2 papers · 0 benchmarks
CaFFe (CAlving Fronts and where to Find thEm)
The temporal variability in calving front positions of marine-terminating glaciers permits inference on the frontal ablation.
2 papers · 2 benchmarks
CalCROP21 is a georeferenced multi-spectral dataset of satellite Imagery and crop labels.
2 papers · 0 benchmarks
The Calandra dataset provides the data from a pair of tactile sensors attached to a jaw gripper (left and right) alongside the RGB images.
2 papers · 0 benchmarks
The Caltech Cars dataset consists of 126 rear-view photographs captured within parking lots.
2 papers · 1 benchmark
The dataset contains two subsets of synthetic, semantically segmented road-scene images, which have been created for developing and applying the methodology described in the paper "A Sim2Real Deep Learning Approach for the Transformation…
2 papers · 2 benchmarks
CamGes (Cambridge Hand Gesture Dataset)
The size of the data set is about 1GB.
2 papers · 1 benchmark
carecall is a Korean dialogue dataset for role-satisfying dialogue systems.
2 papers · 0 benchmarks
This dataset includes music time information i.e.
2 papers · 0 benchmarks
Catalan TimeBank 1.0 was developed by researchers at Barcelona Media and consists of Catalan texts in the AnCora corpus annotated with temporal and event information according to the TimeML specification language.
2 papers · 1 benchmark
Causal Triplet is a causal representation learning benchmark featuring not only visually more complex scenes, but also two crucial desiderata commonly overlooked in previous works: 1) An actionable counterfactual setting, where only…
2 papers · 0 benchmarks
Celeb-HQ Face Gender Recognition Dataset This dataset is curated for the face gender classification task.
2 papers · 0 benchmarks
Celeb-HQ Facial Identity Recognition Dataset This dataset is curated for the facial identity classification task.
2 papers · 0 benchmarks
Hierarchical multi-label classification dataset for functional genomics
2 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.