Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 162 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 7729–7776 of 12,172
COCO-Mixed dataset includes 897 images with annotations of both known and unknown categories.
1 paper · 1 benchmark
COCO-N Medium introduces a stochastic benchmark that simulates common real-world scenarios with noticeable label inaccuracies in the COCO dataset.
1 paper · 1 benchmark
COCO-OOC goes beyond standard object detection to ask the question: Which objects are out-of-context (OOC)?
1 paper · 1 benchmark
COCO-OOD dataset contains only unknown categories, consisting of 504 images with fine-grained annotations of 1655 unknown objects.
1 paper · 1 benchmark
The COCO-WAN benchmark is designed to assess the impact of weakly annotations (combined with auto-annotation tools) noise on instance segmentation models.
1 paper · 1 benchmark
CODD (Cooperative Driving Dataset)
The Cooperative Driving dataset is a synthetic dataset generated using CARLA that contains lidar data from multiple vehicles navigating simultaneously through a diverse set of driving scenarios.
1 paper · 0 benchmarks
COFFE (COFFE: A Code Efficiency Benchmark for Code Generation)
COFFE COFFE is a Python benchmark for evaluating the time efficiency of LLM-generated code.
1 paper · 0 benchmarks
The causal reasoning dataset is generated using the Causal Reasoning in Closed Daily Activities (COLD) framework that helps evaluate large language models (LLMs) on their causal reasoning abilities within real-world, everyday activities.
1 paper · 0 benchmarks
COLLIE-v1 is a dataset with 2080 instances comprising 13 constraint structures designed for text generation under constraints.
1 paper · 0 benchmarks
COMFORT (Consistent Multilingual Frame of Reference Test)
COMFORT is an evaluation protocol to systematically assess the spatial reasoning capabilities of VLMs.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
The COPA-HR dataset (Choice of plausible alternatives in Croatian) is a translation of the English COPA dataset by following the XCOPA dataset translation methodology.
1 paper · 0 benchmarks
DynOPETs is a real-world RGB-D dataset designed for object pose estimation and tracking in dynamic scenes with moving cameras.
1 paper · 0 benchmarks
COQE (Containers Of liQuid contEnt)
Contains more than 5,000 images of 10,000 liquid containers in context labelled with volume, amount of content, bounding box annotation, and corresponding similar 3D CAD models.
1 paper · 0 benchmarks
CORBEL (Conveyor belt pressure signal dataset))
Dataset included measuring static tension under 2 kg load in different points of the CB and measurements in dynamic conditions.
1 paper · 1 benchmark
This dataset is crated via finding the citation links between papers in CORD19 Dataset.
1 paper · 0 benchmarks
CORE-MM is an Open-ended VQA benchmark dataset specifically designed for MLLMs, with a focus on complex reasoning tasks.
1 paper · 1 benchmark
CORRONA CERTAIN (Comparative Effectiveness Registry to Study Therapies for Arthritis and Inflammatory Conditions)
CERTAIN, or the Comparative Effectiveness Registry to Study Therapies for Arthritis and Inflammatory Conditions, is designed as a prospective nested substudy under our larger RA registry.
1 paper · 0 benchmarks
COSTRA 1.0 is a dataset of complex sentence transformations.
1 paper · 0 benchmarks
These datasets were used in the paper 'Evaluation of Thematic Coherence in Microblogs' (ACL, 2021).
1 paper · 0 benchmarks
This case surveillance public use dataset has 12 elements for all COVID-19 cases shared with CDC and includes demographics, any exposure history, disease severity indicators and outcomes, presence of any underlying medical conditions and…
1 paper · 0 benchmarks
A survey of Israelis about their attitudes towards COVID-19 contact tracing apps
1 paper · 0 benchmarks
This is a large public COVID-19 (SARS-CoV-2) lung CT scan dataset, containing total of 8,439 CT scans which consists of 7,495 positive cases (COVID-19 infection) and 944 negative ones (normal and non-COVID-19).
1 paper · 0 benchmarks
The dataset contains Tweet IDs along with the location and tweet timestamp.
1 paper · 0 benchmarks
The data contains CSV files with anonymized user names, tweet texts, vaccine stance, cumulative score for the vaccine stance, location, and topic information.
1 paper · 0 benchmarks
COVID-19-TweetIDs (Tracking Social Media Discourse About the COVID-19 Pandemic: Development of a Public Coronavirus Twitter Data Set)
Since the inception of our collection, we have actively maintained and updated our GitHub repository on a weekly basis.
1 paper · 0 benchmarks
The Covid19-CountryImage dataset is a Twitter dataset which contains COVID-19-related tweets.
1 paper · 0 benchmarks
COVIDx CXR-3 is an open access benchmark dataset that we generated, comprising 30,882 CXR images across 17,026 patient cases.
1 paper · 1 benchmark
COVMis-Stance is a stance detection dataset for COVID-19 misinformation.
1 paper · 0 benchmarks
we introduce COph100, a novel and challenging dataset known as the Comprehensive Ophthalmology Retinal Image Registration dataset for infants with a wide range of image quality issues constituting the public "RIDIRP" database.
1 paper · 0 benchmarks
We present a new simulated dataset for pedestrian action anticipation collected using the CARLA simulator.
1 paper · 0 benchmarks
CPCXR (COVID-19 Posteroanterior Chest X-Ray fused)
The COVID-19 Posteroanterior Chest X-Ray fused (CPCXR) dataset is generated by the fusion of three publicly available datasets: COVID-19 cxr image, Radiological Society of North America (RSNA), and U.S.
1 paper · 0 benchmarks
A large-scale database including substantial CU partition data for HEVC intra- and inter-modes.
1 paper · 0 benchmarks
CPM-Real is a dataset consisting of 3895 images representing real - makeup styles.
1 paper · 0 benchmarks
CPM-Synt-1 is a dataset consisting of 5555 images with synthesis - makeup images with pattern segmentation mask
1 paper · 1 benchmark
CPM-Synt-2 is a dataset consisting of 1625 images with synthesis - triplets: makeup, non-makeup, ground-truth.
1 paper · 1 benchmark
CPMC (crawled persian medical corpus)
a 90 million token medical corpus crawled from medical websites
1 paper · 0 benchmarks
CPNet (CorresPondenceNet)
CPNet dataset has a collection of 25 categories, 2,334 models based on ShapeNetCore, which includes 1,000+ correspondence sets with 104,861 points.
1 paper · 0 benchmarks
In this repository you can find all the elaborate results that were used for the simulated evaluation of an innovative, optimized for real-life use, STC-based, multi-robot Coverage Path Planning (mCPP) algorithm.
1 paper · 0 benchmarks
CPSC2019 (The 2nd China Physiological Signal Challenge (CPSC 2019))
Introduction The China Physiological Signal Challenge 2019 (CPSC 2019) aims to encourage the development of algorithms for challenging QRS detection and heart rate (HR) estimation from short-term single-lead ECG recordings usually with low…
1 paper · 0 benchmarks
CPSC2020 (The 3rd China Physiological Signal Challenge 2020)
Introduction Abnormality of cardiac conduction system can induce arrhythmia.
1 paper · 0 benchmarks
CPSC2021 (The 4th China Physiological Signal Challenge 2021)
Introduction The 4th China Physiological Signal Challenge 2021 (CPSC 2021) aims to encourage the development of algorithms for searching the paroxysmal atrial fibrillation (PAF) events from dynamic ECG recordings.
1 paper · 0 benchmarks
Intrusion alert dataset captured through the Collegiate Penetration Testing Competition (CPTC) 2018.
1 paper · 0 benchmarks
This is the price data that supports the cost estimation of data center providing flexibility, for the following paper "AI-focused HPC Data Centers Can Provide More Power Grid Flexibility and at Lower Cost".
1 paper · 0 benchmarks
Histological images of colorectal cancer, derived from the TCGA database
1 paper · 0 benchmarks
CRED (Crowd Reaction Estimation Dataset)
In the realm of social media, understanding and predicting post reach is a significant challenge.
1 paper · 0 benchmarks
CRIM13 (Caltech Resident-Intruder Mouse 13)
The Caltech Resident-Intruder Mouse dataset (CRIM13) consists of 237x2 videos (recorded with synchronized top and side view) of pairs of mice engaging in social behavior, catalogued into thirteen different actions.
1 paper · 0 benchmarks
Provides two large-scale multi-step benchmarks for biometric identification, where the visual appearance of different classes are highly relevant.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.