Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 238 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 11377–11424 of 12,172
The CHAINS dataset is utilized to evaluate the proposed system.
0 papers · 0 benchmarks
The CICIoMT2024 dataset is a comprehensive dataset designed for cybersecurity research focused on the Internet of Medical Things (IoMT).
0 papers · 0 benchmarks
CIDII Dataset (Correct Information and Disinformation about Islamic Issues)
The CIDII dataset is a binary classification, consisting of two classes of correct information and disinformation related to Islamic issues.
0 papers · 0 benchmarks
This eye tracking video database can be used to validate visual attention models.
0 papers · 0 benchmarks
Technologically Assisted Reviews in Empirical Medicine.
0 papers · 0 benchmarks
CLO-43SD is a dataset for multi-class species identification in avian flight calls.
0 papers · 0 benchmarks
CLO-SWTH is a dataset for species-specific flight call identification for the Swainson’s Thrush.
0 papers · 0 benchmarks
CLO-WTSP is a dataset for species-specific flight call identification for the White-Throated Sparrow.
0 papers · 0 benchmarks
The CMU Wilderness Multilingual Speech Dataset is a dataset of over 700 different languages providing audio, aligned text and word pronunciations.
0 papers · 0 benchmarks
CNTD (Chinese and Naxi text detection)
Chinese and Naxi scene text detection data set, labelme to json.
0 papers · 0 benchmarks
COCA (The Corpus of Contemporary American English)
The Corpus of Contemporary American English (COCA) is a large and balanced corpus of American English.
0 papers · 1 benchmark
COCO-Facet is a benchmark for attribute-focused text-to-image retrieval, comprising 9,112 queries with 100 candidate images for each.
0 papers · 0 benchmarks
The COCONut dataset is a modernized segmentation dataset that builds upon the established COCO benchmark.
0 papers · 0 benchmarks
Applications of unmanned aerial vehicle (UAV) in logistics, agricultural automation, urban management, and emergency response are highly dependent on oriented object detection (OOD) to enhance visual perception.
0 papers · 0 benchmarks
A survey of Israelis about their attitudes towards COVID-19 contact tracing apps
0 papers · 0 benchmarks
With the emergence of the COVID-19 pandemic, the political and the medical aspects of disinformation merged as the problem got elevated to a whole new level to become the first global infodemic.
0 papers · 0 benchmarks
CRIC Cervix (Center for Recognition and Inspection of Cells (CRIC) Cervix collection)
he Center for Recognition and Inspection of Cells (CRIC) platform enables the creation of CRIC Cervix collection, currently with 400 images (1,376 x 1,020 pixels) curated from conventional Pap smears, with manual classification of 11,534…
0 papers · 0 benchmarks
The CTV-Dataset (CTV stands for Cyclist Top-View) is a trajectories dataset for cyclist behaviour in mixed-traffic environments (aka.
0 papers · 0 benchmarks
CUHK occlusion dataset includes 1,063 images with occluded pedestrians.
0 papers · 0 benchmarks
CUHK Square data set is for transfer learning research on adapting generic pedestrian detectors.
0 papers · 0 benchmarks
CVGL Camera Calibration Dataset consists of 49 camera configurations with town 1 having 25 configurations while town 2 having 24 configurations.
0 papers · 0 benchmarks
We present sentence aligned parallel corpora across 10 Indian Languages - Hindi, Telugu, Tamil, Malayalam, Gujarati, Urdu, Bengali, Oriya, Marathi, Punjabi, and English - many of which are categorized as low resource.
0 papers · 0 benchmarks
This publicly available data is synthesised audio for woodwind quartets including renderings of each instrument in isolation.
0 papers · 0 benchmarks
This dataset focuses only on the robbery category, presenting a new weakly labelled dataset that contains 486 new real–world robbery surveillance videos acquired from public sources.
0 papers · 0 benchmarks
Car crash dataset RUSSIA 2022-2023 is a big driving video dataset that contains over 500 high-resolution videos of various driving scenarios.
0 papers · 0 benchmarks
This dataset is a collection of 4,000 images of cars in multiple scenes that are ready to use for optimizing the accuracy of computer vision models.
0 papers · 0 benchmarks
Carioca 1 dataset consists of 200 audios having frequency of 44 KHz
0 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
0 papers · 0 benchmarks
ChaBuD (Change detection for Burned area Delineation)
The dataset comprises patches of size 512x512 pixels collected from Sentinel-2 L2A satellite mission.
0 papers · 0 benchmarks
ChaLearn Pose is a subset of the ChaLearn 2013 Multi-modal gesture dataset from Escalera et al.
0 papers · 0 benchmarks
Dataset of chaotic Chua, Lorenz, Lorenz96, Mackey-Glass with tau=17, Mackey-Glass with tau=30, Rossler, Sprott systems.
0 papers · 0 benchmarks
This is a dataset of paraphrases created by ChatGPT.
0 papers · 0 benchmarks
Dataset Description Our dataset contains questions from a well-known software testing book Introduction to Software Testing 2nd Edition by Ammann and Offutt.
0 papers · 0 benchmarks
Chernobyl is a collection of 620 audio clips collected from unattended remote monitoring equipment in the Chernobyl Exclusion Zone (CEZ).
0 papers · 0 benchmarks
The photo fixation of cherry fruitlets was done in the LatHort orchard in Dobele, at the development of fruit (BBCH stage 72).
0 papers · 0 benchmarks
The photo fixation of cherry fruits was done in the LatHort orchard in Dobele, at the beginning of fruit coloration (BBCH stage 81).
0 papers · 0 benchmarks
ChineseSquad (中文机器阅读理解数据集) is a dataset specifically designed for Chinese machine reading comprehension.
0 papers · 0 benchmarks
The Cityscapes-Motion dataset is a supplement to the semantic annotations provided by the Cityscapes dataset, containing 2975 training images and 500 validation images.
0 papers · 0 benchmarks
CloudSEN12 is a LARGE dataset (~1 TB) for cloud semantic understanding that consists of 49,400 image patches (IP) that are evenly spread throughout all continents except Antarctica.
0 papers · 0 benchmarks
Key Points - Purpose: Captures crossroad navigation under cloudy weather conditions.
0 papers · 0 benchmarks
A large corpus of discourse annotations and relations on ~10K forum threads.
0 papers · 0 benchmarks
CodeFuseEval is a Code Generation benchmark that combines the multi-tasking scenarios of CodeFuse Model with the benchmarks of HumanEval-x and MBPP.
0 papers · 0 benchmarks
CodeSCAN (ScreenCast ANalysis for Video Programming Tutorials)
CodeSCAN is the first large-scale and diverse dataset of coding screenshots with pixel-perfect annotations.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.