Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 239 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 11425–11472 of 12,172
Colorectal-Liver-Metastases (Colorectal-Liver-Metastases | Preoperative CT and Survival Data for Patients Undergoing Resection of Colorectal Liver Metastases)
This collection consists of DICOM images and DICOM Segmentation Objects (DSOs) for 197 patients with Colorectal Liver Metastases (CRLM).
0 papers · 0 benchmarks
A list of all proceedings retrieved from the two-stage keyword (key first, key second in the csv file) filtering approach and the list of all evaluated and reviewed papers by four authors to identify the relevant papers.
0 papers · 0 benchmarks
Corpus of domain names scraped from Common Crawl and manually annotated to add word boundaries (e.g.
0 papers · 0 benchmarks
A fundamental component of human vision is our ability to parse complex visual scenes and judge the relations between their constituent objects.
0 papers · 0 benchmarks
This dataset, which can be used for vision-based deep learning methods, can been implemented to detect and analyze damages in concrete structures.
0 papers · 0 benchmarks
This dataset is an extremely challenging set of over 20,000+ original Construction vehicle images captured and crowdsourced from over 600+ urban and rural areas, where each image is manually reviewed and verified by computer vision…
0 papers · 0 benchmarks
ContraCAT (Contrastive Coreference Analytical Templates (for Machine Translation))
Current approaches to context-aware MT rely on a set of surface heuristics to translate pronouns, which break down when translations require real reasoning.
0 papers · 0 benchmarks
This dataset is the images of corn seeds considering the top and bottom view independently (two images for one corn seed: top and bottom).
0 papers · 0 benchmarks
CornHub (Instance-Segmentation Dataset of Corn Cobs)
🌽 CornHub: Instance-Segmentation Dataset of Corn Cobs Version: 1.0.0 Date: 2025-05-18 License: CC BY 4.0 Author: Sebastian Borukało 📦 Description CornHub is an instance-segmentation dataset of corn cobs in real field conditions.
0 papers · 0 benchmarks
This corpus contains a large metadata-rich collection of fictional conversations extracted from raw movie scripts: - 220,579 conversational exchanges between 10,292 pairs of movie characters - involves 9,035 characters from 617 movies - in…
0 papers · 0 benchmarks
A corpus of movie quotes, annotated with memorability information, in which one is able to control for both the speaker and the setting of the quotes.
0 papers · 0 benchmarks
This dataset includes CSV files that contain IDs and sentiment scores of the tweets related to the COVID-19 pandemic.
0 papers · 0 benchmarks
The study of material corrosion is an important research area, with corrosion degradation of metallic structures causing expenses up to 4% of the global domestic product annually along with major safety risks worldwide.
0 papers · 0 benchmarks
The Couples Therapy corpus contains audio, video recordings and manual transcriptions of conversations between 134 real-life couples attending marital therapy.
0 papers · 0 benchmarks
A dataset that consists of the demographics, triage category, symptoms, and comorbidities of COVID-19 patients.
0 papers · 0 benchmarks
The COVID-19 CT dataset is constructed by Shenzhen Research Institute of Big Data (SRIBD), Future Network of Intelligence Institute (FNii) and CUHKSZ-JD Joint AI Lab, Chinese University of Hongkong, Shenzhen, China, which contains 368…
0 papers · 0 benchmarks
This dataset consists of images of Cracked screen like cracked mobile screen.
0 papers · 0 benchmarks
Crossref is an essential organization in the scholarly publishing domain.
0 papers · 0 benchmarks
Crowd 11 (A Dataset for Fine Grained Crowd Behaviour Analysis)
This dataset defines a total of 11 crowd motion patterns and it is composed of over 6000 video sequences with an average length of 100 frames per sequence.
0 papers · 0 benchmarks
This dataset is an extremely challenging set of over 3000+ original Crowd images captured and crowdsourced from over 300+ urban and rural areas, where each image is manually reviewed and verified by computer vision professionals at…
0 papers · 0 benchmarks
This dataset collects transparency disclosures about the sexual exploitation of children by social media and their reports about such activity and material to the national clearinghouse, the National Center for Missing and Exploited…
0 papers · 0 benchmarks
Description: CytoImageNet is an extensive collection of microscopy images, carefully curated to aid in the development of fast and automated methods for analyzing biological data.
0 papers · 0 benchmarks
The data originate from the journalistic domain in the Czech language.
0 papers · 0 benchmarks
DAHLIA (DAily Human Life Activity)
DAHLIA dataset [1] is devoted to human activity recognition, which is a major issue for adapting smart-home services such as user assistance.
0 papers · 0 benchmarks
DPB-5L is a Multilingual KG dataset containing 5 KGs in English, French, Japanese, Greek, and Spanish.
0 papers · 0 benchmarks
DBR dataset is an environmental audio dataset created for the Bachelor's Seminar in Signal Processing in Tampere University of Technology.
0 papers · 0 benchmarks
DCASE 2021 TASK1A dataset consists of audio examples from 10 different audio scenes.
0 papers · 0 benchmarks
| Name | Purpose | |------|---------| | FM100P | Evaluation of the single palette sorting | | KHTP | Evaluation of the palette pair sorting | | LHSP | Evaluation of the palette similarity measurement | | Perceptual Study | Perceptual Study…
0 papers · 0 benchmarks
This is the DCF template provided by ValueInvesting.io, a high performing value investing platform.
0 papers · 0 benchmarks
DDoS-detect (Towards Resource-Efficient DDoS Detection in IoT: Leveraging Feature Engineering of System and Network Usage Metrics)
The Internet of Things (IoT) is omnipresent, exposing a large number of devices that often lack security controls to the public Internet.
0 papers · 0 benchmarks
The Internet of Things (IoT) is omnipresent, exposing a large number of devices that often lack security controls to the public Internet.
0 papers · 0 benchmarks
Dataset Summary The Deep Evaluation of Audio Representations (DEAR) dataset is a benchmark designed to assess general-purpose audio foundation models on properties critical for hearable devices.
0 papers · 0 benchmarks
The database consists of 89 colour fundus images of which 84 contain at least mild non-proliferative signs (Microaneurysms) of the diabetic retinopathy, and 5 are considered as normal which do not contain any signs of the diabetic…
0 papers · 0 benchmarks
This is an image splicing dataset including different types of preprocessing and postprocessing techniques.
0 papers · 0 benchmarks
DMS (Dense Material Segmentation Dataset)
The Dense Material Segmentation Dataset (DMS) consists of 3 million polygon labels of material categories (metal, wood, glass, etc) for 44 thousand RGB images.
0 papers · 0 benchmarks
In bioinformatics, the issue of mutation discovery and type determination remains a significant concern.
0 papers · 0 benchmarks
Dota 2 is a popular computer game with two teams of 5 players.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.