Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 88 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 4177–4224 of 12,172
BHSD (A 3D Multi-class Brain Hemorrhage Segmentation Dataset)
Intracranial hemorrhage (ICH) is a pathological condition characterized by bleeding inside the skull or brain, which can be attributed to various factors.
5 papers · 0 benchmarks
BIRD (Blocksworld Image Reasoning Dataset)
Blocksworld Image Reasoning Dataset (BIRD) contains images of wooden blocks in different configurations, and the sequence of moves to rearrange one configuration to the other.
5 papers · 1 benchmark
BLUEX is a valuable benchmark dataset designed to evaluate language models in Portuguese.
5 papers · 0 benchmarks
BRACE (The Breakdancing Competition Dataset for Dance Motion Synthesis)
BRACE is a dataset for audio-conditioned dance motion synthesis challenging common assumptions for this task: - strong music-dance correlation - controlled motion data - simple poses and movements To address these issues: - We focus on…
5 papers · 2 benchmarks
BTS3.1 (Expanding Accurate Person Recognition to New Altitudes and Ranges: The BRIAR Dataset)
Large, multimodal biometric dataset: It contains still images and videos of over 1,000 people captured at various ranges (up to 1,000 meters) and elevations (up to 400 meters) using a diverse set of cameras (commercial, military-grade,…
5 papers · 2 benchmarks
BUP20 (Sweet Pepper 2020 University of Bonn)
Video sequences from a glasshouse environment in Campus Kleinaltendorf(CKA), University of Bonn, captured by PATHoBot, a glasshouse monitoring robot.
5 papers · 0 benchmarks
This data set includes beat and bar annotations of the ballroom dataset, introduced by Gouyon et al.
5 papers · 3 benchmarks
This dataset consists of images and annotations in Bengali.
5 papers · 1 benchmark
The BanglaWriting dataset contains single-page handwritings of 260 individuals of different personalities and ages.
5 papers · 1 benchmark
Using the proposed beam-splitter acquisition system, we have collected a new real-world video deblurring dataset (BSD).
5 papers · 1 benchmark
BioCoder is a benchmark developed to evaluate existing pre-trained models in generating bioinformatics code.
5 papers · 0 benchmarks
BioNLI (Biomedical Natural Language Inference)
BioNLI is a dataset in biomedical natural language inference.
5 papers · 1 benchmark
BIWI 3D corpus comprises a total of 1109 sentences uttered by 14 native English speakers (6 males and 8 females).
5 papers · 1 benchmark
The English data for voice building was obtained, prepared and provided the the challenge by Lessac Technologies Inc., having originally came from the publishers Voice Factory International Inc.
5 papers · 1 benchmark
BRATS 2014 is a brain tumor segmentation dataset.
5 papers · 1 benchmark
This brain anatomy segmentation dataset has 1300 2D US scans for training and 329 for testing.
5 papers · 1 benchmark
Jiyoung Lee, Seungryong Kim, Sunok Kim, Jungin Park, Kwanghoon Sohn; Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp.
5 papers · 2 benchmarks
CBC (Complete Blood Count)
The complete blood count (CBC) dataset contains 360 blood smear images along with their annotation files splitting into Training, Testing, and Validation sets.
5 papers · 0 benchmarks
The dataset offers tag and mask annotations for image-text pairs from the CC3M validation set.
5 papers · 2 benchmarks
CCPM (Chinese Classical Poetry Matching)
Introduction CCPM is a large Chinese classical poetry matching dataset that can be used for poetry matching, understanding and translation.
5 papers · 0 benchmarks
Chinese dataset on COVID-19 misinformation.
5 papers · 0 benchmarks
CHI3D is a lab-based accurate 3D motion capture dataset with 631 sequences containing 2,525 contact events,728,664 ground truth 3d poses, as well as FlickrCI3D, a dataset of 11,216 images, with 14,081 processed pairs of people, and 81,233…
5 papers · 0 benchmarks
CHiME-Home is a dataset for sound source recognition in a domestic environment.
5 papers · 0 benchmarks
CICIDS2018 includes seven different attack scenarios: Brute-force, Heartbleed, Botnet, DoS, DDoS, Web attacks, and infiltration of the network from inside.
5 papers · 0 benchmarks
CLEVR-X is a dataset that extends the CLEVR dataset with natural language explanations in the context of VQA.
5 papers · 1 benchmark
CLIC (Challenge on Learned Image Compression)
CLIC is a dataset for learned image compression.
5 papers · 0 benchmarks
CLIP (CLIP: A Dataset for Extracting Action Items for Physicians from Hospital Discharge Notes)
We created a dataset of clinical action items annotated over MIMIC-III.
5 papers · 0 benchmarks
A database of over 1.4 billion 3x3 convolution filters extracted from hundreds of diverse CNN models with relevant meta information.
5 papers · 0 benchmarks
COME15K is an RGB-D saliency detection dataset which contains 15,625 image pairs with high quality polygon-/scribble-/object-/instance-/rank-level annotations.
5 papers · 0 benchmarks
CQADupStack is a benchmark dataset for community question-answering research.
5 papers · 1 benchmark
Request access: cadpath.ai@impdiagnostics.com The CRC dataset contains 1133 colorectal biopsy and polypectomy slides and is the result of our ongoing efforts to contribute to CRC diagnosis with a reference dataset.
5 papers · 0 benchmarks
CSFCube is an expert annotated test collection to evaluate models trained to perform faceted Query by Example.
5 papers · 0 benchmarks
CSRC (Children Speech Recognition Challenge)
CSRC is a collection of data for Children Speech Recognition.
5 papers · 0 benchmarks
CURE-OR (Challenging Unreal and Real Environments for Object Recognition)
CURE-OR is a large-scale, controlled, and multi-platform object recognition dataset denoted as Challenging Unreal and Real Environments for Object Recognition.
5 papers · 0 benchmarks
We present a novel approach to reference-based super-resolution (RefSR) with the focus on real-world dual-camera super-resolution (DCSR).
5 papers · 0 benchmarks
Cata7 is the first cataract surgical instrument dataset for semantic segmentation.
5 papers · 0 benchmarks
Chinese Text in the Wild is a dataset of Chinese text with about 1 million Chinese characters from 3850 unique ones annotated by experts in over 30000 street view images.
5 papers · 0 benchmarks
ChronoMagic with 2265 metamorphic time-lapse videos, each accompanied by a detailed caption.
5 papers · 0 benchmarks
ClimateSet (ClimateSet - : A Large-Scale Climate Model Dataset for Machine Learning)
Climate models are critical tools for analyzing climate change and projecting its future impact.
5 papers · 0 benchmarks
ClovaCall is a new large-scale Korean call-based speech corpus under a goal-oriented dialog scenario from more than 11,000 people.
5 papers · 0 benchmarks
CoDEx comprises a set of knowledge graph completion datasets extracted from Wikidata and Wikipedia that improve upon existing knowledge graph completion benchmarks in scope and level of difficulty.
5 papers · 1 benchmark
CoDEx comprises a set of knowledge graph completion datasets extracted from Wikidata and Wikipedia that improve upon existing knowledge graph completion benchmarks in scope and level of difficulty.
5 papers · 1 benchmark
CoP3D is a collection of crowd-sourced videos showing around 4,200 distinct pets.
5 papers · 0 benchmarks
Coached Conversational Preference Elicitation is a dataset consisting of 502 English dialogs with 12,000 annotated utterances between a user and an assistant discussing movie preferences in natural language.
5 papers · 0 benchmarks
CodRED is the first human-annotated cross-document relation extraction (RE) dataset, aiming to test the RE systems’ ability of knowledge acquisition in the wild.
5 papers · 0 benchmarks
ComFact is a benchmark for commonsense fact linking, where models are given contexts and trained to identify situationally-relevant commonsense knowledge from KGs.
5 papers · 0 benchmarks
ComPhy (Compositional Physical Reasoning Dataset)
Compositional Physical Reasoning is a dataset for understanding object-centric and relational physics properties hidden from visual appearances.
5 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.