Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 126 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 6001–6048 of 12,172
BiGe (Bielefeld Gesture Corpus)
The BiGe corpus is comprised of 54.360 shots of interest extracted from TED and TEDx talks.
2 papers · 0 benchmarks
BiRdQA is a bilingual multiple-choice question answering dataset with 6614 English riddles and 8751 Chinese riddles.
2 papers · 0 benchmarks
BiasCorp is a dataset for racism detection containing 139,090 comments and news segment from three specific sources - Fox News, BreitbartNews and YouTube.
2 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
BioVid (BioVid Heat Pain Database)
To advance methods for pain assessment, in particular automatic assessment methods, the BioVid Heat Pain Database was collected in a collaboration of the Neuro-Information Technology group of the University of Magdeburg and the Medical…
2 papers · 0 benchmarks
Biographical (Biographical: A Semi-Supervised Relation Extraction Dataset)
Biographical is a semi-supervised dataset for RE.
2 papers · 0 benchmarks
BirdClef 2018 is a bird soundscape dataset based on the contributions of the Xeno-canto network.
2 papers · 0 benchmarks
BirdClef 2019 is a bird soundscape dataset.
2 papers · 0 benchmarks
Human keypoint dataset of anime/manga-style character illustrations.
2 papers · 0 benchmarks
BlendMimic3D (A Synthetic Dataset for Human Pose Estimation)
BlendMimic3D is a pioneering synthetic dataset developed using Blender, designed to enhance Human Pose Estimation (HPE) research.
2 papers · 0 benchmarks
The Blog Authorship Corpus consists of the collected posts of 19,320 bloggers gathered from blogger.com in August 2004.
2 papers · 0 benchmarks
Dataset created in the paper "Learning to Count Objects in Images" by Victor Lempitsky and Andrew Zisserman exists as a benchmark to have a dataset useful for cell enumeration.
2 papers · 0 benchmarks
Bone Age (The RSNA Pediatric Bone Age Machine Learning Challenge)
At RSNA 2017 there was a contest to correctly identify the age of a child from an X-ray of their hand.
2 papers · 1 benchmark
BoostCLIR is a bilingual (Japanese-English) corpus of patent abstracts, extracted from the MAREC patent data, and the data from the NTCIR PatentMT workshop collections, accompanied with relevance judgements for the task of patent prior-art…
2 papers · 0 benchmarks
Continuous control tasks in the Box2D simulator.
2 papers · 0 benchmarks
This dataset is a composition of scenes taken by SPOT sensor in 2005 over four counties in the State of Minas Gerais, Brazil: Arceburgo, Guaranesia, Guaxupé and Monte Santo.
2 papers · 0 benchmarks
The breast lesion detection in ultrasound videos dataset uses a clip-level and video-level feature aggregated network (CVA-Net) and consists of 188 ultrasound videos, of which 113 are labeled malignant and 75 benign.
2 papers · 0 benchmarks
Several datasets are fostering innovation in higher-level functions for everyone, everywhere.
2 papers · 0 benchmarks
Differential fluorescent staining is an effective tool widely adopted for the visualization, segmentation and quantification of cells and cellular substructures as a part of standard microscopic imaging protocols.
2 papers · 0 benchmarks
The BugHunter dataset is an automatically constructed and freely available bug dataset containing code elements (files, classes, methods) with a wide set of code metrics and bug information.
2 papers · 0 benchmarks
BugRepo maintains a collection of bug reports that are publicly available for research purposes.
2 papers · 0 benchmarks
A dataset containing 2,221 questions from matriculation exams for twelfth grade in various subjects -history, biology, geography and philosophy-, and 412 additional questions from online quizzes in history.
2 papers · 0 benchmarks
One of the first datasets (if not the first) to highlight the importance of bias and diversity in the community, which started a revolution afterwards.
2 papers · 0 benchmarks
C2A: Combination to Application Dataset Overview This repository contains the code and information for the paper "UAV-Enhanced Combination to Application: Comprehensive Analysis and Benchmarking of a Human Detection Dataset for Disaster…
2 papers · 1 benchmark
CA4P-483 is a dataset designed to facilitate the sequence labeling tasks and regulation compliance identification between privacy policies and software.
2 papers · 0 benchmarks
CADSketchNet is an annotated collection of sketches of 3D CAD models.
2 papers · 0 benchmarks
CAP (Consented Activities of People)
The Consented Activities of People (CAP) dataset is a fine grained activity dataset for visual AI research curated using the Visym Collector platform.
2 papers · 0 benchmarks
CAR (Cityscapes Attributes Recognition)
CAR contains visual attributes for objects in the Cityscapes dataset.
2 papers · 0 benchmarks
CARBEN (Composite Adversarial Robustness Benchmark)
Prior literature on adversarial attack methods has mainly focused on attacking with and defending against a single threat model, e.g., perturbations bounded in Lp ball.
2 papers · 0 benchmarks
CARWC (Consolidated and refined world cup dataset)
Consolidates the world cup 2014 (WC14) and time-series world cup (TSWC) datasets and refines their homography annotations.
2 papers · 0 benchmarks
Introduction Iris is considered one of the most accurate and reliable biometric modality.
2 papers · 0 benchmarks
CAVES (A Dataset to facilitate Explainable Classification and Summarization of Concerns towards COVID Vaccines)
CAVES is the first large-scale dataset containing about 10k COVID-19 anti-vaccine tweets labelled into various specific anti-vaccine concerns in a multi-label setting.
2 papers · 0 benchmarks
CommonCrawl News is a dataset containing news articles from news sites all over the world.
2 papers · 0 benchmarks
Traffic signs are one of the most important information that guide cars to travel, and the detection of traffic signs is an important component of autonomous driving and intelligent transportation systems.
2 papers · 1 benchmark
Our CCTV-Pipe dataset consists of 16 defect categories including structural and functional defects in the pipe.
2 papers · 0 benchmarks
CD&S (Corn Disease and Severity)
The Corn Disease and Severity (CD&S) dataset consists of 511, 524, and 562, field acquired raw images, corresponding to three common foliar corn diseases, namely Northern Leaf Blight (NLB), Gray Leaf Spot (GLS), and Northern Leaf Spot.
2 papers · 0 benchmarks
A video database for testing change detection algorithms.
2 papers · 0 benchmarks
CENTER-TBI (Collaborative European NeuroTrauma Effectiveness Research in TBI)
The CENTER-TBI database contains prospectively collected data of more than 4,500 patients with TBI in Europe.
2 papers · 0 benchmarks
Dataset link: https://github.com/leggedrobotics/cerberusdarpasubtdatasets
2 papers · 0 benchmarks
CEREBRUM-7T (Fast and Fully-volumetric Brain Segmentation of 7 Tesla MR Volumes)
Ultra-high field MRI enables sub-millimetre resolution imaging of human brain, allowing to disentangle complex functional circuits across different cortical depths.
2 papers · 0 benchmarks
CHAMMI (CHAMMI: A benchmark for channel-adaptive models in microscopy imaging)
We present a cellular microscopic image dataset for investigating channel-adaptive models.
2 papers · 0 benchmarks
The CHILI-100K dataset is a large-scale graph dataset (with overall >183M nodes, >1.2B edges) of nanomaterials generated from experimentally determined crystal structures.
2 papers · 8 benchmarks
The CHILI-3K dataset is a medium-scale graph dataset (with overall >6M nodes, >49M edges) of mono-metallic oxide nanomaterials generated from 12 selected crystal types.
2 papers · 8 benchmarks
Cantonese In-car Audio-Visual Speech Recognition (CI-AVSR) is a dataset for in-car command recognition in the Cantonese language with both video and audio data.
2 papers · 0 benchmarks
CIC (Catalonia Independence Corpus)
The dataset is annotated with stance towards one topic, namely, the independence of Catalonia.
2 papers · 3 benchmarks
CIC IoT Dataset 2022 This project aims to generate a state-of-the-art dataset for profiling, behavioural analysis, and vulnerability testing of different IoT devices with different protocols such as IEEE 802.11, Zigbee-based and Z-Wave.
2 papers · 0 benchmarks
CID (Campus Image Dataset)
The CID (Campus Image Dataset) is a dataset captured in low-light env with the help of Android programming.
2 papers · 1 benchmark
CII-Bench (Chinese Image Implication understanding Benchmark)
We introduce the Chinese Image Implication Understanding Benchmark CII-Bench, a new benchmark measuring the higher-order perceptual, reasoning and comprehension abilities of MLLMs when presented with complex Chinese implication images.
2 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.