Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 128 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 6097–6144 of 12,172
ChEMBL is a manually curated database of bioactive molecules with drug-like properties.
2 papers · 0 benchmarks
CheGeKa is a Jeopardy!-like Russian QA dataset collected from the official Russian quiz database ChGK.
2 papers · 1 benchmark
Annotated audio files (separate combined annotation file) of lung sounds as recorded from various vantage points of the chest wall.
2 papers · 1 benchmark
Set of landmark annotations for JSRT, Montgomery, Shenzhen and a subset of Padchest datasets
2 papers · 0 benchmarks
The airborne hyperspectral dataset was taken by Headwall Hyperspec-VNIR-C imaging sensor over agricultural and urban areas in Chikusei, Ibaraki, Japan, on July 29, 2014 between the times 9:56 to 10:53 UTC+9.
2 papers · 1 benchmark
We introduce ChinaTravel, the first open-ended benchmark grounded in authentic Chinese travel requirements collected from 1,154 human participants.
2 papers · 0 benchmarks
Classifiers are function words that are used to express quantities in Chinese and are especially difficult for language learners.
2 papers · 0 benchmarks
Chinese Gigaword corpus consists of 2.2M of headline-document pairs of news stories covering over 284 months from two Chinese newspapers, namely the Xinhua News Agency of China (XIN) and the Central News Agency of Taiwan (CNA).
2 papers · 0 benchmarks
CholecT40 is the first endoscopic dataset introduced to enable research on fine-grained action recognition in laparoscopic surgery.
2 papers · 1 benchmark
CholecTrack20 (Multi-Perspective Multi-Class Multi-Object Tracking Dataset For Surgical Tools)
CholecTrack20 is a surgical video dataset focusing on laparoscopic cholecystectomy and designed for surgical tool tracking, featuring 20 annotated videos.
2 papers · 0 benchmarks
Description - Venue: NeurIPS 2024 D&B Spotlight - Repository: Code, Page, Data - Paper: arxiv.org/abs/2406.18522 - Point of Contact: Shenghai Yuan Citation If you find our paper and code useful in your research, please consider giving a…
2 papers · 0 benchmarks
CinemAirSim is an extension of the well-known drone simulator, AirSim, with a cinematic camera as well as extended its API to control all of its parameters in real time, including various filming lenses and common cinematographic…
2 papers · 0 benchmarks
Cinescale (CineScale: A dataset of cinematic shot scale in movies)
We provide a database containing shot scale annotations (i.e., the apparent distance of the camera from the subject of a filmed scene) for more than 792,000 image frames.
2 papers · 0 benchmarks
The Circa (meaning ‘approximately’) dataset aims to help machine learning systems to solve the problem of interpreting indirect answers to polar questions.
2 papers · 0 benchmarks
CiteWorth is a a large, contextualized, rigorously cleaned labelled dataset for cite-worthiness detection built from a massive corpus of extracted plain-text scientific documents.
2 papers · 0 benchmarks
City Street: We collected a multi-view video dataset of a busy city street using 5 synchronized cameras.
2 papers · 0 benchmarks
City-Networks, a transductive learning dataset for testing long-range dependencies in Graph Neural Networks (GNNs).
2 papers · 4 benchmarks
BEV Crowd-Counting dataset extended from CityUHK-X
2 papers · 0 benchmarks
The training and validation data are subsets of the training split of the Cityscapes dataset.
2 papers · 1 benchmark
The paper introduces three benchmarking tasks inspired by animal learning.
2 papers · 0 benchmarks
The Climate Change Claims dataset for generating fact checking summaries contains claims broadly related to climate change and global warming from climatefeedback.org.
2 papers · 0 benchmarks
The dataset was created to address the crucial need for effective Extreme Weather Events Detection (EWED), an increasingly urgent task due to the rising frequency of such events driven by global warming.
2 papers · 0 benchmarks
This dataset is created from MIMIC-III (Medical Information Mart for Intensive Care III) and contains simulated patient admission notes.
2 papers · 4 benchmarks
A test dataset that annotated articles in 2020 following the CoNLL-2003 NER task.
2 papers · 1 benchmark
CoVaxFrames includes 113 Vaccine Hesitancy Framings found on Twitter about the COVID-19 vaccines.
2 papers · 0 benchmarks
CoVaxLies v2 includes 47 Misinformation Targets (MisTs) found on Twitter about the COVID-19 vaccines.
2 papers · 0 benchmarks
It contains the dataset of class comments extracted from various projects of three programming languages Java, Pharo, and Python
2 papers · 0 benchmarks
CodeQueries Benchmark dataset consists of instances of semantic queries, code context and code spans in the context corresponding to the semantic queries.
2 papers · 0 benchmarks
CodeSyntax is a large-scale dataset of programs annotated with the syntactic relationships in their corresponding abstract syntax trees.
2 papers · 0 benchmarks
Colorectal Adenoma contains 177 whole slide images (156 contain adenoma) gathered and labelled by pathologists from the Department of Pathology, The Chinese PLA General Hospital.
2 papers · 0 benchmarks
The Common Crawl corpus contains petabytes of data collected over 12 years of web crawling.
2 papers · 0 benchmarks
Comparative Question Completion is a dataset to evaluate what do large Language Models learn.
2 papers · 0 benchmarks
Concise has two datasets of 2000 sentences each, that were annotated by two and five human annotators, respectively.
2 papers · 0 benchmarks
State-level data for the US economy through the lens of consumer spending (Credit/Debit Spending) .
2 papers · 1 benchmark
ContactArt is a dataset for learning hand-object interaction priors for hand and articulated object pose estimation.
2 papers · 0 benchmarks
A new dataset of contour drawings.
2 papers · 0 benchmarks
The dataset is constructed from an Amazon review corpus by integrating both user-agent dialogue and custom knowledge graphs for recommendation.
2 papers · 0 benchmarks
Covid-HeRA is a dataset for health risk assessment and severity-informed decision making in the presence of COVID19 misinformation.
2 papers · 0 benchmarks
Includes 3000 animated sequences rendered using styles randomly selected from 40 textured line styles and 38 shading styles, spanning the range between flat cartoon fill and wildly sketchy shading.
2 papers · 0 benchmarks
Underground hacking forums
2 papers · 0 benchmarks
The appearance of the world varies dramatically not only from place to place but also from hour to hour and month to month.
2 papers · 1 benchmark
Given an English article, generate a short summary in the target language.
2 papers · 0 benchmarks
A benchmark dataset for training and evaluating global cloud classification models.
2 papers · 0 benchmarks
The Curiosity dataset consists of 14K dialogs (with 181K utterances) with fine-grained knowledge groundings, dialog act annotations, and other auxiliary annotation.
2 papers · 0 benchmarks
Curlie dataset is a dataset with more than 1M websites in 92 languages with relative labels collected from Curlie, the largest multilingual crowdsourced Web directory.
2 papers · 0 benchmarks
As social media usage becomes increasingly prevalent in every age group, a vast majority of citizens rely on this essential medium for day-to-day communication.
2 papers · 0 benchmarks
Archive of Global Tropical Cyclone Tracks Tracks from 1980 to May 2019.
2 papers · 0 benchmarks
Cylinder in Crossflow is a synthetic dataset that involves unsteady laminar flow past a cylinder that generates vortex shedding pattern known as a von Kármán vortex street.
2 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.