Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 109 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 5185–5232 of 12,172
BUAA-MIHR dataset is a remote photoplethysmography (rPPG) dataset.
3 papers · 0 benchmarks
BUSTER (BUSiness Transaction Entity Recognition dataset.)
BUSiness Transaction Entity Recognition dataset.
3 papers · 0 benchmarks
This dataset contains Bangla handwritten numerals, basic characters and compound characters.
3 papers · 2 benchmarks
23,000 cropped images of tree bark, for 23 species of trees around Quebec City, Canada.
3 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
3 papers · 0 benchmarks
BenchIE: a benchmark and evaluation framework for comprehensive evaluation of OIE systems for English, Chinese and German.
3 papers · 1 benchmark
Benchmark for AMR Metrics based on Overt Objectives (Bamboo), the first benchmark to support empirical assessment of graph-based MR similarity metrics.
3 papers · 1 benchmark
This dataset contains images of individual hand-written Bengali characters.
3 papers · 1 benchmark
Contains 3,689,229 English news articles on politics, gathered from 11 United States (US) media outlets covering a broad ideological spectrum.
3 papers · 0 benchmarks
Due to the highly variable sample size of the original BirdClef2020 dataset and the issues that it presents with reproducibility, we propose a pruned version of the set, where samples longer than 180s are removed along with classes with…
3 papers · 1 benchmark
The BirdVox-full-night dataset contains 6 audio recordings, each about ten hours in duration.
3 papers · 0 benchmarks
BnB is a large-scale and diverse in-domain VLN (Vision and Language Navigation) dataset.
3 papers · 0 benchmarks
Bongard-OpenWorld is a new benchmark for evaluating real-world few-shot reasoning for machine vision.
3 papers · 1 benchmark
In order to collect thermal aerial data, we used FLIR's Boson thermal imager (8.7 mm focal length, 640p resolution, and 50° horizontal field of view)\footnote{\url{https://www.flir.es/products/boson/}}.
3 papers · 0 benchmarks
✔️Abstract A Brain tumor is considered as one of the aggressive diseases, among children and adults.
3 papers · 0 benchmarks
The BuzzFeed-Webis Fake News Corpus 16 comprises the output of 9 publishers in a week close to the US elections.
3 papers · 0 benchmarks
CANNOT (Compilation of ANnotated, Negation-Oriented Text-pairs)
Dataset Summary CANNOT is a dataset that focuses on negated textual pairs.
3 papers · 0 benchmarks
Visible and thermal images have been acquired using a thermographic camera TESTO 880-3, equipped with an uncooled detector with a spectral sensitivity range from 8 to 14 μm and provided with a germanium optical lens, and an approximate…
3 papers · 0 benchmarks
CASIA-Face-Africa is a face image database which contains 38,546 images of 1,183 African subjects.
3 papers · 0 benchmarks
CATT (CATT Arabic Diacritization Benchmark Dataset)
The CATT benchmark dataset comprises 742 sentences, which were scraped from an internet news source in 2023.
3 papers · 1 benchmark
CC-19 is a small new dataset related to the latest family of coronavirus i.e.
3 papers · 0 benchmarks
The CER initiated the Smart Metering Project in 2007 with the purpose of undertaking trials to assess the performance of Smart Meters, their impact on consumers’ energy consumption and the economic case for a wider national rollout.
3 papers · 0 benchmarks
CFC (Caltech Fish Counting Dataset)
Caltech Fish Counting Dataset (CFC) is a large-scale dataset for detecting, tracking, and counting fish in sonar videos.
3 papers · 0 benchmarks
This paper presents CG-Eval, the first comprehensive evaluation of the generation capabilities of large Chinese language models across a wide range of academic disciplines.
3 papers · 0 benchmarks
CHAIRS is a large-scale motion-captured f-AHOI dataset, consisting of 17.3 hours of versatile interactions between 46 participants and 81 articulated and rigid sittable objects.
3 papers · 0 benchmarks
CHQ-Summ (Consumer Healthcare Question Summarization)
Contains 1507 domain-expert annotated consumer health questions and corresponding summaries.
3 papers · 0 benchmarks
CICEROv2 (Contextualized Commonsense Inference in Dialogues (V2))
The CICEROv2 dataset can be found in the data directory.
3 papers · 1 benchmark
Consists of two pedestrian trajectory datasets, CITR dataset and DUT dataset, so that the pedestrian motion models can be further calibrated and verified, especially when vehicle influence on pedestrians plays an important role.
3 papers · 0 benchmarks
The CLCD dataset consists of 600 pairs image of cropland change samples, with 360 pairs for training, 120 pairs for validation and 120 pairs for testing.
3 papers · 1 benchmark
A dataset with two separate domains, i.e., the "Banking'' domain and the "Credit cards'' domain with both general Out-of-Scope (OOD-OOS) queries and In-Domain but Out-of-Scope (ID-OOS) queries, where ID-OOS queries are semantically similar…
3 papers · 0 benchmarks
CMMD (The Chinese Mammography Database)
Breast carcinoma is the second largest cancer in the world among women.
3 papers · 1 benchmark
This dataset contains plot summaries for 16,559 books extracted from Wikipedia, along with aligned metadata from Freebase, including book author, title, and genre.
3 papers · 0 benchmarks
CNewSum is a large-scale Chinese news summarization dataset which consists of 304,307 documents and human-written summaries for the news feed.
3 papers · 0 benchmarks
A novel dataset that represents complex conversational interactions between two individuals via 3D pose.
3 papers · 0 benchmarks
COOLL (Controlled On/Off Loads Library)
Controlled On/Off Loads Library (COOLL) is a dataset of high-sampled electrical current and voltage measurements representing individual appliances consumption.
3 papers · 0 benchmarks
Semi-Structured Explanations for COPA (COPA-SSE) is a new crowdsourced dataset of 9,747 semi-structured, English common sense explanations for COPA questions.
3 papers · 0 benchmarks
COPEN (COnceptual knowledge Probing bENchmark)
COPEN is a COnceptual knowledge Probing benchmark that aims to analyze the conceptual understanding capabilities of Pre-trained Language Models (PLMs).
3 papers · 0 benchmarks
We introduce FUNSD-r and CORD-r in Token Path Prediction, the revised VrD-NER datasets to reflect the real-world scenarios of NER on scanned VrDs.
3 papers · 1 benchmark
CORSMAL is a dataset for estimating the position and orientation in 3D (or 6D pose) of an object from a single view.
3 papers · 0 benchmarks
A benchmark dataset with 960 pairs of Chinese wOrd Similarity, where all the words have two morphemes in three Part of Speech (POS) tags with their human annotated similarity rather than relatedness.
3 papers · 0 benchmarks
COSIAN (a collection of singing voice annotation)
COSIAN is an annotation collection of Japanese popular (J-POP) songs, focusing on singing style and expression of famous solo-singers.
3 papers · 0 benchmarks
The COUNTER (COrpus of Urdu News TExt Reuse) corpus contains 600 source-derived document pairs collected from the field of journalism.
3 papers · 0 benchmarks
COVID-CQ is a stance data set of user-generated content on Twitter in the context of COVID-19.
3 papers · 0 benchmarks
COVID-Q consists of COVID-19 questions which have been annotated into a broad category (e.g.
3 papers · 0 benchmarks
CPP (Chinese Polyphones with Pinyin)
A benchmark dataset that consists of 99,000+ sentences for Chinese polyphone disambiguation.
3 papers · 1 benchmark
The high-quality multi-turn dialogue dataset, which has a total of 3,134 multi-turn consultation dialogues.
3 papers · 0 benchmarks
CREMP is a resource generated for the rapid development and evaluation of machine learning models for macrocyclic peptides.
3 papers · 0 benchmarks
This dataset is constructed and based on the online free-access fictions that are tagged with sci-fi, urban novel, love story, youth, etc.
3 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.