Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 156 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 7441–7488 of 12,172
Data collected from two budget surveys (FY2021 in 2020 and FY2022 in 2021) in collaboration with the City of Austin budget department.
1 paper · 0 benchmarks
Auto-KWS is a dataset for customized keyword spotting, the task of detecting spoken keywords.
1 paper · 0 benchmarks
AutoFR Dataset is broken down by each site that we crawl within a zip file.
1 paper · 0 benchmarks
Dataset proposed by ACM MM 2023 paper "AutoPoster: A Highly Automatic and Content-aware Design System for Advertising Poster Generation" We gather 76537 advertising posters from an e-commerce advertising platform.
1 paper · 0 benchmarks
Temporal Dataset for Indoor and In-Vehicle Thermal Comfort Estimation Abstract Thermal comfort estimation is essential for enhancing user experience in static indoor environments and dynamic in-vehicle scenarios.
1 paper · 0 benchmarks
Collects all the courses from XuetangX5, one of the largest MOOCs in China, and this results in 1951 courses.
1 paper · 0 benchmarks
Forecasting future world events is a challenging but valuable task.
1 paper · 0 benchmarks
This is a benchmark for neural paraphrase detection, to differentiate between original and machine-generated content.
1 paper · 0 benchmarks
Raw antibody microarray data.
1 paper · 0 benchmarks
Logging—used for system events and security breaches to more informational yet essential aspects of software features—is pervasive.
1 paper · 0 benchmarks
This dataset is used to evaluate a predictive consent model for users’ information shared in social media.
1 paper · 0 benchmarks
Dataset Overview: 998 images and 4,208 annotations focusing on interaction with in-vehicle infotainment (IVI) systems.
1 paper · 0 benchmarks
The Autonomous-driving StreAming Perception (ASAP) benchmark is a benchmark to evaluate the online performance of vision-centric perception in autonomous driving.
1 paper · 0 benchmarks
A multi-tasking oral ulcer dataset (Autooral dataset) is proposed.
1 paper · 1 benchmark
For more details see https://huggingface.co/datasets/jpwahle/autoregressive-paraphrase-dataset
1 paper · 0 benchmarks
AuxAD is a a distantly supervised dataset for acronym disambiguation.
1 paper · 0 benchmarks
AuxAI is a distantly supervised dataset for acronym identification.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
AviationQA is introduced in the paper titled- There is No Big Brother or Small Brother: Knowledge Infusion in Language Models for Link Prediction and Question Answering The paper is accepted in the main conference of ICON 2022.
1 paper · 1 benchmark
A syllogism is a common form of deductive reasoning that requires precisely two premises and one conclusion.
1 paper · 0 benchmarks
AzSLD (AzSLD - Azerbaijani Sign Language Dataset)
The Azerbaijani Sign Language Dataset (AzSLD) is a comprehensive, large dataset designed to facilitate the development and evaluation of machine learning models for the recognition and translation of Azerbaijani Sign Language (AzSL).
1 paper · 0 benchmarks
B-XAIC consists of 50K small molecules represented as graphs and includes 7 graph classification tasks, each with ground truth labels and corresponding explanations.
1 paper · 0 benchmarks
BAH (Behavioural Ambivalence/Hesitancy)
Recognizing complex emotions linked to ambivalence and hesitancy (A/H) can play a critical role in the personalization and effectiveness of digital behaviour change interventions.
1 paper · 0 benchmarks
To collect the metrics data, we deploy three benchmark microservice systems: Online Boutique, Sock Shop, and Train Ticket, on a Kubernetes cluster consisting of one master node and five worker nodes.
1 paper · 0 benchmarks
BASIR (BASIR_Budget_Assisted_Sectoral_Impact_Ranking)
Government fiscal policies, particularly annual union budgets, exert significant influence on financial markets.
1 paper · 0 benchmarks
This dataset is for evaluating the task of Black-box Multi-agent Integration which focuses on combining the capabilities of multiple black-box conversational agents at scale.
1 paper · 1 benchmark
Since robust foreground/background separation and segmentation of cellular objects (i.e.,identification of which pixels below to which objects) strongly depends on image quality, focus artifacts are detrimental to data quality.
1 paper · 0 benchmarks
BBBC041 (P. vivax (malaria) infected human blood smears)
P.
1 paper · 0 benchmarks
An annotated subset of the W3C email corpus.
1 paper · 0 benchmarks
BCSD (Bank Check Segmentation Dataset)
The dataset consists of images of 158 filled out bank checks containing various complex backgrounds, and handwritten text and signatures in the respective fields, along with both pixel-level and patch-level segmentation masks for the…
1 paper · 0 benchmarks
BCSS (Breast Cancer Semantic Segmentation)
The BCSS dataset contains over 20,000 segmentation annotations of tissue regions from breast cancer images from The Cancer Genome Atlas (TCGA).
1 paper · 0 benchmarks
BCWS (Bilingual Contextual Word Similarity)
Dataset for evaluating English-Chinese Bilingual Contextual Word Similarity.
1 paper · 0 benchmarks
BD-TypoSAT (Building Damage Typology Satellite Dataset)
On Sunday, August 29, 2021, Hurricane Ida struck parts of Louisiana and Mississippi with wind gusts reaching up to 172 mph, leaving more than a million customers without electricity, including the entire New Orleans area.
1 paper · 0 benchmarks
BDD-QA is distinguished by its encompassing range of traffic actions, crafted to rigorously evaluate a model's decision-making abilities in traffic scenario.
1 paper · 0 benchmarks
BDD100K-weather is a dataset which is inherited from BDD100K using image attribute labels for Out-of-Distribution object detection.
1 paper · 0 benchmarks
This dataset provides a comprehensive resource for detecting and evaluating bias across multiple NLP tasks.
1 paper · 0 benchmarks
BEAMetrics (Benchmark to Evaluate Automatic Metrics) is resource to make research into new metrics for evaluation of generated language easier to evaluate.
1 paper · 0 benchmarks
BEAR (Benchmark on video Action Recognition)
BEAR (Benchmark on video Action Recognition) is a collection of 18 video datasets grouped into 5 categories (anomaly, gesture, daily, sports, and instructional), which covers a diverse set of real-world applications.
1 paper · 0 benchmarks
BEAR-probe (Benchmark for Evaluating Associative Reasoning)
The BEAR dataset and its larger version, BEARbig, are benchmarks for evaluating common factual knowledge contained in language models.
1 paper · 0 benchmarks
BERSt (Basic Emotion Random phrase Shouts)
BERSt Dataset We release the BERSt Dataset for various speech recognition tasks including Automatic Speech Recognition (ASR) and Speech Emotion Recogniton (SER) Overview 4526 single phrase recordings (~3.75h) 98 professional actors 19…
1 paper · 1 benchmark
BFN (Backdoored Face-Networks Dataset)
This database is a database of backdoored neural networks intended for face recognition.
1 paper · 0 benchmarks
BFRD (Bengali Fake Review Dataset)
This is a binary dataset used for Bengali fake review detection in the paper "Bengali Fake Reviews: A Benchmark Dataset and Detection System" accepted in Neurocomputing, a journal published by Elsevier.
1 paper · 0 benchmarks
We recorded gun sounds by changing the type and position of guns to diversify distances and angles in the PUBG environment.
1 paper · 0 benchmarks
BH-rPPG dataset (stands for Beihang University Remote PhotoPlethysmoGraphy) is a dataset consists of 3 lighting conditions with uneven distribution which collected in indoor environment.
1 paper · 0 benchmarks
BIDCD (Bosch Industrial Depth Completion Dataset)
Bosch Industrial Depth Completion Dataset (BIDCD) is an RGBD dataset for of static table-top scenes with industrial objects.
1 paper · 0 benchmarks
This dataset is a BIDS-compatible version of the CHB-MIT Scalp EEG Database.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.