Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 157 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 7489–7536 of 12,172
This dataset is a BIDS compatible version of the Siena Scalp EEG Database.
1 paper · 0 benchmarks
BIOSED-ACPD: BIOacoustic Sound Event Detection - Adaptive Change Point Detection dataset Description.
1 paper · 0 benchmarks
BIRD (Big Impulse Response Dataset) is an open dataset that consists of 100,000 multichannel room impulse responses (RIRs) generated from simulations using the Image Method, making it the largest multichannel open dataset currently…
1 paper · 0 benchmarks
The BIRDeep Audio Annotations dataset is a collection of bird vocalizations from Doñana National Park, Spain.
1 paper · 0 benchmarks
BKEE (BKEE: Pioneering Event Extraction in the Vietnamese Language)
A novel event extraction dataset for Vietnamese.
1 paper · 0 benchmarks
BLANCA (Benchmarks for LANguage models on Coding Artifacts) is a collection of benchmarks that assess code understanding based on tasks such as predicting the best answer to a question in a forum post, finding related forum posts, or…
1 paper · 0 benchmarks
The BLEBeacon dataset is a collection of Bluetooth Low Energy (BLE) advertisement packets/traces generated from BLE beacons carried by people following their daily routine inside a university building for a whole month.
1 paper · 0 benchmarks
BLM-17m is a labeled dataset for topic detection that contains 17 million tweets.
1 paper · 0 benchmarks
BLN600 (BLN600: A Parallel Corpus of Machine/Human Transcribed Nineteenth Century Newspaper Texts)
A publicly available corpus of nineteenth-century newspaper text focused on crime in London, derived from the Gale British Library Newspapers corpus parts 1 and 2.
1 paper · 0 benchmarks
BLP (Blackout Poetry Dataset)
A blackout poetry dataset constructed from publicly available short stories and large poems.
1 paper · 1 benchmark
Although research on author profiling has quite progressed in abundant resources languages, it is still infancy for limited resources languages such as Bengali.
1 paper · 1 benchmark
This is a dataset for Bengali Captioning from Images.
1 paper · 0 benchmarks
BODMAS (Blue Hexagon Open Dataset for Malware AnalysiS)
We collaborate with Blue Hexagon to release a dataset containing timestamped malware samples and well-curated family information for research purposes.
1 paper · 0 benchmarks
a large-scale, slow event-related human fMRI study incorporating 5,000 real-world images as stimuli.
1 paper · 0 benchmarks
BOOM (Benchmark of Observability Metrics)
BOOM (Benchmark of Observability Metrics) is a large-scale, real-world time series dataset designed for evaluating models on forecasting tasks in complex observability environments.
1 paper · 0 benchmarks
BOTH57M is a body-hand dataset with body-level text prompts and finger-level text prompts.
1 paper · 0 benchmarks
BPAEC (bovine pulmonary artery endothelial cells)
This dataset contains the confocal fluorescence microscopy images of nucleus, actin and mitochondria, where each clear image corresponds to 6 out-of-focus images with different degree of blurring.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
BPCIS (Bacterial Phase Contrast for Instance Segementation)
BPCIS is collection of 364 bacterial phase contrast images and corresponding label matrices for instance segmentation.
1 paper · 0 benchmarks
Brown Pedestrian Odometry Dataset (BPOD) is a dataset for benchmarking visual odometry algorithms in head-mounted pedestrian settings.
1 paper · 0 benchmarks
BP^C (A Benchmark Dataset for Causal Business Process Reasoning)
Large Language Models (LLMs) are increasingly used for boosting organizational efficiency and automating tasks.
1 paper · 0 benchmarks
BPersona-chat is an evaluation dataset based on the English multiturn chat corpus Persona-chat and the Japanese multiturn chat corpus JPersona-chat.
1 paper · 0 benchmarks
BRISC (BRISC: Annotated Dataset for Brain Tumor Segmentation and Classification)
BRISC is a high-quality, expert-annotated MRI dataset curated for brain tumor segmentation and classification.
1 paper · 1 benchmark
BS-Objaverse 660k Dataset is a set of GPT4-Vision-powered multi-modal captions data.
1 paper · 0 benchmarks
BTAT (Blockchain Transaction-based Attacks dataset)
The Synthesis Blockchain Intrusion Detection System dataset We encourage you to also perform reproducible research!.
1 paper · 0 benchmarks
BTS (Building Timeseries Dataset: Empowering Large-Scale Building Analytics)
The Building TimeSeries (BTS) dataset covers three buildings over a three-year period, comprising more than ten thousand timeseries data points with hundreds of unique ontologies.
1 paper · 0 benchmarks
BU-BIL (Boston University Biomedical Image Library)
BU-BIL is an image library which includes six datasets that represent three imaging modalities and six object types.
1 paper · 0 benchmarks
BUPTCampus is a video-based visible-infrared dataset with approximately pixel-level aligned tracklet pairs and single-camera auxiliary samples.
1 paper · 0 benchmarks
The BWB corpus consists of Chinese novels translated by experts into English, and the annotated test set is designed to probe the ability of machine translation systems to model various discourse phenomena.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Baidu PersonaChat, which is a personalization dataset collected and open-sourced by Baidu, is similar to ConvAI2, although it’s Chinese.
1 paper · 0 benchmarks
The dataset contains a total of 253,070 records, with 18 features.
1 paper · 0 benchmarks
This is a link to the source code of the Baking-Large domain introduced in the paper.
1 paper · 0 benchmarks
The Balanced Choice of Plausible Alternatives dataset is a benchmark for training machine learning models that are robust to superficial cues/spurious correlations.
1 paper · 1 benchmark
A Bambara dialectal dataset dedicated for Sentiment Analysis, available freely for Natural Language Processing research purposes
1 paper · 0 benchmarks
The dataset consists of images of bananas and apples.
1 paper · 0 benchmarks
We provide a Mikolov-style word-analogy evaluation set specifically for Bangla, with a sample size of 16678, as well as a translated and curated version of the Mikolov dataset, which contains 10594 samples for cross-lingual research.
1 paper · 0 benchmarks
BanglaBook (Large-scale Bangla Dataset for Sentiment Analysis from Book Reviews)
This repository contains the code, data, and models of the paper titled "BᴀɴɢʟᴀBᴏᴏᴋ: A Large-scale Bangla Dataset for Sentiment Analysis from Book Reviews" published in the Findings of the Association for Computational Linguistics: ACL…
1 paper · 1 benchmark
BanglaEmotion (BanglaEmotion: A Benchmark Dataset for Bangla Textual Emotion Analysis)
BanglaEmotion is a manually annotated Bangla Emotion corpus, which incorporates the diversity of fine-grained emotion expressions in social-media text.
1 paper · 0 benchmarks
A Bilingual Dataset for Bangla and English Voice Commands Colloquial Bangla has adopted many English words due to colonial influence.
1 paper · 1 benchmark
Millions of people around the world have low or no vision.
1 paper · 0 benchmarks
The dataset identifies the shortcomings of existing benchmarks in evaluating the problem of compositional generalization, which underscores the need for the development of datasets tailored to assess compositional generalization in open…
1 paper · 1 benchmark
This is the dataset release for the ACM MobiCom 2023 paper "BatMobility: Towards Flying Without Seeing for Autonomous Drones".
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.