Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 159 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 7585–7632 of 12,172
The dataset contains 50 hours for high quality speech samples from a native speaker and 2 more hours of lower quality recordings from a different speaker
1 paper · 0 benchmarks
Overview This is a dataset of blood cells photos.
1 paper · 0 benchmarks
T(BCD) dataset consisted of a total of 364 blood smear images with annotations.
1 paper · 0 benchmarks
Bluesky Social Dataset Pollution of online social spaces caused by rampaging d/misinformation is a growing societal concern.
1 paper · 0 benchmarks
A real-world low-light camera motion blur dataset for evaluating deblurring radiance fields methods.
1 paper · 0 benchmarks
The first large-scale dataset for training and evaluating novel-view synthesis from blurred images.
1 paper · 0 benchmarks
The linked repository holds data from a controlled single-obstacle avoidance experiment recorded in a motion laboratory.
1 paper · 0 benchmarks
This dataset comprises extensive multi-modal data related to the experimental study of ultrasonically excited pulsating fluid jets used for bone cement removal.
1 paper · 0 benchmarks
The dataset contains anonymised hotel checkins.
1 paper · 0 benchmarks
The Books3 dataset emerged as part of a broader effort to train AI models for natural language understanding and generation.
1 paper · 1 benchmark
Boombox is a multi-modal dataset for visual reconstruction from acoustic vibrations.
1 paper · 0 benchmarks
Boreal Forest Fire (Boreal Forest Fire: UAV-collected Wildfire Detection and Smoke Segmentation Dataset)
This dataset consists of annotated images and videos of smoke resulting from prescribed burning events in Finnish boreal forests.
1 paper · 0 benchmarks
BorealTC (Boreal Terrain Classification Dataset)
Recorded with a Husky A200 wheeled UGV, BorealTC contains 116 min of Inertial Measurement Unit (IMU), motor current, and wheel odometry data, focusing on typical boreal forest terrains, notably snow, ice, and silty loam.
1 paper · 1 benchmark
This dataset is parallel text for Bornholmsk and Danish.
1 paper · 0 benchmarks
The dataset provided is a collection of real-world industrial vibration data collected from a brownfield CNC milling machine.
1 paper · 0 benchmarks
With the remarkable capability to reach the public instantly, social media has become integral in sharing scholarly articles to measure public response.
1 paper · 0 benchmarks
The BottleCap dataset contains over 1100 color images and 7 types of real defects.
1 paper · 1 benchmark
RGB-D instance segmentation box dataset.
1 paper · 2 benchmarks
Box-Jenkins gas furnace, a well-known time series forecasting problem
1 paper · 0 benchmarks
This is a dataset for multi-document summarization in Portuguese, what means that it has examples of multiple documents (input) related to human-written summaries (output).
1 paper · 0 benchmarks
BraTS PEDs 2023 (The Brain Tumor Segmentation (BraTS) Challenge 2023: Focus on Pediatrics (CBTN-CONNECT-DIPGR-ASNR-MICCAI BraTS-PEDs))
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
BraTS-Africa (Brain Tumor Segmentation (BraTS) Challenge: Sub Saharan Africa)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 1 benchmark
Multimodal Brain Tumor Segmentation Challenge 2019
1 paper · 0 benchmarks
BraTs Peds 2024 (The Brain Tumor Segmentation in Pediatrics (BraTS-PEDs) Challenge (CBTN-CONNECT-DIPGR-ASNR-MICCAI BraTS-PEDs) 2024)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 1 benchmark
Brain2Text (Data for: A high-performance speech neuroprosthesis)
From the dataset paper: Brain-computer interfaces (BCIs) can restore communication to people who have lost the ability to move or speak.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
This dataset contains both the artificial and real flower images of bramble flowers.
1 paper · 0 benchmarks
- For each BDLO, dynamic trajectory data is captured in real-world settings using a motion capture system operating at 100 Hz when robots grasp the BDLO’s ends.
1 paper · 0 benchmarks
BrazilDAM is a multi sensor and multitemporal dataset that consists of multispectral images of ore tailings dams throughout Brazil.
1 paper · 0 benchmarks
See https://www.kaggle.com/datasets/olistbr/brazilian-ecommerce .
1 paper · 0 benchmarks
Brazilian Protest is a dataset for event filtering that focuses on protests in multi-modal social media data, with most of the text in Portuguese.
1 paper · 0 benchmarks
BreastDICOM4 ([MIMBCD-UI] UTA4: Medical Imaging DICOM Files Dataset)
Several datasets are fostering innovation in higher-level functions for everyone, everywhere.
1 paper · 1 benchmark
Several datasets are fostering innovation in higher-level functions for everyone, everywhere.
1 paper · 0 benchmarks
The dataset contains around 180K rendered images with 100K classified as anomaly and 80K normal.
1 paper · 0 benchmarks
BuGL is a large-scale cross-language dataset for bug localization in code.
1 paper · 0 benchmarks
A multi-physics dataset of boiling processes.
1 paper · 0 benchmarks
BuckTales (A multi-UAV dataset for multi-object tracking and re-identification of wild antelopes)
The first and large scale dataset to solve multi-object tracking and Re-identification problem with wild animals using UAVs.
1 paper · 0 benchmarks
Dataset of 5,591 labeled issue tickets.
1 paper · 0 benchmarks
The original paper contains a high-level explanation of the dataset characteristics, and potential use cases of the dataset.
1 paper · 0 benchmarks
Bukva (Bukva: Russian Sign Language Alphabet)
We introduce a video dataset Bukva for Russian Dactyl Recognition task.
1 paper · 1 benchmark
BurnMD (A Fire Projection and Mitigation Modeling Dataset)
A dataset composed of 308 medium sized fires from the years 2018-2021, complete with both time series airborne based inference and ground operational estimation of fire extent, and operational mitigation data such as control line…
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.