Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 237 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 11329–11376 of 12,172
BI2015a MOABB (P300 dataset BI2015a from a "Brain Invaders" experiment.)
0 papers · 0 benchmarks
BI2015b MOABB (P300 dataset BI2015b from a "Brain Invaders" experiment.)
0 papers · 0 benchmarks
BIGOS (Benchmark Intended Grouping of Open Speech)
The Benchmark Intended Grouping of Open Speech (BIGOS) is a novel corpus specifically designed for Polish Automatic Speech Recognition (ASR) systems.
0 papers · 0 benchmarks
BMI/OpenBMI dataset for MI.
0 papers · 0 benchmarks
BMS-26 (Berkeley Motion Segmentation)
The Berkeley Motion Segmentation Dataset (BMS-26) is a dataset for motion segmentation, which consists of 26 video sequences with pixel-accurate segmentation annotation of moving objects.
0 papers · 0 benchmarks
BNCI 2014-001 Motor Imagery dataset Dataset IIa from BCI Competition 4 [1].
0 papers · 0 benchmarks
Dataset Description This data set consists of EEG data from 9 subjects.
0 papers · 0 benchmarks
Dataset description This data set consists of EEG data from 9 subjects of a study published in [1].
0 papers · 0 benchmarks
Dataset description We acquired the EEG from three Laplacian derivations, 3.5 cm (center-to- center) around the electrode positions (according to International 10-20 System of Electrode Placement) C3 (FC3, C5, CP3 and C1), Cz (FCz, C1, CPz…
0 papers · 0 benchmarks
Dataset description We provide EEG data recorded from nine users with disability (spinal cord injury and stroke) on two different days (sessions).
0 papers · 0 benchmarks
Datasets for Bangla Natural Language Processing tasks.
0 papers · 0 benchmarks
BSTLD (Bosch Small Traffic Lights Dataset)
This dataset contains 13427 camera images at a resolution of 1280x720 pixels and contains about 24000 annotated traffic lights.
0 papers · 0 benchmarks
Reflectance measurements of Bidirectional Texture Functions (BTFs) Database contains both flat samples: as well as 3D geometry with texture mapped BTFs: furthermore, there are some multispectral BTFs:
0 papers · 0 benchmarks
Bach chorales is a univariate time series based on chorales, where the task is to learn generative grammar.
0 papers · 0 benchmarks
A Filipino multi-modal language dataset for text+visual tasks.
0 papers · 0 benchmarks
This dataset presents a novel, multi-variate time series specifically designed for advancing research in spatio-temporal forecasting.
0 papers · 0 benchmarks
This repository contains datasets and baselines for benchmarking Chinese text recognition.
0 papers · 0 benchmarks
"Bend the Truth" dataset contains news in six different domains: technology, education, business, sports, politics, and entertainment.
0 papers · 0 benchmarks
This dataset contains 44,001 Bengali comments, curated to detect cyberbullying using Natural Language Processing (NLP) techniques.
0 papers · 0 benchmarks
The dataset consists of 3265 text samples corresponding to the concatenation of lines spoken by fictional characters.
0 papers · 0 benchmarks
This dataset consists of odometer or speedometer images of bike and car vehicles.
0 papers · 0 benchmarks
Annotated and original images of billboards in Japanese street scapes
0 papers · 0 benchmarks
The BirdVox-DCASE-20k dataset contains 20,000 ten-second audio recordings.
0 papers · 0 benchmarks
Objective This study introduces the BlendedICU dataset, a massive dataset of international intensive care data.
0 papers · 0 benchmarks
Supplemental Rcode with original results and images
0 papers · 0 benchmarks
Our networks are saved in GEXF as follows: • graphdimension1.gexf: Feed member interaction network saved in DiGraph object, where an edge has attributes sign and time and a node is Feed member.
0 papers · 0 benchmarks
This dataset comprises fractured and non-fractured X-ray images covering all anatomical body regions, including lower limb, upper limb, lumbar, hips, knees, etc.
0 papers · 0 benchmarks
This dataset consists of both fractured and non-fractured X-ray images encompassing various anatomical regions of the body, such as the lower limb, upper limb, lumbar region, hips, knees, and more.
0 papers · 0 benchmarks
This dataset consists of images of bottles and cups.
0 papers · 0 benchmarks
Boxy (Boxy Vehicles Dataset)
A large vehicle detection dataset with almost two million annotated vehicles for training and evaluating object detection methods for self-driving cars on freeways.
0 papers · 0 benchmarks
The Burmese Handwritten Digit Dataset (BHDD) is a dataset project specifically created for recognizing handwritten Burmese digits.
0 papers · 0 benchmarks
CAMEO (Continuous automated model evaluation)
Xavier Robin, Juergen Haas, Rafal Gumienny, Anna Smolinski, Gerardo Tauriello, and Torsten Schwede.Continuous automated model evaluation (cameo)—perspectives on the future of fully automated evaluation of structure prediction…
0 papers · 0 benchmarks
Medical report generation (MRG), which aims to automatically generate a textual description of a specific medical image (e.g., a chest X-ray), has recently received increasing research interest.
0 papers · 0 benchmarks
CBCT Walnut (Cone-Beam X-Ray CT Data Collection Designed for Machine Learning)
The scans are performed using a custom-built, highly flexible X-ray CT scanner, the FleX-ray scanner, developed by XRE nvand located in the FleX-ray Lab at the Centrum Wiskunde & Informatica (CWI) in Amsterdam, Netherlands.
0 papers · 0 benchmarks
CBLPRD-330k (China-Balanced-License-Plate-Recognition-Dataset-330k)
A high-quality, balanced dataset of 330,000 images featuring various types of Chinese license plates.
0 papers · 0 benchmarks
We introduce CCI4.0, a large-scale bilingual pre-training dataset engineered for superior data quality and diverse human-like reasoning trajectory.
0 papers · 0 benchmarks
CCIC (Concrete Crack Images for Classification)
The dataset contains concrete images having cracks.
0 papers · 0 benchmarks
CEAHB2021-5 (Chinese Ethnic Ancient Handwritten Books database)
Ancient books script identification of Chinese ethnic minorities with deep convolutional neural networks via multi-branch and spatial pyramid pooling Automatic classification of ancient books is an important component of the digital…
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.