Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 236 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 11281–11328 of 12,172
Allergen30 is created with the goal of building a robust detection model that can assist people in avoiding possible allergic reactions.
0 papers · 0 benchmarks
A new dataset for sentiment analysis, scraped from Allociné.fr user reviews.
0 papers · 0 benchmarks
This dataset provides full historical daily stock price for Alphabet.
0 papers · 0 benchmarks
Dataset for the paper entitled "Efficient Concept Drift Handling for Batch Android Malware Detection Models".
0 papers · 0 benchmarks
AntM2C (Ant-Group Multi-Scenario Multi-Modal CTR dataset)
We release a large-scale Multi-Scenario Multi-Modal CTR dataset named AntM2C, built from real industrial data from Alipay.
0 papers · 0 benchmarks
The photo fixation of apple fruitlets was done in the LatHort orchard in Dobele, at the development of fruit (BBCH stage 76-78).
0 papers · 0 benchmarks
The photo fixation of apple fruits was done in the LatHort orchard in Dobele, at the maturity of fruit and seed (BBCH stage 81-85).
0 papers · 0 benchmarks
Dataset contains images with apples infected by scab.
0 papers · 0 benchmarks
Dataset contains images with apple leaves infected by scab.
0 papers · 0 benchmarks
Global epidemics, like COVID-19, have substantial impacts on almost all countries in multiple aspects, such as economy, hospitalization, lifestyle, etc1, 2.
0 papers · 0 benchmarks
A dataset to identify the meters of Arabic poems.
0 papers · 0 benchmarks
Contain Arabic handwritten digits images (60000 training and 10000 testing images).
0 papers · 0 benchmarks
The Arabic Speech Corpus (1.5 GB) is a Modern Standard Arabic (MSA) speech corpus for speech synthesis.
0 papers · 0 benchmarks
It includes 150M (150,211,934) Arabic Web pages.
0 papers · 0 benchmarks
ArcBench is a logically challenging dataset of 158 English question–answer pairs, derived from the RoR-Bench benchmark.
0 papers · 0 benchmarks
The Arena-Hard benchmark is a high-quality benchmarking tool for Language Learning Models (LLMs) developed by LMSYS Org¹.
0 papers · 0 benchmarks
This dataset contains video shots for two different classes: tigers and cars.
0 papers · 0 benchmarks
AstroSmartphoneDataset, is a collection of astronomical images captured with smartphones with Google Camera (GCam), a Google mobile application allowing to take long-exposure astronomical images with a large field of view.
0 papers · 0 benchmarks
The subset of audio samples from the AudioSet ontology which are licensed with Creative Commons.
0 papers · 0 benchmarks
Data files with the information required to replicate all the experiments reported in the paper: Linares López, Carlos; Herman, Ian, 2024.
0 papers · 0 benchmarks
This dataset is comprised of the dynamic analysis reports generated by CAPEv2, from both malware and goodware.
0 papers · 0 benchmarks
Dataset for automatic keyphrase extraction task.
0 papers · 0 benchmarks
This dataset contains 10089 Bengali comments and its tag( Nostalgic and Non-nostalgic)
0 papers · 0 benchmarks
Dataset for reproducing the code of the work: Estimation of Semiconductor Power Losses Through Automatic Thermal Modeling.
0 papers · 0 benchmarks
This dataset is an extremely challenging set of over 8000+ original Fire and Smoke images captured and crowdsourced from over 1200+ urban and rural areas, where each image is manually reviewed and verified by computer vision professionals…
0 papers · 0 benchmarks
An axial turbine is a simplest hydrulic machine which is suitable for low-head conditions.
0 papers · 0 benchmarks
BABEL is a multilingual corpus of conversational telephone speech from IARPA, which includes Asian and African languages.
0 papers · 0 benchmarks
BACC-18 (Bengali Authorship Classification Corpus-18)
The developed BACC-18 contains the text of 18 famous authors of Bengali literature.
0 papers · 0 benchmarks
BAVL (Blind Audio-Visual Localization (BAVL))
Blind Audio-Visual Localization (BAVL) Dataset consists of 20 audio-visual recordings of sound sources, which could be talking faces or music instruments.
0 papers · 0 benchmarks
A new large-scale baseball video dataset which is produced semi-automatically by using play-by-play texts available online.
0 papers · 0 benchmarks
BEHAVIOR is a benchmark with the 100 household activities that represent a new challenge for embodied AI solutions.
0 papers · 0 benchmarks
BI2012 MOABB (P300 dataset BI2012 from a "Brain Invaders" experiment.)
0 papers · 0 benchmarks
BI2013a MOABB (P300 dataset BI2013a from a "Brain Invaders" experiment.)
0 papers · 0 benchmarks
BI2014a MOABB (P300 dataset BI2014a from a "Brain Invaders" experiment.)
0 papers · 0 benchmarks
BI2014b MOABB (P300 dataset BI2014b from a "Brain Invaders" experiment.)
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.