Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 108 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 5137–5184 of 12,172
The ASR-GLUE benchmark is a collection of 6 different NLU (Natural Language Understanding) tasks for evaluating the performance of models under automatic speech recognition (ASR) error across 3 different levels of background noise and 6…
3 papers · 0 benchmarks
A crowdsourced multi-reference corpus where each simplification was produced by executing several rewriting transformations.
3 papers · 0 benchmarks
ATM'22 is a multi-site, multi-domain dataset for pulmonary airway segmentation.
3 papers · 0 benchmarks
This is a simple audio-visual dataset artificially assembled from independent visual and audio datasets.
3 papers · 0 benchmarks
Audio-visual question answering aims to answer questions regarding both audio and visual modalities in a given video.
3 papers · 0 benchmarks
All Words Open IE (AW-OIE) is an open information extraction dataset derived from Question-Answer Meaning Representation (QAMR) dataset.
3 papers · 0 benchmarks
The Action-Camera Parking Dataset contains 293 images captured at a roughly 10-meter height using a GoPro Hero 6 camera.
3 papers · 1 benchmark
ActivityNet Adverbs is a subset from the ActivityNet dataset with extracted verb-adverb annotations.
3 papers · 2 benchmarks
AdobeVFR real (Adobe Visual Font Recognition real-world images dataset)
Subset of AdobeVFR.
3 papers · 1 benchmark
AdobeVFR syn (Adobe Visual Font Recognition synthetic dataset)
Subset of AdobeVFR.
3 papers · 1 benchmark
AeroPath (AeroPath: An airway segmentation benchmark dataset with challenging pathology)
Public benchmark dataset (AeroPath), consisting of 27 CT images from patients with pathologies ranging from emphysema to large tumors, with corresponding trachea and bronchi annotations.
3 papers · 0 benchmarks
AfriQA is a cross-lingual QA dataset with a focus on African languages.
3 papers · 0 benchmarks
The AQI dataset is collected from 12 observing stations around Beijing from year 2013 to 2017.
3 papers · 0 benchmarks
The Aircraft Context Dataset, a composition of two inter-compatible large-scale and versatile image datasets focusing on manned aircraft and UAVs, is intended for training and evaluating classification, detection and segmentation models in…
3 papers · 0 benchmarks
The Algonauts dataset provides human brain responses to a set of 1,102 3-s long video clips of everyday events.
3 papers · 0 benchmarks
Ali-CCP (Alibaba Click and Conversion Prediction)
This data set is provided by Alimama
3 papers · 0 benchmarks
We design an all-day semantic segmentation benchmark all-day CityScapes.
3 papers · 1 benchmark
A comprehensive multi-task benchmark for the Polish language understanding, accompanied by an online leaderboard.
3 papers · 0 benchmarks
Amateur Drawings is a dataset collected via the public demo of Animated Drawings, containing over 178,000 amateur drawings and corresponding user-accepted character bounding boxes, segmentation masks, and joint location annotations.
3 papers · 0 benchmarks
This dataset includes reviews (ratings, text, helpfulness votes), product metadata (descriptions, category information, price, brand, and image features), and links (also viewed/also bought graphs).
3 papers · 1 benchmark
This dataset includes reviews (ratings, text, helpfulness votes), product metadata (descriptions, category information, price, brand, and image features), and links (also viewed/also bought graphs).
3 papers · 1 benchmark
Is a new open-domain question answering task which involves predicting a set of question-answer pairs, where every plausible answer is paired with a disambiguated rewrite of the original question.
3 papers · 0 benchmarks
It contains about 28K medium quality animal images belonging to 10 categories: dog, cat, horse, spyder, butterfly, chicken, sheep, cow, squirrel, and elephant.
3 papers · 1 benchmark
A dataset for 2D pose estimation of anime/manga images.
3 papers · 0 benchmarks
AnoVox is a large-scale benchmark for ANOmaly detection in autonomous driving.
3 papers · 0 benchmarks
For a detailed description, we refer to Section 3 in our research article.
3 papers · 0 benchmarks
The data consist of 70 records, divided into a learning set of 35 records (a01 through a20, b01 through b05, and c01 through c10), and a test set of 35 records (x01 through x35), all of which may be downloaded from this page.
3 papers · 1 benchmark
A benchmark Arabic dataset for commonsense understanding and validation as well as a baseline research and models trained using the same dataset.
3 papers · 0 benchmarks
A dataset (in English; and also extended to Hindi) with human-written navigation and assembling instructions, and the corresponding ground truth trajectories.
3 papers · 0 benchmarks
Arxiv GR-QC (General Relativity and Quantum Cosmology collaboration network)
Arxiv GR-QC (General Relativity and Quantum Cosmology) collaboration network is from the e-print arXiv and covers scientific collaborations between authors papers submitted to General Relativity and Quantum Cosmology category.
3 papers · 0 benchmarks
ArzEn (Corpus of Egyptian Arabic-English Code-switching)
Corpus of Egyptian Arabic-English Code-switching (ArzEn) is a spontaneous conversational speech corpus, obtained through informal interviews held at the German University in Cairo.
3 papers · 0 benchmarks
Atlas is a dataset for e-commerce clothing product categorization.
3 papers · 0 benchmarks
AugMod (AugMod: pythagore-mod-reco)
Context A radio signal consists in two channels, channel I (for 'In phase') and channel Q (for 'Quadrature') and can be assimilated as a stream of complex numbers.
3 papers · 0 benchmarks
The AxonEM dataset consists of two 30x30x30 um^3 EM image volumes from the human and mouse cortex, respectively.
3 papers · 0 benchmarks
This is a set of files representing part of the workload of Microsoft's Azure Functions offering, collected in July of 2019.
3 papers · 0 benchmarks
BAFMD (Bias-Aware Face Mask Detection Dataset)
BAFMD contains images posted on Twitter during the pandemic from around the world with more images from underrepresented race and age groups to mitigate the problem for the face mask detection task.
3 papers · 0 benchmarks
BB-MAS (Behavioural Biometrics Multi-device and multi-Activity data from Same users)
BB-MAS is a behavioural biometrics dataset.
3 papers · 0 benchmarks
This image set is part of a high-throughput chemical screen on U2OS cells, with examples of 200 bioactive compounds.
3 papers · 0 benchmarks
BCOPA-CE (A Balanced COPA Test Set with cause-effect as alternatives)
We provide the BCOPA-CE test set, which has balanced token distribution in the correct and wrong alternatives and increases the difficulty of being aware of cause and effect.
3 papers · 0 benchmarks
BGP (Border Gateway Protocol (BGP) Network)
Border Gateway Protocol (BGP) Network describes the Internet's inter-domain structure, where nodes represent the autonomous systems and edges are the business relationships between nodes.
3 papers · 1 benchmark
In an effort to catalog insect biodiversity, we propose a new large dataset of hand-labelled insect images, the BIOSCAN-1M Insect Dataset.
3 papers · 1 benchmark
BLEFF (Blender Forward Facing Dataset)
Synthetic (Blender) Dataset for forward facing scenes Toe vaualte NVS quality and camera parameter accuracy.
3 papers · 1 benchmark
BRIGHT is the first open-access, globally distributed, event-diverse multimodal dataset specifically curated to support AI-based disaster response.
3 papers · 1 benchmark
BRIND is a short name of BSDS-RIND is the first public benchmark that dedicated to studying simultaneously the four edge types, namely Reflectance Edge (RE), Illumination Edge (IE), Normal Edge (NE) and Depth Edge (DE)
3 papers · 1 benchmark
BRUSH (Brown University Stylus Handwriting)
The BRUSH dataset (BRown University Stylus Handwriting) contains 27,649 online handwriting samples from a total of 170 writers.
3 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.