Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 53 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 2497–2544 of 12,172
2021 Hotel-ID is a dataset for hotel recognition to help raise awareness of human trafficking and generate novel approaches.
14 papers · 0 benchmarks
4DFAB is a large scale database of dynamic high-resolution 3D faces which consists of recordings of 180 subjects captured in four different sessions spanning over a five-year period (2012 - 2017), resulting in a total of over 1,800,000 3D…
14 papers · 0 benchmarks
ACL Anthology Reference Corpus (ACL ARC) is a collection of 10,920 academic papers from the ACL Anthology.
14 papers · 3 benchmarks
A large-scale dataset named AIC (AI Challenger) with three sub-datasets, human keypoint detection (HKD), large-scale attribute dataset (LAD) and image Chinese captioning (ICC).
14 papers · 1 benchmark
ARAD-1K (Ntire 2022 spectral recovery challenge and data set)
The dataset used for NTIRE 2022 Spectral Recovery Challenge
14 papers · 1 benchmark
AVSD (Audio-Visual Scene-Aware Dialog)
The Audio Visual Scene-Aware Dialog (AVSD) dataset, or DSTC7 Track 3, is a audio-visual dataset for dialogue understanding.
14 papers · 1 benchmark
Development of a benchmark corpus to support the automatic extraction of drug-related adverse effects from medical case reports.
14 papers · 3 benchmarks
This datasets is a subset of the Amazon reviews dataset which contain Men related products
14 papers · 2 benchmarks
Amazon-Sports is a sub-category of the Amazon dataset, which contains a series of product reviews crawled from Amazon.com.
14 papers · 1 benchmark
AmazonQA consists of 923k questions, 3.6M answers and 14M reviews across 156k products.
14 papers · 0 benchmarks
ArSarcasm is a new Arabic sarcasm detection dataset.
14 papers · 0 benchmarks
BabyLM is a dataset for small scale language modeling, human language acquisition, low-resource NLP, and cognitive modeling.
14 papers · 0 benchmarks
BioLAMA is a benchmark comprised of 49K biomedical factual knowledge triples for probing biomedical Language Models.
14 papers · 0 benchmarks
BoLD (Body Language Dataset)
Yu Luo, Jianbo Ye, Reginald B.
14 papers · 1 benchmark
A novel dataset for traffic accidents analysis.
14 papers · 0 benchmarks
Benchmark for HMER and OHMER Source: CROHME 2014
14 papers · 1 benchmark
Source: ICFHR2016 CROHME: Competition on Recognition of Online Handwritten Mathematical Expressions
14 papers · 1 benchmark
The Collaborative Drawing game (CoDraw) dataset contains ~10K dialogs consisting of ~138K messages exchanged between human players in the CoDraw game.
14 papers · 0 benchmarks
A SemEval shared task in which participants must extract definitions from free text using a term-definition pair corpus that reflects the complex reality of definitions in natural language.
14 papers · 0 benchmarks
The contest of binarization using a popular document database was organized called as Document Image Binarization Contest (DIBCO) from 2009 to 2019, except for 2015.
14 papers · 0 benchmarks
Contains biases but is two orders of magnitude larger than those used previously.
14 papers · 0 benchmarks
DPD dataset has two versions - single view and dual-view.
14 papers · 0 benchmarks
DadaGP is a new symbolic music dataset comprising 26,181 song scores in the GuitarPro format covering 739 musical genres, along with an accompanying tokenized format well-suited for generative sequence models such as the Transformer.
14 papers · 0 benchmarks
DailyTalk is a high-quality conversational speech dataset designed for Text-to-Speech.
14 papers · 0 benchmarks
The Dakshina dataset is a collection of text in both Latin and native scripts for 12 South Asian languages.
14 papers · 0 benchmarks
DeepStab is a dataset for online video stabilization consisting of synchronized steady/unsteady video pairs collected via a well designed hand-held hardware.
14 papers · 0 benchmarks
DuLeMon (Baidu Long-term Memory Conversation)
DuLeMon is a large-scale Chinese Long-term Memory Conversation dataset, which simulates long-term memory conversations and focuses on the ability to actively construct and utilize the user's and the bot's persona in a long-term interaction.
14 papers · 0 benchmarks
ELPV (A dataset of functional and defective solar cells extracted from EL images of solar modules)
The dataset contains 2,624 samples of 300×300 pixels 8-bit grayscale images of functional and defective solar cells with varying degree of degradations extracted from 44 different solar modules.
14 papers · 0 benchmarks
Fashionpedia consists of two parts: (1) an ontology built by fashion experts containing 27 main apparel categories, 19 apparel parts, 294 fine-grained attributes and their relationships; (2) a dataset with everyday and celebrity event…
14 papers · 0 benchmarks
Node classification on Film with the fixed 48%/32%/20% splits provided by Geom-GCN.
14 papers · 2 benchmarks
Fisheye dataset comprises of synthetically generated fisheye sequences and fisheye video sequences captured with an actual fisheye camera designed for fisheye motion estimation.
14 papers · 0 benchmarks
Food2K is a large food recognition dataset with 2,000 categories and over 1 million images.
14 papers · 0 benchmarks
GTOS (Ground Terrain in Outdoor Scenes)
The database consists of over 30,000 images covering 40 classes of outdoor ground terrain under varying weather and lighting conditions.
14 papers · 0 benchmarks
GlobalOpinionQA consists of questions and answers from cross-national surveys designed to capture diverse opinions on global issues across different countries.
14 papers · 0 benchmarks
GraphQuestions is a characteristic-rich dataset designed for factoid question answering.
14 papers · 2 benchmarks
We present a comprehensive framework for egocentric interaction recognition using markerless 3D annotations of two hands manipulating objects.
14 papers · 2 benchmarks
HAA500 (Human-Centric Atomic Action Dataset)
HAA500 is a manually annotated human-centric atomic action dataset for action recognition on 500 classes with over 591k labeled frames.
14 papers · 1 benchmark
The data has been produced using Monte Carlo simulations.
14 papers · 1 benchmark
The human-Related version of the ShanghaiTech Campus, was first presented by Morais et al.
14 papers · 1 benchmark
HiFiMask is a large-scale High-Fidelity Mask dataset, namely CASIA-SURF HiFiMask (briefly HiFiMask).
14 papers · 0 benchmarks
The IIIT5K dataset contains 5,000 text instance images: 2,000 for training and 3,000 for testing.
14 papers · 1 benchmark
Dataset provided by the Image Matching Workshop https://www.cs.ubc.ca/research/image-matching-challenge/current/
14 papers · 1 benchmark
The IndoNLU benchmark is a collection of resources for training, evaluating, and analyzing natural language understanding systems for Bahasa Indonesia.
14 papers · 2 benchmarks
Inter-X is a large-scale dataset containing ~11K interaction sequences, more than 8.1M frames and 34K fine-grained human textual descriptions.
14 papers · 1 benchmark
The KITTI-Depth dataset includes depth maps from projected LiDAR point clouds that were matched against the depth estimation from the stereo cameras.
14 papers · 0 benchmarks
KUAKE-QIC (Query Intent Classification Dataset)
KUAKE Query Intent Classification, a dataset for intent classification, is used for the KUAKE-QIC task.
14 papers · 1 benchmark
LAV-DF (Localized Audio Visual DeepFake Dataset)
Localized Audio Visual DeepFake Dataset (LAV-DF).
14 papers · 1 benchmark
LIVE-YT-HFR comprises of 480 videos having 6 different frame rates, obtained from 16 diverse contents.
14 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.