12,172 datasets listed, ordered by the archive's paper count. Page 12 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
ORL (Our Database of Faces)
The ORL Database of Faces contains 400 images from 40 distinct subjects.
129 papers · 1 benchmark
TabFact is a large-scale dataset which consists of 117,854 manually annotated statements with regard to 16,573 Wikipedia tables, their relations are classified as ENTAILED and REFUTED.
129 papers · 2 benchmarks
VOT2018 is a dataset for visual object tracking.
129 papers · 1 benchmark
Europarl (European Parliament Proceedings Parallel Corpus)
A corpus of parallel text in 21 European languages from the proceedings of the European Parliament.
128 papers · 1 benchmark
The Free Music Archive (FMA) is a large-scale dataset for evaluating several tasks in Music Information Retrieval.
128 papers · 2 benchmarks
LFPW (Labeled Face Parts in the Wild)
The Labeled Face Parts in-the-Wild (LFPW) consists of 1,432 faces from images downloaded from the web using simple text queries on sites such as google.com, flickr.com, and yahoo.com.
128 papers · 0 benchmarks
The Meta-Dataset benchmark is a large few-shot learning benchmark and consists of multiple datasets of different data distributions.
128 papers · 2 benchmarks
We observe that satellite imagery is a powerful source of information as it contains more structured and uniform data, compared to traditional images.
127 papers · 1 benchmark
LogiQA consists of 8,678 QA instances, covering multiple types of deductive reasoning.
127 papers · 1 benchmark
NELL-995 KG Completion Dataset
127 papers · 2 benchmarks
SMAP (Soil Moisture Active Passive)
Soil Moisture Active Passive (SMAP) dataset is a dataset of soil samples and telemetry information using the Mars rover by NASA.
127 papers · 2 benchmarks
WNUT 2017 (WNUT 2017 Emerging and Rare entity recognition)
This shared task focuses on identifying unusual, previously-unseen entities in the context of emerging discussions.
127 papers · 2 benchmarks
WikiHow is a dataset of more than 230,000 article and summary pairs extracted and constructed from an online knowledge base written by different human authors.
127 papers · 2 benchmarks
BLiMP (Benchmark of Linguistic Minimal Pairs)
BLiMP is a challenge set for evaluating what language models (LMs) know about major grammatical phenomena in English.
126 papers · 0 benchmarks
DFDC (Deepfake Detection Challenge)
The DFDC (Deepfake Detection Challenge) is a dataset for deepface detection consisting of more than 100,000 videos.
126 papers · 1 benchmark
FBMS (Freiburg-Berkeley Motion Segmentation)
The Freiburg-Berkeley Motion Segmentation Dataset (FBMS-59) is an extension of the BMS dataset with 33 additional video sequences.
126 papers · 1 benchmark
LSMDC (Large Scale Movie Description Challenge)
This dataset contains 118,081 short video clips extracted from 202 movies.
126 papers · 3 benchmarks
PASCAL VOC 2007 is a dataset for image recognition.
126 papers · 13 benchmarks
SUN3D contains a large-scale RGB-D video database, with 8 annotated sequences.
126 papers · 0 benchmarks
Aff-Wild is a large-scale in-the-wild dataset for valence-arousal estimation from videos with a variety of head poses, illumination conditions and occlusions.
125 papers · 0 benchmarks
FreiHAND is a 3D hand pose dataset which records different hand actions performed by 32 people.
125 papers · 1 benchmark
The LUNA challenges provide datasets for automatic nodule detection algorithms using the largest publicly available reference database of chest CT scans, the LIDC-IDRI data set.
125 papers · 2 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
125 papers · 0 benchmarks
The Yelp2018 dataset is adopted from the 2018 edition of the yelp challenge.
125 papers · 2 benchmarks
FER+ (Face Expression Recognition Plus dataset)
The FER+ dataset is an extension of the original FER dataset, where the images have been re-labelled into one of 8 emotion types: neutral, happiness, surprise, sadness, anger, disgust, fear, and contempt.
124 papers · 3 benchmarks
The MSRA-TD500 dataset is a text detection dataset that contains 300 training images and 200 test images.
124 papers · 1 benchmark
The Microsoft Academic Graph is a heterogeneous graph containing scientific publication records, citation relationships between those publications, as well as authors, institutions, journals, conferences, and fields of study.
124 papers · 0 benchmarks
NCLT (North Campus Long-Term Vision and LiDAR)
The NCLT dataset is a large scale, long-term autonomy dataset for robotics research collected on the University of Michigan’s North Campus.
124 papers · 0 benchmarks
CARER (Contextualized Affect Representations for Emotion Recognition)
CARER is an emotion dataset collected through noisy labels, annotated via distant supervision as in (Go et al., 2009).
123 papers · 2 benchmarks
JFT-300M is an internal Google dataset used for training image classification models.
123 papers · 1 benchmark
PubLayNet is a dataset for document layout analysis by automatically matching the XML representations and the content of over 1 million PDF articles that are publicly available on PubMed Central.
123 papers · 1 benchmark
BBQ (Bias Benchmark for QA)
Bias Benchmark for QA (BBQ) is a dataset consisting of question-sets constructed by the authors that highlight attested social biases against people belonging to protected classes along nine different social dimensions relevant for U.S.
122 papers · 0 benchmarks
LibriMix is an open-source alternative to wsj0-2mix.
122 papers · 1 benchmark
MAWPS (MAth Word ProblemS)
MAWPS is an online repository of Math Word Problems, to provide a unified testbed to evaluate different algorithms.
122 papers · 1 benchmark
Multi-News, consists of news articles and human-written summaries of these articles from the site newser.com.
122 papers · 5 benchmarks
ETHD is a multi-view stereo benchmark / 3D reconstruction benchmark that covers a variety of indoor and outdoor scenes.
121 papers · 3 benchmarks
The GENIA corpus is the primary collection of biomedical literature compiled and annotated within the scope of the GENIA project.
121 papers · 7 benchmarks
Permuted MNIST is an MNIST variant that consists of 70,000 images of handwritten digits from 0 to 9, where 60,000 images are used for training, and 10,000 images for test.
121 papers · 1 benchmark
The dataset for the SemEval-2010 Task 8 is a dataset for multi-way classification of mutually exclusive semantic relations between pairs of nominals.
121 papers · 1 benchmark
GTEA (Georgia Tech Egocentric Activity)
The Georgia Tech Egocentric Activities (GTEA) dataset contains seven types of daily activities such as making sandwich, tea, or coffee.
120 papers · 2 benchmarks
Raindrop is a set of image pairs, where each pair contains exactly the same background scene, yet one is degraded by raindrops and the other one is free from raindrops.
120 papers · 1 benchmark
SHAPES (Swarm Heuristics based Adaptive and Penalized Estimation of Splines)
SHAPES is a dataset of synthetic images designed to benchmark systems for understanding of spatial and logical relations among multiple objects.
120 papers · 1 benchmark
SIQA (Social Interaction QA)
Social Interaction QA (SIQA) is a question-answering benchmark for testing social commonsense intelligence.
120 papers · 1 benchmark
SimpleQuestions is a large-scale factoid question answering dataset.
120 papers · 2 benchmarks
The 'shape bias' dataset was introduced in Geirhos et al.
120 papers · 1 benchmark
SEED (SJTU Emotion EEG Dataset)
The SEED dataset contains subjects' EEG signals when they were watching films clips.
119 papers · 4 benchmarks
The MAESTRO dataset contains over 200 hours of paired audio and MIDI recordings from ten years of International Piano-e-Competition.
118 papers · 1 benchmark
VATEX is multilingual, large, linguistically complex, and diverse dataset in terms of both video and natural language descriptions.
118 papers · 3 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.