12,172 datasets listed, ordered by the archive's paper count. Page 5 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
COPA (Choice of Plausible Alternatives)
The Choice Of Plausible Alternatives (COPA) evaluation provides researchers with a tool for assessing progress in open-domain commonsense causal reasoning.
329 papers · 1 benchmark
Multiple choice question answering based on the United States Medical License Exams (USMLE).
328 papers · 1 benchmark
The Multi-domain Wizard-of-Oz (MultiWOZ) dataset is a large-scale human-human conversational corpus spanning over seven domains, containing 8438 multi-turn dialogues, with each dialogue averaging 14 turns.
328 papers · 8 benchmarks
Animal FacesHQ (AFHQ) is a dataset of animal faces consisting of 15,000 high-quality images at 512 × 512 resolution.
327 papers · 6 benchmarks
MSVD (Microsoft Research Video Description Corpus)
The Microsoft Research Video Description Corpus (MSVD) dataset consists of about 120K sentences collected during the summer of 2010.
327 papers · 3 benchmarks
IMDB-BINARY is a movie collaboration dataset that consists of the ego-networks of 1,000 actors/actresses who played roles in movies in IMDB.
326 papers · 2 benchmarks
SMAC (The StarCraft Multi-Agent Challenge)
The StarCraft Multi-Agent Challenge (SMAC) is a benchmark that provides elements of partial observability, challenging dynamics, and high-dimensional observation spaces.
324 papers · 5 benchmarks
AffectNet is a large facial expression dataset with around 0.4 million images manually labeled for the presence of eight (neutral, happy, angry, sad, fear, surprise, disgust, contempt) facial expressions along with the intensity of valence…
323 papers · 4 benchmarks
This is a public domain speech dataset consisting of 13,100 short audio clips of a single speaker reading passages from 7 non-fiction books.
323 papers · 2 benchmarks
The PASCAL Context dataset is an extension of the PASCAL VOC 2010 detection challenge, and it contains pixel-wise labels for all training images.
323 papers · 6 benchmarks
ETT (Electricity Transformer Temperature)
The Electricity Transformer Temperature (ETT) is a crucial indicator in the electric power long-term deployment.
321 papers · 20 benchmarks
The THUMOS14 (THUMOS 2014) dataset is a large-scale video dataset that includes 1,010 videos for validation and 1,574 videos for testing from 20 classes.
318 papers · 18 benchmarks
The tieredImageNet dataset is a larger subset of ILSVRC-12 with 608 classes (779,165 images) grouped into 34 higher-level nodes in the ImageNet human-curated hierarchy.
317 papers · 7 benchmarks
DTU (DTU MVS dataset - 2014)
DTU MVS 2014 is a multi-view stereo dataset, which is an order of magnitude larger in number of scenes and with a significant increase in diversity.
313 papers · 2 benchmarks
The MPQA Opinion Corpus contains 535 news articles from a wide variety of news sources manually annotated for opinions and other private states (i.e., beliefs, emotions, sentiments, speculations, etc.).
313 papers · 3 benchmarks
BEIR (Benchmarking IR) is a heterogeneous benchmark containing different information retrieval (IR) tasks.
311 papers · 10 benchmarks
DRIVE (Digital Retinal Images for Vessel Extraction)
The Digital Retinal Images for Vessel Extraction (DRIVE) dataset is a dataset for retinal vessel segmentation.
311 papers · 2 benchmarks
PPI (Protein-Protein Interactions (PPI))
protein roles—in terms of their cellular functions from gene ontology—in various protein-protein interaction (PPI) graphs, with each graph corresponding to a different human tissue [41].
309 papers · 2 benchmarks
The CodeSearchNet Corpus is a large dataset of functions with associated documentation written in Go, Java, JavaScript, PHP, Python, and Ruby from open source projects on GitHub.
308 papers · 12 benchmarks
DAVIS17 is a dataset for video object segmentation.
308 papers · 12 benchmarks
HAR (Human Activity Recognition Using Smartphones)
The Human Activity Recognition Dataset has been collected from 30 subjects performing six different activities (Walking, Walking Upstairs, Walking Downstairs, Sitting, Standing, Laying).
307 papers · 3 benchmarks
Manga109 has been compiled by the Aizawa Yamasaki Matsui Laboratory, Department of Information and Communication Engineering, the Graduate School of Information Science and Technology, the University of Tokyo.
300 papers · 12 benchmarks
The Multi-PIE (Multi Pose, Illumination, Expressions) dataset consists of face images of 337 subjects taken under different pose, illumination and expressions.
299 papers · 1 benchmark
DOTA (Dataset for Object deTection in Aerial Images)
DOTA is a large-scale dataset for object detection in aerial images.
293 papers · 2 benchmarks
The LAMBADA (LAnguage Modeling Broadened to Account for Discourse Aspects) benchmark is an open-ended cloze task which consists of about 10,000 passages from BooksCorpus where a missing target word is predicted in the last sentence of each…
293 papers · 1 benchmark
MOT17 (Multiple Object Tracking 17)
The Multiple Object Tracking 17 (MOT17) dataset is a dataset for multiple object tracking.
291 papers · 2 benchmarks
StrategyQA is a question answering benchmark where the required reasoning steps are implicit in the question, and should be inferred using a strategy.
291 papers · 1 benchmark
DocVQA consists of 50,000 questions defined on 12,000+ document images.
290 papers · 3 benchmarks
The 20BN-SOMETHING-SOMETHING V2 dataset is a large collection of labeled video clips that show humans performing pre-defined basic actions with everyday objects.
290 papers · 7 benchmarks
MELD (Multimodal EmotionLines Dataset)
Multimodal EmotionLines Dataset (MELD) has been created by enhancing and extending EmotionLines dataset.
289 papers · 3 benchmarks
MPI-INF-3DHP is a 3D human body pose estimation dataset consisting of both constrained indoor and complex outdoor scenes.
289 papers · 5 benchmarks
Clothing1M contains 1M clothing images in 14 classes.
288 papers · 4 benchmarks
WMT 2014 is a collection of datasets used in shared tasks of the Ninth Workshop on Statistical Machine Translation.
288 papers · 9 benchmarks
The Adversarial Natural Language Inference (ANLI, Nie et al.) is a new large-scale NLI benchmark dataset, collected via an iterative, adversarial human-and-model-in-the-loop procedure.
287 papers · 3 benchmarks
DUTS is a saliency detection dataset containing 10,553 training images and 5,019 test images.
286 papers · 5 benchmarks
HowTo100M is a large-scale dataset of narrated videos with an emphasis on instructional videos where content creators teach complex tasks with an explicit intention of explaining the visual content on screen.
286 papers · 1 benchmark
AirSim is a simulator for drones, cars and more, built on Unreal Engine.
285 papers · 0 benchmarks
CoQA (Conversational Question Answering Challenge)
CoQA is a large-scale dataset for building Conversational Question Answering systems.
281 papers · 2 benchmarks
AudioCaps is a dataset of sounds with event descriptions that was introduced for the task of audio captioning, with sounds sourced from the AudioSet dataset.
279 papers · 6 benchmarks
The efforts to create a non-trivial and publicly available dataset for action recognition was initiated at the KTH Royal Institute of Technology in 2004.
279 papers · 2 benchmarks
Charts are very popular for analyzing data.
278 papers · 1 benchmark
The Shanghaitech dataset is a large-scale crowd counting dataset.
277 papers · 5 benchmarks
The task of PubMedQA is to answer research questions with yes/no/maybe (e.g.: Do preoperative statins reduce atrial fibrillation after coronary artery bypass grafting?) using the corresponding abstracts.
276 papers · 3 benchmarks
LaSOT (Large-scale Single Object Tracking)
LaSOT is a high-quality benchmark for Large-scale Single Object Tracking.
275 papers · 3 benchmarks
MSMT17 (Multi Scene Multi Time dataset for person re-id)
MSMT17 is a multi-scene multi-time person re-identification dataset.
275 papers · 6 benchmarks
In particular, MUTAG is a collection of nitroaromatic compounds and the goal is to predict their mutagenicity on Salmonella typhimurium.
274 papers · 3 benchmarks
The NewsQA dataset is a crowd-sourced machine reading comprehension dataset of 120,000 question-answer pairs.
272 papers · 1 benchmark
PASCAL-S is a dataset for salient object detection consisting of a set of 850 images from PASCAL VOC 2010 validation set with multiple salient objects on the scenes.
271 papers · 3 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.