12,172 datasets listed, ordered by the archive's paper count. Page 4 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Object Tracking Benchmark (OTB) is a visual tracking benchmark that is widely used to evaluate the performance of a visual tracking algorithm.
416 papers · 1 benchmark
The CASIA-WebFace dataset is used for face verification and face identification tasks.
415 papers · 2 benchmarks
The Replica Dataset is a dataset of high quality reconstructions of a variety of indoor spaces.
414 papers · 4 benchmarks
GTA5 (Grand Theft Auto 5)
The GTA5 dataset contains 24966 synthetic images with pixel level semantic annotation.
412 papers · 7 benchmarks
RACE (ReAding Comprehension dataset from Examinations)
The ReAding Comprehension dataset from Examinations (RACE) dataset is a machine reading comprehension dataset consisting of 27,933 passages and 97,867 questions from English exams, targeting Chinese students aged 12-18.
412 papers · 3 benchmarks
MVTecAD (MVTEC ANOMALY DETECTION DATASET)
MVTec AD is a dataset for benchmarking anomaly detection methods with a focus on industrial inspection.
402 papers · 4 benchmarks
Caltech-256 is an object recognition dataset containing 30,607 real-world images, of different sizes, spanning 257 classes (256 object classes and an additional clutter class).
401 papers · 4 benchmarks
DailyDialog is a high-quality multi-turn open-domain English dialog dataset.
399 papers · 2 benchmarks
DeepFashion is a dataset containing around 800K diverse fashion images with their rich annotations (46 categories, 1,000 descriptive attributes, bounding boxes and landmark information) ranging from well-posed product images to…
397 papers · 5 benchmarks
The 3D Poses in the Wild dataset is the first dataset in the wild with accurate 3D poses for evaluation.
395 papers · 4 benchmarks
WN18RR is a link prediction dataset created from WN18, which is a subset of WordNet.
394 papers · 3 benchmarks
Objaverse is a large dataset of objects with 800K+ (and growing) 3D models with descriptive captions, tags, and animations.
393 papers · 2 benchmarks
Speech Commands is an audio dataset of spoken words designed to help train and evaluate keyword spotting systems .
392 papers · 4 benchmarks
The GoPro dataset for deblurring consists of 3,214 blurred images with the size of 1,280×720 that are divided into 2,103 training images and 1,111 test images.
390 papers · 4 benchmarks
The ESC-50 dataset is a labeled collection of 2000 environmental audio recordings suitable for benchmarking methods of environmental sound classification.
387 papers · 4 benchmarks
Argoverse is a tracking benchmark with over 30K scenarios collected in Pittsburgh and Miami.
386 papers · 6 benchmarks
MMBench is a multi-modality benchmark.
384 papers · 1 benchmark
DROP (Discrete Reasoning Over Paragraphs)
Discrete Reasoning Over Paragraphs DROP is a crowdsourced, adversarially-created, 96k-question benchmark, in which a system must resolve references in a question, perhaps to multiple input positions, and perform discrete operations over…
382 papers · 3 benchmarks
The CiteSeer dataset consists of 3312 scientific publications classified into one of six classes.
381 papers · 13 benchmarks
GTSRB (German Traffic Sign Recognition Benchmark)
The German Traffic Sign Recognition Benchmark (GTSRB) contains 43 classes of traffic signs, split into 39,209 training images and 12,630 test images.
374 papers · 5 benchmarks
PROTEINS is a dataset of proteins that are classified as enzymes or non-enzymes.
371 papers · 1 benchmark
Netflix Prize consists of about 100,000,000 ratings for 17,770 movies given by 480,189 users.
370 papers · 1 benchmark
FaceForensics++ is a forensics dataset consisting of 1000 original video sequences that have been manipulated with four automated face manipulation methods: Deepfakes, Face2Face, FaceSwap and NeuralTextures.
368 papers · 2 benchmarks
OK-VQA (Outside Knowledge Visual Question Answering)
Outside Knowledge Visual Question Answering (OK-VQA) includes more than 14,000 questions that require external knowledge to answer.
368 papers · 2 benchmarks
AMASS is a large database of human motion unifying different optical marker-based motion capture datasets by representing them within a common framework and parameterization.
366 papers · 1 benchmark
The Arcade Learning Environment (ALE) is an object-oriented framework that allows researchers to develop AI agents for Atari 2600 games.
366 papers · 57 benchmarks
Visual Question Answering (VQA) v2.0 is a dataset containing open-ended questions about images.
366 papers · 7 benchmarks
The DeepMind Control Suite (DMCS) is a set of simulated continuous control environments with a standardized structure and interpretable rewards.
364 papers · 3 benchmarks
SVAMP (Simple Variations on Arithmetic Math word Problems)
A challenge set for elementary-level Math Word Problems (MWP).
362 papers · 2 benchmarks
WSC (Winograd Schema Challenge)
The Winograd Schema Challenge was introduced both as an alternative to the Turing Test and as a test of a system’s ability to do commonsense reasoning.
361 papers · 2 benchmarks
YAGO (Yet Another Great Ontology)
Yet Another Great Ontology (YAGO) is a Knowledge Graph that augments WordNet with common knowledge facts extracted from Wikipedia, converting WordNet from a primarily linguistic resource to a common knowledge base.
359 papers · 5 benchmarks
BIG-Bench Hard (BBH) is a subset of the BIG-Bench, a diverse evaluation suite for language models.
352 papers · 2 benchmarks
Automatic image captioning is the task of producing a natural-language utterance (usually a sentence) that correctly reflects the visual content of an image.
352 papers · 2 benchmarks
BIG-bench (Beyond the Imitation Game Benchmark)
The Beyond the Imitation Game Benchmark (BIG-bench) is a collaborative benchmark intended to probe large language models and extrapolate their future capabilities.
349 papers · 121 benchmarks
XNLI (Cross-lingual Natural Language Inference)
The Cross-lingual Natural Language Inference (XNLI) corpus is the extension of the Multi-Genre NLI (MultiNLI) corpus to 15 languages.
349 papers · 7 benchmarks
The NUS-WIDE dataset contains 269,648 images with a total of 5,018 tags collected from Flickr.
348 papers · 3 benchmarks
SICK (Sentences Involving Compositional Knowledge)
The Sentences Involving Compositional Knowledge (SICK) dataset is a dataset for compositional distributional semantics.
348 papers · 5 benchmarks
Mip-NeRF 360 (Unbounded Anti-Aliased Neural Radiance Fields)
Mip-NeRF 360 is an extension to the Mip-NeRF that uses a non-linear parameterization, online distillation, and a novel distortion-based regularize to overcome the challenge of unbounded scenes.
347 papers · 1 benchmark
BookCorpus is a large collection of free novel books written by unpublished authors, which contains 11,038 books (around 74M sentences and 1G words) of 16 different sub-genres (e.g., Romance, Historical, Adventure, etc.).
344 papers · 1 benchmark
The DukeMTMC-reID (Duke Multi-Tracking Multi-Camera ReIDentification) dataset is a subset of the DukeMTMC for image-based person re-ID.
344 papers · 7 benchmarks
LLFF (Local Light Field Fusion)
Local Light Field Fusion (LLFF) is a practical and robust deep learning solution for capturing and rendering novel views of complex real-world scenes for virtual exploration.
340 papers · 4 benchmarks
MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
339 papers · 2 benchmarks
Science Question Answering (ScienceQA) is a new benchmark that consists of 21,208 multimodal multiple choice questions with diverse science topics and annotations of their answers with corresponding lectures and explanations.
339 papers · 1 benchmark
The Common Objects in COntext-stuff (COCO-stuff) dataset is a dataset for scene understanding tasks like semantic segmentation, object detection and image captioning.
338 papers · 17 benchmarks
The SST-5, also known as the Stanford Sentiment Treebank with 5 labels, is a dataset used for sentiment analysis.
338 papers · 2 benchmarks
ScanObjectNN is a newly published real-world dataset comprising of 2902 3D objects in 15 categories.
337 papers · 6 benchmarks
RCV1 (Reuters Corpus Volume 1)
The RCV1 dataset is a benchmark dataset on text categorization.
336 papers · 6 benchmarks
The fastMRI dataset includes two types of MRI scans: knee MRIs and the brain (neuro) MRIs, and containing training, validation, and masked test sets.
332 papers · 5 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.