12,172 datasets listed, ordered by the archive's paper count. Page 3 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
The CheXpert dataset contains 224,316 chest radiographs of 65,240 patients with both frontal and lateral views available.
628 papers · 3 benchmarks
TruthfulQA is a benchmark to measure whether a language model is truthful in generating answers to questions.
607 papers · 3 benchmarks
The iNaturalist 2017 dataset (iNat) contains 675,170 training and validation images from 5,089 natural fine-grained categories.
603 papers · 12 benchmarks
The Cora dataset consists of 2708 scientific publications classified into one of seven classes.
602 papers · 18 benchmarks
ImageNet-C is an open source data set that consists of algorithmically generated corruptions (blur, noise) applied to the ImageNet test-set.
602 papers · 4 benchmarks
DBpedia (from "DB" for "database") is a project aiming to extract structured content from the information created in the Wikipedia project.
597 papers · 4 benchmarks
The Urban100 dataset contains 100 images of urban scenes.
591 papers · 25 benchmarks
VoxCeleb2 is a large scale speaker recognition dataset obtained automatically from open-source media.
564 papers · 5 benchmarks
The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified Good and Featured articles on Wikipedia.
560 papers · 2 benchmarks
LVIS is a dataset for long tail instance segmentation.
551 papers · 12 benchmarks
VGGFace2 (Vggface2: A dataset for recognising faces across pose and age)
VGGFace2 is a large-scale face recognition dataset.
539 papers · 3 benchmarks
D4RL is a collection of environments for offline reinforcement learning.
538 papers · 2 benchmarks
SYNTHIA (SYNTHetic Collection of Imagery and Annotations)
The SYNTHIA dataset is a synthetic dataset that consists of 9400 multi-viewpoint photo-realistic frames rendered from a virtual city and comes with pixel-level semantic annotations for 13 classes.
538 papers · 10 benchmarks
VGG-Face2 (Vggface2: A dataset for recognising faces across pose and age)
VGGFace2 is a large-scale face recognition dataset.
533 papers · 0 benchmarks
The Set14 dataset is a dataset consisting of 14 images commonly used for testing performance of Image Super-Resolution models.
532 papers · 8 benchmarks
This dataset contains 3.3K expert-level pairwise human preferences for model responses generated by 6 models in response to 80 MT-bench questions.
531 papers · 1 benchmark
CNN/Daily Mail is a dataset for text summarization.
530 papers · 8 benchmarks
The Places205 dataset is a large-scale scene-centric dataset with 205 common scene categories.
525 papers · 1 benchmark
The DreamBooth dataset is a collection of images used for fine-tuning text-to-image diffusion models for subject-driven generation¹.
523 papers · 1 benchmark
FGVC-Aircraft contains 10,200 images of aircraft, with 100 images for each of 102 different aircraft model variants, most of which are airplanes.
520 papers · 12 benchmarks
The Universal Dependencies (UD) project seeks to develop cross-linguistically consistent treebank annotation of morphology and syntax for multiple languages.
520 papers · 5 benchmarks
FEVER (Fact Extraction and VERification)
FEVER is a publicly available dataset for fact extraction and verification against textual sources.
498 papers · 3 benchmarks
The MPII Human Pose Dataset for single person pose estimation is composed of about 25K images of which 15K are training samples, 3K are validation samples and 7K are testing samples (which labels are withheld by the authors).
495 papers · 4 benchmarks
Common corruptions dataset for CIFAR10
494 papers · 2 benchmarks
S3DIS (Stanford 3D Indoor Scene Dataset (S3DIS))
The Stanford 3D Indoor Scene Dataset (S3DIS) dataset contains 6 large-scale indoor areas with 271 rooms.
488 papers · 9 benchmarks
The WN18 dataset has 18 relations scraped from WordNet for roughly 41,000 synsets, resulting in 141,442 triplets.
485 papers · 3 benchmarks
The CommonsenseQA is a dataset for commonsense question answering task.
483 papers · 1 benchmark
ImageNet-R(endition) contains art, cartoons, deviantart, graffiti, embroidery, graphics, origami, paintings, patterns, plastic objects, plush objects, sculptures, sketches, tattoos, toys, and video game renditions of ImageNet classes.
481 papers · 5 benchmarks
The Waymo Open Dataset is comprised of high resolution sensor data collected by autonomous vehicles operated by the Waymo Driver in a wide variety of conditions.
481 papers · 16 benchmarks
The SUN RGBD dataset contains 10335 real RGB-D images of room scenes.
477 papers · 11 benchmarks
NTU RGB+D is a large-scale dataset for RGB-D human action recognition.
476 papers · 9 benchmarks
TextVQA is a dataset to benchmark visual reasoning based on text in images.
476 papers · 3 benchmarks
This CSTR VCTK Corpus includes speech data uttered by 110 English speakers with various accents.
476 papers · 6 benchmarks
Datasets drive vision progress, yet existing driving datasets are impoverished in terms of visual content and supported tasks to study multitask learning for autonomous driving.
469 papers · 16 benchmarks
The Pile is a 825 GiB diverse, open source language modelling data set that consists of 22 smaller, high-quality datasets combined together.
467 papers · 1 benchmark
The Matterport3D dataset is a large RGB-D dataset for scene understanding in indoor environments.
461 papers · 4 benchmarks
USPS is a digit dataset automatically scanned from envelopes by the U.S.
459 papers · 2 benchmarks
FB15k-237 is a link prediction dataset created from FB15k.
452 papers · 3 benchmarks
Perceptual Similarity is a dataset of human perceptual similarity judgments.
452 papers · 0 benchmarks
Common Voice is an audio dataset that consists of a unique MP3 and corresponding text file.
449 papers · 143 benchmarks
FrameNet is a linguistic knowledge graph containing information about lexical and predicate argument semantics of the English language.
444 papers · 0 benchmarks
The Set5 dataset is a dataset consisting of 5 images (“baby”, “bird”, “butterfly”, “head”, “woman”) commonly used for testing performance of Image Super-Resolution models.
444 papers · 9 benchmarks
The RefCOCO dataset is a referring expression generation (REG) dataset used for tasks related to understanding natural language expressions that refer to specific objects in images.
439 papers · 11 benchmarks
The ImageNet-A dataset consists of real-world, unmodified, and naturally occurring examples that are misclassified by ResNet models.
431 papers · 5 benchmarks
The Charades dataset is composed of 9,848 videos of daily indoors activities with an average length of 30 seconds, involving interactions with 46 objects classes in 15 types of indoor scenes and containing a vocabulary of 30 verbs leading…
428 papers · 6 benchmarks
WebText is an internal OpenAI corpus created by scraping web pages with emphasis on document quality.
425 papers · 0 benchmarks
SuperGLUE is a benchmark dataset designed to pose a more rigorous test of language understanding than GLUE.
423 papers · 0 benchmarks
CUHK03 (Chinese University of Hong Kong Re-identification)
The CUHK03 consists of 14,097 images of 1,467 different identities, where 6 campus cameras were deployed for image collection and each identity is captured by 2 campus cameras.
419 papers · 8 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.