Home › Datasets › modality › Images
Images datasets
archive 2025-07-28
3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 2 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Images datasets 49–96 of 3,239
SYNTHIA (SYNTHetic Collection of Imagery and Annotations)
The SYNTHIA dataset is a synthetic dataset that consists of 9400 multi-viewpoint photo-realistic frames rendered from a virtual city and comes with pixel-level semantic annotations for 13 classes.
538 papers · 10 benchmarks
VGG-Face2 (Vggface2: A dataset for recognising faces across pose and age)
VGGFace2 is a large-scale face recognition dataset.
533 papers · 0 benchmarks
The Set14 dataset is a dataset consisting of 14 images commonly used for testing performance of Image Super-Resolution models.
532 papers · 8 benchmarks
The Places205 dataset is a large-scale scene-centric dataset with 205 common scene categories.
525 papers · 1 benchmark
FGVC-Aircraft contains 10,200 images of aircraft, with 100 images for each of 102 different aircraft model variants, most of which are airplanes.
520 papers · 12 benchmarks
The MPII Human Pose Dataset for single person pose estimation is composed of about 25K images of which 15K are training samples, 3K are validation samples and 7K are testing samples (which labels are withheld by the authors).
495 papers · 4 benchmarks
Common corruptions dataset for CIFAR10
494 papers · 2 benchmarks
ImageNet-R(endition) contains art, cartoons, deviantart, graffiti, embroidery, graphics, origami, paintings, patterns, plastic objects, plush objects, sculptures, sketches, tattoos, toys, and video game renditions of ImageNet classes.
481 papers · 5 benchmarks
The Waymo Open Dataset is comprised of high resolution sensor data collected by autonomous vehicles operated by the Waymo Driver in a wide variety of conditions.
481 papers · 16 benchmarks
The SUN RGBD dataset contains 10335 real RGB-D images of room scenes.
477 papers · 11 benchmarks
TextVQA is a dataset to benchmark visual reasoning based on text in images.
476 papers · 3 benchmarks
USPS is a digit dataset automatically scanned from envelopes by the U.S.
459 papers · 2 benchmarks
Perceptual Similarity is a dataset of human perceptual similarity judgments.
452 papers · 0 benchmarks
The Set5 dataset is a dataset consisting of 5 images (“baby”, “bird”, “butterfly”, “head”, “woman”) commonly used for testing performance of Image Super-Resolution models.
444 papers · 9 benchmarks
The RefCOCO dataset is a referring expression generation (REG) dataset used for tasks related to understanding natural language expressions that refer to specific objects in images.
439 papers · 11 benchmarks
The ImageNet-A dataset consists of real-world, unmodified, and naturally occurring examples that are misclassified by ResNet models.
431 papers · 5 benchmarks
CUHK03 (Chinese University of Hong Kong Re-identification)
The CUHK03 consists of 14,097 images of 1,467 different identities, where 6 campus cameras were deployed for image collection and each identity is captured by 2 campus cameras.
419 papers · 8 benchmarks
The CASIA-WebFace dataset is used for face verification and face identification tasks.
415 papers · 2 benchmarks
The Replica Dataset is a dataset of high quality reconstructions of a variety of indoor spaces.
414 papers · 4 benchmarks
GTA5 (Grand Theft Auto 5)
The GTA5 dataset contains 24966 synthetic images with pixel level semantic annotation.
412 papers · 7 benchmarks
MVTecAD (MVTEC ANOMALY DETECTION DATASET)
MVTec AD is a dataset for benchmarking anomaly detection methods with a focus on industrial inspection.
402 papers · 4 benchmarks
Caltech-256 is an object recognition dataset containing 30,607 real-world images, of different sizes, spanning 257 classes (256 object classes and an additional clutter class).
401 papers · 4 benchmarks
DeepFashion is a dataset containing around 800K diverse fashion images with their rich annotations (46 categories, 1,000 descriptive attributes, bounding boxes and landmark information) ranging from well-posed product images to…
397 papers · 5 benchmarks
The 3D Poses in the Wild dataset is the first dataset in the wild with accurate 3D poses for evaluation.
395 papers · 4 benchmarks
The GoPro dataset for deblurring consists of 3,214 blurred images with the size of 1,280×720 that are divided into 2,103 training images and 1,111 test images.
390 papers · 4 benchmarks
Argoverse is a tracking benchmark with over 30K scenarios collected in Pittsburgh and Miami.
386 papers · 6 benchmarks
GTSRB (German Traffic Sign Recognition Benchmark)
The German Traffic Sign Recognition Benchmark (GTSRB) contains 43 classes of traffic signs, split into 39,209 training images and 12,630 test images.
374 papers · 5 benchmarks
FaceForensics++ is a forensics dataset consisting of 1000 original video sequences that have been manipulated with four automated face manipulation methods: Deepfakes, Face2Face, FaceSwap and NeuralTextures.
368 papers · 2 benchmarks
OK-VQA (Outside Knowledge Visual Question Answering)
Outside Knowledge Visual Question Answering (OK-VQA) includes more than 14,000 questions that require external knowledge to answer.
368 papers · 2 benchmarks
Visual Question Answering (VQA) v2.0 is a dataset containing open-ended questions about images.
366 papers · 7 benchmarks
Automatic image captioning is the task of producing a natural-language utterance (usually a sentence) that correctly reflects the visual content of an image.
352 papers · 2 benchmarks
The NUS-WIDE dataset contains 269,648 images with a total of 5,018 tags collected from Flickr.
348 papers · 3 benchmarks
The DukeMTMC-reID (Duke Multi-Tracking Multi-Camera ReIDentification) dataset is a subset of the DukeMTMC for image-based person re-ID.
344 papers · 7 benchmarks
MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
339 papers · 2 benchmarks
Science Question Answering (ScienceQA) is a new benchmark that consists of 21,208 multimodal multiple choice questions with diverse science topics and annotations of their answers with corresponding lectures and explanations.
339 papers · 1 benchmark
The Common Objects in COntext-stuff (COCO-stuff) dataset is a dataset for scene understanding tasks like semantic segmentation, object detection and image captioning.
338 papers · 17 benchmarks
The fastMRI dataset includes two types of MRI scans: knee MRIs and the brain (neuro) MRIs, and containing training, validation, and masked test sets.
332 papers · 5 benchmarks
Animal FacesHQ (AFHQ) is a dataset of animal faces consisting of 15,000 high-quality images at 512 × 512 resolution.
327 papers · 6 benchmarks
AffectNet is a large facial expression dataset with around 0.4 million images manually labeled for the presence of eight (neutral, happy, angry, sad, fear, surprise, disgust, contempt) facial expressions along with the intensity of valence…
323 papers · 4 benchmarks
The PASCAL Context dataset is an extension of the PASCAL VOC 2010 detection challenge, and it contains pixel-wise labels for all training images.
323 papers · 6 benchmarks
The tieredImageNet dataset is a larger subset of ILSVRC-12 with 608 classes (779,165 images) grouped into 34 higher-level nodes in the ImageNet human-curated hierarchy.
317 papers · 7 benchmarks
DRIVE (Digital Retinal Images for Vessel Extraction)
The Digital Retinal Images for Vessel Extraction (DRIVE) dataset is a dataset for retinal vessel segmentation.
311 papers · 2 benchmarks
DAVIS17 is a dataset for video object segmentation.
308 papers · 12 benchmarks
Manga109 has been compiled by the Aizawa Yamasaki Matsui Laboratory, Department of Information and Communication Engineering, the Graduate School of Information Science and Technology, the University of Tokyo.
300 papers · 12 benchmarks
The Multi-PIE (Multi Pose, Illumination, Expressions) dataset consists of face images of 337 subjects taken under different pose, illumination and expressions.
299 papers · 1 benchmark
DOTA (Dataset for Object deTection in Aerial Images)
DOTA is a large-scale dataset for object detection in aerial images.
293 papers · 2 benchmarks
DocVQA consists of 50,000 questions defined on 12,000+ document images.
290 papers · 3 benchmarks
The 20BN-SOMETHING-SOMETHING V2 dataset is a large collection of labeled video clips that show humans performing pre-defined basic actions with everyday objects.
290 papers · 7 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.