Home › Datasets › modality › Images

Images datasets

archive 2025-07-28

3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 5 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Images datasets 193–240 of 3,239

In-Shop (In-shop Clothes Retrieval Benchmark)
In-shop Clothes Retrieval Benchmark evaluates the performance of in-shop Clothes Retrieval.
154 papers · 2 benchmarks
CC12M (Conceptual 12M)
Conceptual 12M (CC12M) is a dataset with 12 million image-text pairs specifically meant to be used for vision-and-language pre-training.
153 papers · 0 benchmarks
dSprites (Disentanglement testing Sprites dataset)
dSprites is a dataset of 2D shapes procedurally generated from 6 ground truth independent latent factors.
153 papers · 0 benchmarks
The MegaDepth dataset is a dataset for single-view depth prediction that includes 196 different locations reconstructed from COLMAP SfM/MVS.
152 papers · 0 benchmarks
VRD (Visual Relationship Detection dataset)
The Visual Relationship Dataset (VRD) contains 4000 images for training and 1000 for testing annotated with visual relationships.
151 papers · 5 benchmarks
MOT16 (Multiple Object Tracking 2016)
The MOT16 dataset is a dataset for multiple object tracking.
149 papers · 2 benchmarks
DISFA (Denver Intensity of Spontaneous Facial Action)
The Denver Intensity of Spontaneous Facial Action (DISFA) dataset consists of 27 videos of 4844 frames each, with 130,788 images in total.
148 papers · 3 benchmarks
2D-3D-S (2D-3D-Semantic)
The 2D-3D-S dataset provides a variety of mutually registered modalities from 2D, 2.5D and 3D domains, with instance-level semantic and geometric annotations.
147 papers · 6 benchmarks
SBD (Semantic Boundaries Dataset)
The Semantic Boundaries Dataset (SBD) is a dataset for predicting pixels on the boundary of the object (as opposed to the inside of the object with semantic segmentation).
147 papers · 2 benchmarks
Taskonomy provides a large and high-quality dataset of varied indoor scenes.
147 papers · 2 benchmarks
aPY (Attribute Pascal and Yahoo)
aPY is a coarse-grained dataset composed of 15339 images from 3 broad categories (animals, objects and vehicles), further divided into a total of 32 subcategories (aeroplane, …, zebra).
147 papers · 5 benchmarks
STARE (Structured Analysis of the Retina)
The STARE (Structured Analysis of the Retina) dataset is a dataset for retinal vessel segmentation.
146 papers · 6 benchmarks
The SumMe dataset is a video summarization dataset consisting of 25 videos, each annotated with at least 15 human summaries (390 in total).
146 papers · 3 benchmarks
Set12 is a collection of 12 grayscale images of different scenes that are widely used for evaluation of image denoising methods.
145 papers · 5 benchmarks
VQA-RAD (Visual Question Answering in Radiology)
VQA-RAD consists of 3,515 question–answer pairs on 315 radiology images.
145 papers · 0 benchmarks
fMoW (Functional Map of the World)
Functional Map of the World (fMoW) is a dataset that aims to inspire the development of machine learning models capable of predicting the functional purpose of buildings and land use from temporal sequences of satellite images and a rich…
144 papers · 1 benchmark
MPII Human Pose Dataset is a dataset for human pose estimation.
143 papers · 1 benchmark
NABirds (North America Birds)
NABirds V1 is a collection of 48,000 annotated photographs of the 400 species of birds that are commonly observed in North America.
143 papers · 1 benchmark
Aff-Wild2 is a large-scale in-the-wild database and an extension of the Aff-Wild dataset for affect recognition.
142 papers · 2 benchmarks
CBSD68 (Color BSD68)
Color BSD68 dataset for image denoising benchmarks is part of The Berkeley Segmentation Dataset and Benchmark.
142 papers · 15 benchmarks
The Flickr30K Entities dataset is an extension to the Flickr30K dataset.
142 papers · 2 benchmarks
The Pix3D dataset is a large-scale benchmark of diverse image-shape pairs with pixel-level 2D-3D alignment.
142 papers · 5 benchmarks
CAMO (Camouflaged Object)
Camouflaged Object (CAMO) dataset specifically designed for the task of camouflaged object segmentation.
139 papers · 2 benchmarks
FC100 (Fewshot-CIFAR100)
The FC100 dataset (Fewshot-CIFAR100) is a newly split dataset based on CIFAR-100 for few-shot learning.
137 papers · 5 benchmarks
Oxford5k (Oxford Buildings)
Oxford5K is the Oxford Buildings Dataset, which contains 5062 images collected from Flickr.
137 papers · 1 benchmark
SEED-Bench consists of 19K multiple choice questions with accurate human annotations (~6 larger than existing benchmarks), which spans 12 evaluation dimensions including the comprehension of both the image and video modality.
137 papers · 0 benchmarks
Cholec80 (Surgical Workflow Dataset)
Cholec80 is an endoscopic video dataset containing 80 videos of cholecystectomy surgeries performed by 13 surgeons.
134 papers · 2 benchmarks
VIPeR (Viewpoint Invariant Pedestrian Recognition)
The Viewpoint Invariant Pedestrian Recognition (VIPeR) dataset includes 632 people and two outdoor cameras under different viewpoints and light conditions.
134 papers · 0 benchmarks
VehicleID (PKU VehicleID)
The “VehicleID” dataset contains CARS captured during the daytime by multiple real-world surveillance cameras distributed in a small city in China.
134 papers · 8 benchmarks
The Replay-Attack Database for face spoofing consists of 1300 video clips of photo and video attack attempts to 50 clients, under different lighting conditions.
133 papers · 1 benchmark
RobustBench is a benchmark of adversarial robustness, which as accurately as possible reflects the robustness of the considered models within a reasonable computational budget.
133 papers · 0 benchmarks
Virtual KITTI is a photo-realistic synthetic video dataset designed to learn and evaluate computer vision models for several video understanding tasks: object detection and multi-object tracking, scene-level and instance-level semantic…
133 papers · 0 benchmarks
CORe50 is a dataset designed for assessing Continual Learning techniques in an Object Recognition context.
132 papers · 0 benchmarks
Panoptic (CMU Panoptic Studio)
CMU Panoptic is a large scale dataset providing 3D pose annotations (1.5 millions) for multiple people engaging social activities.
131 papers · 4 benchmarks
The CityPersons dataset is a subset of Cityscapes which only consists of person annotations.
129 papers · 2 benchmarks
The Make3D dataset is a monocular Depth Estimation dataset that contains 400 single training RGB and depth map pairs, and 134 test samples.
129 papers · 1 benchmark
ORL (Our Database of Faces)
The ORL Database of Faces contains 400 images from 40 distinct subjects.
129 papers · 1 benchmark
VOT2018 is a dataset for visual object tracking.
129 papers · 1 benchmark
LFPW (Labeled Face Parts in the Wild)
The Labeled Face Parts in-the-Wild (LFPW) consists of 1,432 faces from images downloaded from the web using simple text queries on sites such as google.com, flickr.com, and yahoo.com.
128 papers · 0 benchmarks
The Meta-Dataset benchmark is a large few-shot learning benchmark and consists of multiple datasets of different data distributions.
128 papers · 2 benchmarks
FBMS (Freiburg-Berkeley Motion Segmentation)
The Freiburg-Berkeley Motion Segmentation Dataset (FBMS-59) is an extension of the BMS dataset with 33 additional video sequences.
126 papers · 1 benchmark
SUN3D contains a large-scale RGB-D video database, with 8 annotated sequences.
126 papers · 0 benchmarks
Aff-Wild is a large-scale in-the-wild dataset for valence-arousal estimation from videos with a variety of head poses, illumination conditions and occlusions.
125 papers · 0 benchmarks
FreiHAND is a 3D hand pose dataset which records different hand actions performed by 32 people.
125 papers · 1 benchmark
The LUNA challenges provide datasets for automatic nodule detection algorithms using the largest publicly available reference database of chest CT scans, the LIDC-IDRI data set.
125 papers · 2 benchmarks
FER+ (Face Expression Recognition Plus dataset)
The FER+ dataset is an extension of the original FER dataset, where the images have been re-labelled into one of 8 emotion types: neutral, happiness, surprise, sadness, anger, disgust, fear, and contempt.
124 papers · 3 benchmarks
MSRA-TD500 (MSRA Text Detection 500 Database)
The MSRA-TD500 dataset is a text detection dataset that contains 300 training images and 200 test images.
124 papers · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.