Home › Datasets › modality › Images
Images datasets
archive 2025-07-28
3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 3 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Images datasets 97–144 of 3,239
MPI-INF-3DHP is a 3D human body pose estimation dataset consisting of both constrained indoor and complex outdoor scenes.
289 papers · 5 benchmarks
Clothing1M contains 1M clothing images in 14 classes.
288 papers · 4 benchmarks
DUTS is a saliency detection dataset containing 10,553 training images and 5,019 test images.
286 papers · 5 benchmarks
The Shanghaitech dataset is a large-scale crowd counting dataset.
277 papers · 5 benchmarks
PASCAL-S is a dataset for salient object detection consisting of a set of 850 images from PASCAL VOC 2010 validation set with multiple salient objects on the scenes.
271 papers · 3 benchmarks
ImageNet-Sketch data set consists of 50,889 images, approximately 50 images for each of the 1000 ImageNet classes.
268 papers · 3 benchmarks
DensePose-COCO is a large-scale ground-truth dataset with image-to-surface correspondences manually annotated on 50K COCO images and train DensePose-RCNN, to densely regress part-specific UV coordinates within every human region at…
265 papers · 1 benchmark
AwA (Animals with Attributes)
Animals with Attributes (AwA) was a dataset for benchmarking transfer-learning algorithms, in particular attribute base classification.
264 papers · 3 benchmarks
EMNIST (extended MNIST) has 4 times more data than MNIST.
264 papers · 10 benchmarks
BSDS500 (Berkeley Segmentation Dataset 500)
Berkeley Segmentation Data Set 500 (BSDS500) is a standard benchmark for contour detection.
261 papers · 8 benchmarks
NAS-Bench-201 is a benchmark (and search space) for neural architecture search.
260 papers · 4 benchmarks
The VizWiz-VQA dataset originates from a natural visual question answering setting where blind people each took an image and recorded a spoken question about it, together with 10 crowdsourced answers per visual question.
260 papers · 7 benchmarks
The LOL dataset is composed of 500 low-light and normal-light image pairs and divided into 485 training pairs and 15 testing pairs.
257 papers · 2 benchmarks
The MS-Celeb-1M dataset is a large-scale face recognition dataset consists of 100K identities, and each identity has about 100 facial images.
257 papers · 0 benchmarks
Foggy Cityscapes is a synthetic foggy dataset which simulates fog on real scenes.
249 papers · 7 benchmarks
The HPatches is a recent dataset for local patch descriptor evaluation that consists of 116 sequences of 6 images with known homography.
248 papers · 4 benchmarks
The ICDAR 2013 dataset consists of 229 training images and 233 testing images, with word-level annotations provided.
246 papers · 3 benchmarks
IJB-C (IARPA Janus Benchmark-C)
The IJB-C dataset is a video-based face recognition dataset.
246 papers · 3 benchmarks
KITTI-360 is a large-scale dataset that contains rich sensory information and full annotations.
246 papers · 7 benchmarks
SIDD (Smartphone Image Denoising Dataset)
SIDD is an image denoising dataset containing 30,000 noisy images from 10 scenes under different lighting conditions using five representative smartphone cameras.
245 papers · 2 benchmarks
The UTKFace dataset is a large-scale face dataset with long age span (range from 0 to 116 years old).
243 papers · 4 benchmarks
YFCC100M is a that dataset contains a total of 100 million media objects, of which approximately 99.2 million are photos and 0.8 million are videos, all of which carry a Creative Commons license.
243 papers · 0 benchmarks
MathVista (Mathematical Reasoning of in Visual Contexts)
MathVista is a consolidated Mathematical reasoning benchmark within Visual contexts.
242 papers · 0 benchmarks
The LIDC-IDRI dataset contains lesion annotations from four experienced thoracic radiologists.
240 papers · 6 benchmarks
MIMIC-CXR from Massachusetts Institute of Technology presents 371,920 chest X-rays associated with 227,943 imaging studies from 65,079 patients.
240 papers · 3 benchmarks
GOT-10k (Generic Object Tracking Benchmark)
The GOT-10k dataset contains more than 10,000 video segments of real-world moving objects and over 1.5 million manually labelled bounding boxes.
239 papers · 2 benchmarks
ChestX-ray14 is a medical imaging dataset which comprises 112,120 frontal-view X-ray images of 30,805 (collected from the year of 1992 to 2015) unique patients with the text-mined fourteen common disease labels, mined from the text…
237 papers · 6 benchmarks
The Pascal3D+ multi-view dataset consists of images in the wild, i.e., images of object categories exhibiting high variability, captured under uncontrolled settings, in cluttered scenes and under many different poses.
237 papers · 1 benchmark
The Sketch dataset contains over 20,000 sketches evenly distributed over 250 object categories.
237 papers · 1 benchmark
TUM RGB-D is an RGB-D dataset.
235 papers · 1 benchmark
AwA2 (Animals with Attributes 2)
Animals with Attributes 2 (AwA2) is a dataset for benchmarking transfer-learning algorithms, such as attribute base classification and zero-shot learning.
231 papers · 5 benchmarks
DAVIS16 is a dataset for video object segmentation which consists of 50 videos in total (30 videos for training and 20 for testing).
231 papers · 4 benchmarks
Stanford Online Products (SOP) dataset has 22,634 classes with 120,053 product images.
231 papers · 5 benchmarks
CamVid (Cambridge-driving Labeled Video Database)
CamVid (Cambridge-driving Labeled Video Database) is a road/driving scene understanding database which was originally captured as five video sequences with a 960×720 resolution camera mounted on the dashboard of a car.
227 papers · 4 benchmarks
FlyingThings3D is a synthetic dataset for optical flow, disparity and scene flow estimation.
226 papers · 0 benchmarks
HKU-IS is a visual saliency prediction dataset which contains 4447 challenging images, most of which have either low contrast or multiple salient objects.
226 papers · 3 benchmarks
The Middlebury Stereo dataset consists of high-resolution stereo sequences with complex geometry and pixel-accurate ground-truth disparity data.
223 papers · 5 benchmarks
VisDA-2017 is a simulation-to-real dataset for domain adaptation with over 280,000 images across 12 categories in the training, validation and testing domains.
223 papers · 6 benchmarks
The Vimeo-90K is a large-scale high-quality video dataset for lower-level video processing.
220 papers · 3 benchmarks
The DUT-OMRON dataset is used for evaluation of Salient Object Detection task and it contains 5,168 high quality images.
214 papers · 4 benchmarks
TrackingNet is a large-scale tracking dataset consisting of videos in the wild.
210 papers · 2 benchmarks
HAM10000 is a dataset of 10000 training images for detecting pigmented skin lesions.
209 papers · 2 benchmarks
FairFace is a face image dataset which is race balanced.
208 papers · 1 benchmark
AI2 Diagrams (AI2D) is a dataset of over 5000 grade school science diagrams with over 150000 rich annotations, their ground truth syntactic parses, and more than 15000 corresponding multiple choice questions.
207 papers · 1 benchmark
300W (300 Faces-In-The-Wild)
The 300-W is a face dataset that consists of 300 Indoor and 300 Outdoor in-the-wild images.
206 papers · 9 benchmarks
CIFAR100 few-shots (CIFAR-FS) is randomly sampled from CIFAR-100 (Krizhevsky & Hinton, 2009) by using the same criteria with which miniImageNet has been generated.
206 papers · 2 benchmarks
LAION 5B is a large-scale dataset for research purposes consisting of 5,85B CLIP-filtered image-text pairs.
205 papers · 0 benchmarks
COCO Captions contains over one and a half million captions describing over 330,000 images.
203 papers · 4 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.