Home › Datasets › modality › Images
Images datasets
archive 2025-07-28
3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 7 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Images datasets 289–336 of 3,239
The Image Shadow Triplets dataset (ISTD) is a dataset for shadow understanding that contains 1870 image triplets of shadow image, shadow mask, and shadow-free image.
99 papers · 2 benchmarks
The LUNA16 (LUng Nodule Analysis) dataset is a dataset for lung segmentation.
99 papers · 0 benchmarks
The SYSU-MM01 is a dataset collected for the Visible-Infrared Re-identification problem.
99 papers · 2 benchmarks
IDD (Indian Driving Dataset)
IDD is a dataset for road scene understanding in unstructured environments used for semantic segmentation and object detection for autonomous driving.
98 papers · 1 benchmark
Contains 145k captions for 28k images.
98 papers · 1 benchmark
Kuzushiji-MNIST is a drop-in replacement for the MNIST dataset (28x28 grayscale, 70,000 images).
97 papers · 2 benchmarks
The Medical Segmentation Decathlon is a collection of medical image segmentation datasets.
97 papers · 1 benchmark
The ImageCLEF-DA dataset is a benchmark dataset for ImageCLEF 2014 domain adaptation challenge, which contains three domains: Caltech-256 (C), ImageNet ILSVRC 2012 (I) and Pascal VOC 2012 (P).
96 papers · 1 benchmark
Indian Pines is a Hyperspectral image segmentation dataset.
96 papers · 1 benchmark
RAVEN consists of 1,120,000 images and 70,000 RPM (Raven's Progressive Matrices) problems, equally distributed in 7 distinct figure configurations.
96 papers · 0 benchmarks
WikiArt contains painting from 195 different artists.
96 papers · 2 benchmarks
Consists of 8,422 blurry and sharp image pairs with 65,784 densely annotated FG human bounding boxes.
95 papers · 4 benchmarks
As far as we know, there only exists one large camouflaged object testing dataset, the COD10K, while the sizes of other testing datasets are less than 300.
94 papers · 1 benchmark
T-LESS is a dataset for estimating the 6D pose, i.e.
94 papers · 2 benchmarks
The VGG Face dataset is face identity recognition dataset that consists of 2,622 identities.
94 papers · 0 benchmarks
Aachen Day-Night is a dataset designed for benchmarking 6DOF outdoor visual localization in changing conditions.
93 papers · 1 benchmark
The CUHK-PEDES dataset is a caption-annotated pedestrian dataset.
93 papers · 3 benchmarks
JAFFE (Japanese Female Facial Expression)
The JAFFE dataset consists of 213 images of different facial expressions from 10 different Japanese female subjects.
93 papers · 4 benchmarks
SUN360 (Scene UNderstanding 360° panorama)
The goal of the SUN360 panorama database is to provide academic researchers in computer vision, computer graphics and computational photography, cognition and neuroscience, human perception, machine learning and data mining, with a…
93 papers · 1 benchmark
VOC 2012 (The PASCAL Visual Object Classes Challenge 2012)
see detailed use case on code implementation of the paper 'Tell Me Where To Look: Guided Attention Inference Networks'
93 papers · 0 benchmarks
xView is one of the largest publicly available datasets of overhead imagery.
93 papers · 1 benchmark
SIM10k is a synthetic dataset containing 10,000 images, which is rendered from the video game Grand Theft Auto V (GTA5).
92 papers · 3 benchmarks
The UCSD Anomaly Detection Dataset was acquired with a stationary camera mounted at an elevation, overlooking pedestrian walkways.
92 papers · 4 benchmarks
VITON (VITON-Zalando Dataset)
VITON was a dataset for virtual try-on of clothing items.
92 papers · 1 benchmark
A large-scale multi-object tracking dataset for human tracking in occlusion, frequent crossover, uniform appearance and diverse body gestures.
91 papers · 1 benchmark
FaceWarehouse is a 3D facial expression database that provides the facial geometry of 150 subjects, covering a wide range of ages and ethnic backgrounds.
91 papers · 0 benchmarks
The MIT-States dataset has 245 object classes, 115 attribute classes and ∼53K images.
91 papers · 4 benchmarks
The Oxford-IIIT Pet Dataset has 37 categories with roughly 200 images for each class.
90 papers · 5 benchmarks
ST-VQA (Scene Text Visual Question Answering)
ST-VQA aims to highlight the importance of exploiting high-level semantic information present in images as textual cues in the VQA process.
90 papers · 0 benchmarks
Winoground is a dataset for evaluating the ability of vision and language models to conduct visio-linguistic compositional reasoning.
90 papers · 1 benchmark
The COCO-Text dataset is a dataset for text detection and recognition.
89 papers · 2 benchmarks
ImageNet-O consists of images from classes that are not found in the ImageNet-1k dataset.
89 papers · 0 benchmarks
The MSU-MFSD dataset contains 280 video recordings of genuine and attack faces.
89 papers · 1 benchmark
PathVQA consists of 32,799 open-ended questions from 4,998 pathology images where each question is manually checked to ensure correctness.
89 papers · 0 benchmarks
ONCE (One Million Scenes)
ONCE (One millioN sCenEs) is a dataset for 3D object detection in the autonomous driving scenario.
87 papers · 1 benchmark
PPMI (Parkinson’s Progression Markers Initiative)
The Parkinson’s Progression Markers Initiative (PPMI) dataset originates from an observational clinical and longitudinal study comprising evaluations of people with Parkinson’s disease (PD), those people with high risk, and those who are…
87 papers · 3 benchmarks
AbstractReasoning is a dataset for abstract reasoning, where the goal is to infer the correct answer from the context panels based on abstract reasoning.
86 papers · 0 benchmarks
DICM is a dataset for low-light enhancement which consists of 69 images collected with commercial digital cameras.
86 papers · 1 benchmark
DIODE (Dense Indoor and Outdoor Depth)
Diode Dense Indoor/Outdoor DEpth (DIODE) is the first standard dataset for monocular depth estimation comprising diverse indoor and outdoor scenes acquired with the same hardware setup.
86 papers · 2 benchmarks
BigEarthNet consists of 590,326 Sentinel-2 image patches, each of which is a section of i) 120x120 pixels for 10m bands; ii) 60x60 pixels for 20m bands; and iii) 20x20 pixels for 60m bands.
85 papers · 3 benchmarks
CULane is a large scale challenging dataset for academic research on traffic lane detection.
85 papers · 1 benchmark
Structured3D is a large-scale photo-realistic dataset containing 3.5K house designs (a) created by professional designers with a variety of ground truth 3D structure annotations (b) and generate photo-realistic 2D images (c).
85 papers · 7 benchmarks
The PROMISE12 dataset was made available for the MICCAI 2012 prostate segmentation challenge.
84 papers · 2 benchmarks
VITON-HD (High-Resolution VITON-Zalando Dataset)
VITON-HD dataset is a dataset for high-resolution (i.e., 1024x768) virtual try-on of clothing items.
84 papers · 2 benchmarks
NLVR (Natural Language Visual Reasoningnatural language for visual reasoning)
NLVR contains 92,244 pairs of human-written English sentences grounded in synthetic images.
83 papers · 3 benchmarks
ChestX-ray8 is a medical imaging dataset which comprises 108,948 frontal-view X-ray images of 32,717 (collected from the year of 1992 to 2015) unique patients with the text-mined eight common disease labels, mined from the text…
81 papers · 0 benchmarks
LoveDA (Remote Sensing Land-Cover Dataset for Domain Adaptive Semantic Segmentation)
1.
81 papers · 1 benchmark
RaFD (Radboud Faces Database)
The Radboud Faces Database (RaFD) is a set of pictures of 67 models (both adult and children, males and females) displaying 8 emotional expressions.
81 papers · 2 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.