Home › Datasets › modality › Images
Images datasets
archive 2025-07-28
3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 14 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Images datasets 625–672 of 3,239
MetaShift is a collection of 12,868 sets of natural images across 410 classes.
28 papers · 0 benchmarks
OASIS (Open Annotations of Single Image Surfaces)
A dataset for single-image 3D in the wild consisting of annotations of detailed 3D geometry for 140,000 images.
28 papers · 3 benchmarks
PSG dataset has 48749 images with 133 object classes (80 objects and 53 stuff) and 56 predicate classes.
28 papers · 1 benchmark
SegTHOR (Segmentation of THoracic Organs at Risk)
SegTHOR (Segmentation of THoracic Organs at Risk) is a dataset dedicated to the segmentation of organs at risk (OARs) in the thorax, i.e.
28 papers · 0 benchmarks
THuman2.0 Dataset contains 500 high-quality human scans captured by a dense DLSR rig.
28 papers · 1 benchmark
TextZoom is a super-resolution dataset that consists of paired Low Resolution – High Resolution scene text images.
28 papers · 2 benchmarks
Research in massively multilingual image captioning has been severely hampered by a lack of high-quality evaluation datasets.
28 papers · 0 benchmarks
2D-3D Match Dataset is a new dataset of 2D-3D correspondences by leveraging the availability of several 3D datasets from RGB-D scans.
27 papers · 0 benchmarks
DeeperForensics-1.0 represents the largest face forgery detection dataset by far, with 60,000 videos constituted by a total of 17.6 million frames, 10 times larger than existing datasets of the same kind.
27 papers · 0 benchmarks
GID (Gaofen Image Dataset)
Gaofen Image Dataset (GID) is a large-scale land-cover dataset constructed with Gaofen-2 (GF-2) satellite images.
27 papers · 0 benchmarks
The ISIC 2018 dataset was published by the International Skin Imaging Collaboration (ISIC) as a large-scale dataset of dermoscopy images.
27 papers · 1 benchmark
A benchmark dataset for out-of-distribution detection.
27 papers · 1 benchmark
IntrA is an open-access 3D intracranial aneurysm dataset that makes the application of points-based and mesh-based classification and segmentation models available.
27 papers · 2 benchmarks
LHQ (Landscapes High-Quality)
A dataset of 90,000 high-resolution nature landscape images, crawled from Unsplash and Flickr and preprocessed with Mask R-CNN and Inception V3.
27 papers · 4 benchmarks
MedMNIST v2 is a large-scale MNIST-like collection of standardized biomedical images, including 12 datasets for 2D and 6 datasets for 3D.
27 papers · 0 benchmarks
Multi-Modal-CelebA-HQ is a large-scale face image dataset that has 30,000 high-resolution face images selected from the CelebA dataset by following CelebA-HQ.
27 papers · 3 benchmarks
Nighttime Driving is a dataset of road scenes consisting of 35,000 images ranging from daytime to twilight time and to nighttime.
27 papers · 2 benchmarks
The goal of this benchmark is to introduce a standard evaluation metric to measure the accuracy and robustness of 3D face reconstruction methods under variations in viewing angle, lighting, and common occlusions.
27 papers · 1 benchmark
RadarScenes is a real-world radar point cloud dataset for automotive applications.
27 papers · 0 benchmarks
S2Looking is a building change detection dataset that contains large-scale side-looking satellite images captured at varying off-nadir angles.
27 papers · 1 benchmark
The Caltech 101 Silhouettes dataset consists of 4,100 training samples, 2,264 validation samples and 2,307 test samples.
27 papers · 0 benchmarks
UVO (Unidentified Video Objects: A Benchmark for Dense, Open-World Segmentation)
UVO is a new benchmark for open-world class-agnostic object segmentation in videos.
27 papers · 2 benchmarks
VQA-HAT (Human ATtention) is a dataset to evaluate the informative regions of an image depending on the question being asked about it.
27 papers · 0 benchmarks
iNat2021 is a large-scale image dataset collected and annotated by community scientists that contains over 2.7M images from 10k different species.
27 papers · 0 benchmarks
Animal Kingdom is a large and diverse dataset that provides multiple annotated tasks to enable a more thorough understanding of natural animal behaviors.
26 papers · 2 benchmarks
DUDE (Document UnderstanDing of Everything)
DUDE is formulated as an instance of Document Question Answering (DocQA) to evaluate how well current solutions deal with multi-page documents, if they can navigate and reason over the layout, and if they can generalize these skills to…
26 papers · 0 benchmarks
Contains 1024 pairs of high-quality images and covers diverse scenarios.
26 papers · 0 benchmarks
Object detection benchmark for logo detection.
26 papers · 3 benchmarks
Harm-C is a dataset for detecting harmful memes related to Covid-19.
26 papers · 0 benchmarks
MNIST8M is derived from the MNIST dataset by applying random deformations and translations to the dataset.
26 papers · 0 benchmarks
The exact pre-processing steps used to construct the MNIST dataset have long been lost.
26 papers · 2 benchmarks
The SUN09 dataset consists of 12,000 annotated images with more than 200 object categories.
26 papers · 0 benchmarks
Screen2Words is a large-scale screen summarization dataset annotated by human workers.
26 papers · 0 benchmarks
The PASCAL-Scribble Dataset is an extension of the PASCAL dataset with scribble annotations for semantic segmentation.
26 papers · 0 benchmarks
The SentiCap dataset contains several thousand images with captions with positive and negative sentiments.
26 papers · 0 benchmarks
UrbanCars facilitates multi-shortcut learning under the controlled setting with two shortcuts—background and co-occurring object.
26 papers · 0 benchmarks
VALSE (VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena)
We propose VALSE (Vision And Language Structured Evaluation), a novel benchmark designed for testing general-purpose pretrained vision and language (V&L) models for their visio-linguistic grounding capabilities on specific linguistic…
26 papers · 12 benchmarks
VAW (Visual Attributes in the Wild)
VAW is a large scale visual attributes dataset with explicitly labelled positive and negative attributes.
26 papers · 0 benchmarks
VOID (Visual Odometry with Inertial and Depth)
The dataset was collected using the Intel RealSense D435i camera, which was configured to produce synchronized accelerometer and gyroscope measurements at 400 Hz, along with synchronized VGA-size (640 x 480) RGB and depth streams at 30 Hz.
26 papers · 1 benchmark
BAM! (Behance Artistic Media)
The Behance Artistic Media dataset (BAM!) is a large-scale dataset of contemporary artwork from Behance, a website containing millions of portfolios from professional and commercial artists.
25 papers · 0 benchmarks
Chest ImaGenome is a dataset with a scene graph data structure to describe 242,072 images.
25 papers · 0 benchmarks
CrisisMMD is a large multi-modal dataset collected from Twitter during different natural disasters.
25 papers · 0 benchmarks
DIOR-RSVG is a large-scale benchmark dataset of remote sensing data (RSVG).
25 papers · 0 benchmarks
DND (Darmstadt Noise Dataset)
Benchmarking Denoising Algorithms with Real Photographs This dataset consists of 50 pairs of noisy and (nearly) noise-free images captured with four consumer cameras.
25 papers · 2 benchmarks
ELEVATER (Evaluation of Language-augmented Visual Task-level Transfer)
The ELEVATER benchmark is a collection of resources for training, evaluating, and analyzing language-image models on image classification and object detection.
25 papers · 2 benchmarks
This is a dataset for a video super-resolution task.
25 papers · 1 benchmark
ROSE (Retinal OCTA SEgmentation dataset)
Retinal OCTA SEgmentation dataset (ROSE) consists of 229 OCTA images with vessel annotations at either centerline-level or pixel level.
25 papers · 4 benchmarks
TinyPerson is a benchmark for tiny object detection in a long distance and with massive backgrounds.
25 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.