Home › Datasets › modality › Images
Images datasets
archive 2025-07-28
3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 18 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Images datasets 817–864 of 3,239
GSV-Cities is a large-scale dataset for training deep neural network for the task of Visual Place Recognition.
16 papers · 0 benchmarks
GeneCIS benchmark is designed for measuring models’ ability to adapt to a range of similarity conditions, which is zero-shot evaluation only.
16 papers · 1 benchmark
ICFG-PEDES (Identity-Centric and Fine-Grained Person Description Dataset)
One large-scale database for Text-to-Image Person Re-identification, i.e., Text-based Person Retrieval.
16 papers · 3 benchmarks
IDRiD (Indian Diabetic Retinopathy Image Dataset)
Indian Diabetic Retinopathy Image Dataset (IDRiD) dataset consists of typical diabetic retinopathy lesions and normal retinal structures annotated at a pixel level.
16 papers · 3 benchmarks
MinneApple is a benchmark dataset for apple detection and segmentation.
16 papers · 0 benchmarks
The NYU Hand pose dataset contains 8252 test-set and 72757 training-set frames of captured RGBD data with ground-truth hand-pose information.
16 papers · 1 benchmark
NewsCLIPpings is a dataset for detecting mismatched images and captions.
16 papers · 0 benchmarks
SKM-TEA (Stanford Knee MRI with Multi-Task Evaluation)
The SKM-TEA dataset pairs raw quantitative knee MRI (qMRI) data, image data, and dense labels of tissues and pathology for end-to-end exploration and evaluation of the MR imaging pipeline.
16 papers · 0 benchmarks
SMHD (Self-reported Mental Health Diagnoses)
A novel large dataset of social media posts from users with one or multiple mental health conditions along with matched control users.
16 papers · 0 benchmarks
SemArt is a multi-modal dataset for semantic art understanding.
16 papers · 0 benchmarks
SketchGraphs is a dataset of 15 million sketches extracted from real-world CAD models intended to facilitate research in both ML-aided design and geometric program induction.
16 papers · 0 benchmarks
SketchyCOCO dataset consists of two parts: Object-level data Object-level data contains 20198(train18869+val1329) triplets of {foreground sketch, foreground image, foreground edge map} examples covering 14 classes,…
16 papers · 1 benchmark
TextSeg is a large-scale fine-annotated and multi-purpose text detection and segmentation dataset, collecting scene and design text with six types of annotations: word- and character-wise bounding polygons, masks and transcriptions.
16 papers · 1 benchmark
V-D4RL provides pixel-based analogues of the popular D4RL benchmarking tasks, derived from the dmcontrol suite, along with natural extensions of two state-of-the-art online pixel-based continuous control algorithms, DrQ-v2 and DreamerV2,…
16 papers · 0 benchmarks
VehicleX is a large-scale synthetic dataset.
16 papers · 0 benchmarks
VideoLQ consists of videos downloaded from various video hosting sites such as Flickr and YouTube, with a Creative Common license.
16 papers · 1 benchmark
Wukong is a large-scale Chinese cross-modal dataset for benchmarking different multi-modal pre-training methods to facilitate the Vision-Language Pre-training (VLP).
16 papers · 0 benchmarks
XQLFW (Cross-Quality Labeled Faces in the Wild)
An evaluation protocol for face verification focusing on a large intra-pair image quality difference.
16 papers · 1 benchmark
e-SNLI-VE is a large VL (vision-language) dataset with NLEs (natural language explanations) with over 430k instances for which the explanations rely on the image content.
16 papers · 2 benchmarks
4Seasons is adataset covering seasonal and challenging perceptual conditions for autonomous driving.
15 papers · 0 benchmarks
ACRE (Abstract Causal REasoning)
Abstract Causal REasoning (ACRE) is a dataset for the systematic evaluation of current vision systems in causal induction, i.e., identifying unobservable mechanisms that lead to the observable relations among variables.
15 papers · 0 benchmarks
CASIA V2 is a dataset for forgery classification.
15 papers · 0 benchmarks
Fakeddit is a novel multimodal dataset for fake news detection consisting of over 1 million samples from multiple categories of fake news.
15 papers · 0 benchmarks
The GoodsAD dataset contains 6124 images with 6 categories of common supermarket goods.
15 papers · 1 benchmark
HaGRID (HaGRID - HAnd Gesture Recognition Image Dataset)
We introduce a large image dataset HaGRID (HAnd Gesture Recognition Image Dataset) for hand gesture recognition (HGR) systems.
15 papers · 0 benchmarks
The ISIC 2017 dataset was published by the International Skin Imaging Collaboration (ISIC) as a large-scale dataset of dermoscopy images.
15 papers · 0 benchmarks
ISTD+ consists of shadow images, shadow-free images, and shadow masks, with 1,330 training images and 540 testing images from 135 unique background scenes.
15 papers · 1 benchmark
The dataset is constructed from images of defective production items that were provided and annotated by Kolektor Group d.o.o..
15 papers · 1 benchmark
KolektorSDD2 is a surface-defect detection dataset with over 3000 images containing several types of defects, obtained while addressing a real-world industrial problem.
15 papers · 2 benchmarks
MALF (Multi-Attribute Labelled Faces)
The MALF dataset is a large dataset with 5,250 images annotated with multiple facial attributes and it is specifically constructed for fine grained evaluation.
15 papers · 0 benchmarks
MIAP (More Inclusive Annotations for People)
MIAP is a dataset created by obtaining a new set of annotations on a subset of the Open Images dataset, containing bounding boxes and attributes for all of the people visible in those images, as the original Open Images dataset annotations…
15 papers · 0 benchmarks
Existing hate speech datasets contain only textual data.
15 papers · 0 benchmarks
The MSK dataset is a dataset for lesion recognition from the Memorial Sloan-Kettering Cancer Center.
15 papers · 0 benchmarks
MSRA Hands is a dataset for hand tracking.
15 papers · 1 benchmark
OCTID (Optical Coherence Tomography Image Retinal Database)
An open-source Optical Coherence Tomography Image Database containing different retinal OCT images with various pathological conditions.
15 papers · 0 benchmarks
The Paris-Lille-3D is a Benchmark on Point Cloud Classification.
15 papers · 1 benchmark
The goal of PubTables-1M is to create a large, detailed, high-quality dataset for training and evaluating a wide variety of models for the tasks of table detection, table structure recognition, and functional analysis.
15 papers · 0 benchmarks
QED is a linguistically principled framework for explanations in question answering.
15 papers · 1 benchmark
REDS (REalistic and Diverse Scenes dataset
realistic and dynamic scenes)
The realistic and dynamic scenes (REDS) dataset was proposed in the NTIRE19 Challenge.
15 papers · 1 benchmark
RoadAnomaly21 is a dataset for anomaly segmentation, the task of identify the image regions containing objects that have never been seen during training.
15 papers · 0 benchmarks
SYSU-30k contains 30k categories of persons, which is about 20 times larger than CUHK03 (1.3k categories) and Market1501 (1.5k categories), and 30 times larger than ImageNet (1k categories).
15 papers · 2 benchmarks
Syn2Real, a synthetic-to-real visual domain adaptation benchmark meant to encourage further development of robust domain transfer methods.
15 papers · 1 benchmark
UBody is a large-scale Upper-Body dataset with the following annotations: 2D whole-body keypoints 3D SMPLX annotations Frame validity label Person bounding box (bbox) Hand bounding box (bbox)
15 papers · 1 benchmark
This dataset includes 4,500 fully annotated images (over 30,000 license plate characters) from 150 vehicles in real-world scenarios where both the vehicle and the camera (inside another vehicle) are moving.
15 papers · 1 benchmark
Washington RGB-D is a widely used testbed in the robotic community, consisting of 41,877 RGB-D images organized into 300 instances divided in 51 classes of common indoor objects (e.g.
15 papers · 0 benchmarks
2021 Hotel-ID is a dataset for hotel recognition to help raise awareness of human trafficking and generate novel approaches.
14 papers · 0 benchmarks
4DFAB is a large scale database of dynamic high-resolution 3D faces which consists of recordings of 180 subjects captured in four different sessions spanning over a five-year period (2012 - 2017), resulting in a total of over 1,800,000 3D…
14 papers · 0 benchmarks
ELPV (A dataset of functional and defective solar cells extracted from EL images of solar modules)
The dataset contains 2,624 samples of 300×300 pixels 8-bit grayscale images of functional and defective solar cells with varying degree of degradations extracted from 44 different solar modules.
14 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.