Home › Datasets › modality › Images
Images datasets
archive 2025-07-28
3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 10 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Images datasets 433–480 of 3,239
HRF (High-Resolution Fundus)
The HRF dataset is a dataset for retinal vessel segmentation which comprises 45 images and is organized as 15 subsets.
53 papers · 3 benchmarks
Large language models (LLMs), after being aligned with vision models and integrated into vision-language models (VLMs), can bring impressive improvement in image reasoning tasks.
53 papers · 1 benchmark
The ICDAR2003 dataset is a dataset for scene text recognition.
53 papers · 1 benchmark
The MMVP (Multimodal Visual Patterns) Benchmark focuses on identifying "CLIP-blind pairs" – images that appear similar to the CLIP model despite having clear visual differences.
53 papers · 1 benchmark
RICH (Real scenes, Interaction, Contact and Humans)
Inferring human-scene contact (HSC) is the first step toward understanding how humans interact with their surroundings.
53 papers · 1 benchmark
ACDC (Automated Cardiac Diagnosis Challenge)
The goal of the Automated Cardiac Diagnosis Challenge (ACDC) challenge is to: - compare the performance of automatic methods on the segmentation of the left ventricular endocardium and epicardium as the right ventricular endocardium for…
52 papers · 5 benchmarks
ICDAR 2015 was a scene text detection used for the ICDAR 2015 conference.
52 papers · 2 benchmarks
InfographicVQA is a dataset that comprises a diverse collection of infographics along with natural language questions and answers annotations.
52 papers · 1 benchmark
The MultiMNIST dataset is generated from MNIST.
52 papers · 1 benchmark
NJU2K is a large RGB-D dataset containing 1,985 image pairs.
52 papers · 1 benchmark
The Object Discovery dataset was collected by downloading images from Internet for airplane, car and horse.
52 papers · 1 benchmark
Pick-a-Pic dataset was created by logging user interactions with the Pick-a-Pic web application for text-to image generation.
52 papers · 0 benchmarks
The Scene UNderstanding (SUN) database contains 899 categories and 130,519 images.
52 papers · 8 benchmarks
Sprites (2D Video Game Character Sprites)
The Sprites dataset contains 60 pixel color images of animated characters (sprites).
52 papers · 3 benchmarks
The Event-Camera Dataset is a collection of datasets with an event-based camera for high-speed robotics.
51 papers · 2 benchmarks
Fishyscapes is a public benchmark for uncertainty estimation in a real-world task of semantic segmentation for urban driving.
51 papers · 2 benchmarks
SiW (Spoofing in the Wild)
SiW provides live and spoof videos from 165 subjects.
51 papers · 1 benchmark
The xBD dataset contains over 45,000KM2 of polygon labeled pre and post disaster imagery.
51 papers · 2 benchmarks
Lesion segmentation data includes the original image, paired with the expert manual tracing of the lesion boundaries in the form of a binary mask.
50 papers · 1 benchmark
The InterHand2.6M dataset is a large-scale real-captured dataset with accurate GT 3D interacting hand poses, used for 3D hand pose estimation The dataset contains 2.6M labeled single and interacting hand frames.
50 papers · 2 benchmarks
PlotQA is a VQA dataset with 28.9 million question-answer pairs grounded over 224,377 plots on data from real-world sources and questions based on crowd-sourced question templates.
50 papers · 5 benchmarks
The Quick Draw Dataset is a collection of 50 million drawings across 345 categories, contributed by players of the game Quick, Draw!.
50 papers · 0 benchmarks
2018 Data Science Bowl (2018 Data Science Bowl Find the nuclei in divergent images to advance medical discovery)
This dataset contains a large number of segmented nuclei images.
49 papers · 1 benchmark
DENSE (Depth Estimation oN Synthetic Events)
DENSE (Depth Estimation oN Synthetic Events) is a new dataset with synthetic events and perfect ground truth.
49 papers · 1 benchmark
DVQA (Data Visualizations via Question Answering)
DVQA is a synthetic question-answering dataset on images of bar-charts.
49 papers · 1 benchmark
FSS-1000 is a 1000 class dataset for few-shot segmentation.
49 papers · 1 benchmark
Letter (Letter Recognition Data Set)
Letter Recognition Data Set is a handwritten digit dataset.
49 papers · 2 benchmarks
MMKG is a collection of three knowledge graphs for link prediction and entity matching research.
49 papers · 3 benchmarks
The Pavia University dataset is a hyperspectral image dataset which gathered by a sensor known as the reflective optics system imaging spectrometer (ROSIS-3) over the city of Pavia, Italy.
49 papers · 1 benchmark
A new large-scale benchmark consisting of both synthetic and real-world hazy images, called REalistic Single Image DEhazing (RESIDE).
49 papers · 5 benchmarks
WildDeepfake is a dataset for real-world deepfakes detection which consists of 7,314 face sequences extracted from 707 deepfake videos that are collected completely from the internet.
49 papers · 0 benchmarks
3D-FUTURE (3D FUrniture shape with TextURE) is a 3D dataset that contains 20,240 photo-realistic synthetic images captured in 5,000 diverse scenes, and 9,992 involved unique industrial 3D CAD shapes of furniture with high-resolution…
48 papers · 0 benchmarks
JHU-CROWD++ is A large-scale unconstrained crowd counting dataset with 4,372 images and 1.51 million annotations.
48 papers · 1 benchmark
The Objectron dataset is a collection of short, object-centric video clips, which are accompanied by AR session metadata that includes camera poses, sparse point-clouds and characterization of the planar surfaces in the surrounding…
48 papers · 0 benchmarks
SRD (Shadow Removal Dataset)
SRD is a dataset for shadow removal that contains 3088 shadow and shadow-free image pairs.
48 papers · 1 benchmark
CityFlow is a city-scale traffic camera dataset consisting of more than 3 hours of synchronized HD videos from 40 cameras across 10 intersections, with the longest distance between two simultaneous cameras being 2.5 km.
47 papers · 1 benchmark
Gait3D is a large-scale 3D representation-based gait recognition dataset.
47 papers · 2 benchmarks
ScreenSpot Evaluation Benchmark ScreenSpot is an evaluation benchmark for GUI grounding, comprising over 1,200 instructions from various environments, including iOS, Android, macOS, Windows, and Web.
47 papers · 1 benchmark
FLIC (Frames Labelled in Cinema)
The FLIC dataset contains 5003 images from popular Hollywood movies.
46 papers · 2 benchmarks
OpenLane is the first real-world and the largest scaled 3D lane dataset to date.
46 papers · 2 benchmarks
The Stanford Background dataset contains 715 RGB images and the corresponding label images.
46 papers · 0 benchmarks
Synscapes is a synthetic dataset for street scene parsing created using photorealistic rendering techniques, and show state-of-the-art results for training and validation as well as new types of analysis.
46 papers · 1 benchmark
A new large-scale geometry problem-solving dataset - 3,002 multi-choice geometry problems - dense annotations in formal language for the diagrams and text - 27,213 annotated diagram logic forms (literals) - 6,293 annotated text logic forms…
45 papers · 1 benchmark
HICO (Humans Interacting with Common Objects)
HICO is a benchmark for recognizing human-object interactions (HOI).
45 papers · 1 benchmark
LiTS17 (Liver Tumor Segmentation Challenge 2017)
LiTS17 is a liver tumor segmentation benchmark.
45 papers · 3 benchmarks
DailyActivity3D dataset is a daily activity dataset captured by a Kinect device.
45 papers · 1 benchmark
MVTec 3D Anomaly Detection Dataset (MVTec 3D-AD) is a comprehensive 3D dataset for the task of unsupervised anomaly detection and localization.
45 papers · 4 benchmarks
RSTPReid (Real Scenario Text-based Person Re-identification)
RSTPReid contains 20505 images of 4,101 persons from 15 cameras.
45 papers · 2 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.