Home › Datasets › modality › Images
Images datasets
archive 2025-07-28
3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 21 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Images datasets 961–1008 of 3,239
FoodX-251 is a dataset of 251 fine-grained classes with 118k training, 12k validation and 28k test images.
11 papers · 1 benchmark
The GTA Indoor Motion dataset (GTA-IM) that emphasizes human-scene interactions in the indoor environments.
11 papers · 2 benchmarks
Four pathologists from Longhua Hospital Shanghai University of Traditional Chinese Medicine provide 600 images of gastric cancer pathology images at size 2048×2048 pixels.
11 papers · 1 benchmark
Hilti SLAM Challenge is a dataset for Simultaneous Localization and Mapping (SLAM) algorithms due to sparsity, varying illumination conditions, and dynamic objects.
11 papers · 0 benchmarks
The ISIC 2018 dataset was published by the International Skin Imaging Collaboration (ISIC) as a large-scale dataset of dermoscopy images.
11 papers · 0 benchmarks
ImageCoDe (Image Retrieval from Contextual Descriptions)
Given 10 minimally contrastive (highly similar) images and a complex description for one of them, the task is to retrieve the correct image.
11 papers · 1 benchmark
ImageNet-X is a set of human annotations pinpointing failure types for the popular ImageNet dataset.
11 papers · 0 benchmarks
Kaggle EyePACS (Kaggle EyePACS. Diabetic Retinopathy Detection Identify signs of diabetic retinopathy in eye images)
Diabetic retinopathy is the leading cause of blindness in the working-age population of the developed world.
11 papers · 1 benchmark
A large-scale Landmark guided face Parsing dataset (LaPa) for face parsing.
11 papers · 1 benchmark
MultI-Modal In-Context Instruction Tuning (MIMIC-IT) is a dataset for instruction tuning into multi-modal models, motivated by the Flamingo model's upstream interleaved format pretraining dataset.
11 papers · 0 benchmarks
A data-set which consists of over one million images of physical 3D objects with seven factors of variation, such as object color, shape, size and position.
11 papers · 0 benchmarks
Multi-Modal Reading (MMR) Benchmark includes 550 annotated question-answer pairs across 11 distinct tasks involving texts, fonts, visual elements, bounding boxes, spatial relations, and grounding, with carefully designed evaluation metrics.
11 papers · 1 benchmark
MSD (Million Song Dataset)
The Million Song Dataset is a freely-available collection of audio features and metadata for a million contemporary popular music tracks.
11 papers · 2 benchmarks
OpenLane-V2 is the world's first perception and reasoning benchmark for scene structure in autonomous driving.
11 papers · 2 benchmarks
P-DukeMTMC-reID is a modified version based on DukeMTMC-reID dataset.
11 papers · 1 benchmark
A large-scale collection of visually-grounded, task-oriented dialogues in English designed to investigate shared dialogue history accumulating during conversation.
11 papers · 0 benchmarks
Place Pulse is a crowdsourcing effort that aims to map which areas of a city are perceived as safer, livelier, wealthier, more active, beautiful and friendly.
11 papers · 1 benchmark
RGB-D-D is a large-scale dataset for depth map super-resolution (SR).
11 papers · 0 benchmarks
RMFD (Real-World Masked Face Dataset)
Real-World Masked Face Dataset (RMFD) is a large dataset for masked face detection.
11 papers · 0 benchmarks
RailSem19 (RailSem19: A Dataset for Semantic Rail Scene Understanding)
RailSem19 offers 8500 unique images taken from a the ego-perspective of a rail vehicle (trains and trams).
11 papers · 0 benchmarks
A large-scale dataset of ~29.5K rain/rain-free image pairs that covers a wide range of natural rain scenes.
11 papers · 0 benchmarks
So2Sat LCZ42 consists of local climate zone (LCZ) labels of about half a million Sentinel-1 and Sentinel-2 image patches in 42 urban agglomerations (plus 10 additional smaller areas) across the globe.
11 papers · 1 benchmark
Synbols is a dataset generator designed for probing the behavior of learning algorithms.
11 papers · 0 benchmarks
TACO is a growing image dataset of waste in the wild.
11 papers · 0 benchmarks
The first large demoire dataset.
11 papers · 1 benchmark
TUM monoVO is a dataset for evaluating the tracking accuracy of monocular Visual Odometry (VO) and SLAM methods.
11 papers · 0 benchmarks
Talk The Walk is a large-scale dialogue dataset grounded in action and perception.
11 papers · 0 benchmarks
UDIS-D (Unsupervised Deep Image Stitching Dataset)
UDIS-D is a large image dataset for image stitching or image registration.
11 papers · 0 benchmarks
UnrealEgo is a dataset that provides in-the-wild stereo images with a large variety of motions for 3D human pose estimation.
11 papers · 1 benchmark
VinDr-CXR is an open large-scale dataset of chest X-rays with radiologist’s annotations.
11 papers · 0 benchmarks
WiderPerson contains a total of 13,382 images with 399,786 annotations, i.e., 29.87 annotations per image, which means this dataset contains dense pedestrians with various kinds of occlusions.
11 papers · 1 benchmark
These data are the results of a chemical analysis of wines grown in the same region in Italy but derived from three different cultivars.
11 papers · 6 benchmarks
nvBench is a large-scale NL2VIS (natural languagge to visualisations) benchmark, containing 25,750 (NL, VIS) pairs from 750 tables over 105 domains, synthesized from (NL, SQL) benchmarks to support cross-domain NLPVIS (Natural Language…
11 papers · 0 benchmarks
2-PM Vessel is an open-source volumetric brain vasculature dataset obtained with two-photon microscopy at Focused Ultrasound Lab, at Sunnybrook Research Institute (affiliated with University of Toronto by Dr.
10 papers · 0 benchmarks
Novel benchmark which features aspects of natural scenes, e.g.
10 papers · 1 benchmark
We release expert-made scribble annotations for the medical ACDC dataset [1].
10 papers · 1 benchmark
ATD-12K is a large-scale animation triplet dataset, which comprises 12,000 triplets(train10k,test2k) by manually inspect and the test2k with rich annotations, including levels of difficulty, the Regions of Interest (RoIs) on movements, and…
10 papers · 1 benchmark
CDDB (Continual Deepfake Detection Benchmark)
Abstract: There have been emerging a number of benchmarks and techniques for the detection of deepfakes.
10 papers · 0 benchmarks
CLEVR-Dialog is a large diagnostic dataset for studying multi-round reasoning in visual dialog.
10 papers · 0 benchmarks
CalMS21 (Caltech Mouse Social Interactions)
The Caltech Mouse Social Interactions (CalMS21) dataset is a multi-agent dataset from behavioral neuroscience.
10 papers · 0 benchmarks
Chart2Text is a dataset that was crawled from 23,382 freely accessible pages from statista.com in early March of 2020, yielding a total of 8,305 charts, and associated summaries.
10 papers · 0 benchmarks
The CropAndWeed dataset is focused on the fine-grained identification of 74 relevant crop and weed species with a strong emphasis on data variability.
10 papers · 0 benchmarks
CubiCasa5K is a large-scale floorplan image dataset containing 5000 samples annotated into over 80 floorplan object categories.
10 papers · 0 benchmarks
DAWN emphasizes a diverse traffic environment (urban, highway and freeway) as well as a rich variety of traffic flow.
10 papers · 0 benchmarks
DOTA 2.0 (Dataset of Object deTection in Aerial images)
—In the past decade, object detection has achieved significant progress in natural images but not in aerial images, due to the massive variations in the scale and orientation of objects caused by the bird’s-eye view of aerial images.
10 papers · 0 benchmarks
DeepScores contains high quality images of musical scores, partitioned into 300,000 sheets of written music that contain symbols of different shapes and sizes.
10 papers · 0 benchmarks
A new dataset of handwritten text with fine-grained annotations at the character level and report results from an initial user evaluation.
10 papers · 0 benchmarks
Although deep face recognition has achieved impressive results in recent years, there is increasing controversy regarding racial and gender bias of the models, questioning their trustworthiness and deployment into sensitive scenarios.
10 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.