Home › Datasets › modality › Images

Images datasets

archive 2025-07-28

3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 44 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Images datasets 2065–2112 of 3,239

SR-Reg (SynthRAD Registration)
SR-Reg is a brain MR-CT registration dataset, deriving from SynthRAD 2023 (https://synthrad2023.grand-challenge.org/).
2 papers · 1 benchmark
STIR (Scaled and Translated Image Recognition)
While convolutions are known to be invariant to (discrete) translations, scaling continues to be a challenge and most image recognition networks are not invariant to them.
2 papers · 0 benchmarks
SWAX (Sense Wax Attack dataset)
Comprised of real human and wax figure images and videos that endorse the problem of face spoofing detection.
2 papers · 0 benchmarks
Saint Gall dataset contains handwritten historical manuscripts written in Latin that date back to the 9th century.
2 papers · 1 benchmark
ScanBank is a benchmark dataset for figure extraction from scanned electronic theses and dissertations containing 10 thousand scanned page images, manually labeled by humans as to the presence of the 3.3 thousand figures or tables found…
2 papers · 0 benchmarks
SciGen is a challenge dataset for the task of reasoning-aware data-to-text generation consisting of tables from scientific articles and their corresponding descriptions.
2 papers · 0 benchmarks
Large-scale shadows from buildings in a city play an important role in determining the environmental quality of public spaces.
2 papers · 0 benchmarks
The synthetic ShapeNet intrinsic image decomposition dataset used for training the deep CNN models IntrinsicNet and RetiNet of CVPR2018.
2 papers · 0 benchmarks
SmartCity consists of 50 images in total collected from ten city scenes including office entrance, sidewalk, atrium, shopping mall etc..
2 papers · 0 benchmarks
StereoMSI comprises of 350 registered colour-spectral image pairs.
2 papers · 0 benchmarks
A multimodal empathetic dialogue dataset.
2 papers · 0 benchmarks
Text-Vison Cross-Modal Place Recognition Dataset
2 papers · 0 benchmarks
StreetTryOn, the new in-the-wild Virtual Try-On dataset, consists of 12,364 and 2,089 street person images for training and validation, respectively.
2 papers · 1 benchmark
Introduction This dataset supports Ye et al.
2 papers · 0 benchmarks
SuperCaustics is a simulation tool made in Unreal Engine for generating massive computer vision datasets that include transparent objects.
2 papers · 0 benchmarks
SynMirror consists of samples rendered from 3D assets of two widely used 3D object datasets - Objaverse and Amazon Berkeley Objects (ABO) placed in front of a mirror in a virtual blender environment.
2 papers · 0 benchmarks
SynthEVox3D-Tiny (Synthetic Event Camera Voxel 3D Reconstruction Dataset)
Event cameras are sensors that are inspired by biological systems and specialize in capturing changes in brightness.
2 papers · 1 benchmark
TAS-NIR is a VIS+NIR dataset of semantically annotated images in unstructured outdoor environments.
2 papers · 0 benchmarks
TBBR (Thermal Bridges on Building Rooftops)
The dataset of Thermal Bridges on Building Rooftops (TBBR dataset) consists of annotated combined RGB and thermal drone images with a height map.
2 papers · 2 benchmarks
A new text effects dataset with 141,081 text effect/glyph pairs in total.
2 papers · 0 benchmarks
TI1K Dataset (Thumb Index 1000 Hand & Fingertip Detection Dataset)
Thumb Index 1000 (TI1K) is a dataset of 1000 hand images with the hand bounding box, and thumb and index fingertip positions.
2 papers · 0 benchmarks
TNCR Dataset (Table Net Detection and Classification Dataset)
We present TNCR, a new table dataset with varying image quality collected from free open source websites.
2 papers · 0 benchmarks
TRN (Toulouse Road Network)
The Toulouse Road Network dataset describes patches of road maps from the city of Toulouse, represented both as spatial graphs G = (A, X) and as grayscale segmentation images.
2 papers · 1 benchmark
TTE-A&O (Travel Time Estimation: Abakan and Omsk)
The dataset includes two parts corresponding to the cities of Abakan (65524 nodes, 340012 edges) and Omsk (231688 nodes, 1149492 edges).
2 papers · 1 benchmark
TVL Dataset (Touch-Vision-Language Dataset)
Touch is an important sensing modality for humans, but it has not yet been incorporated into a multimodal generative language model.
2 papers · 0 benchmarks
TabLeX is a large-scale benchmark dataset comprising table images generated from scientific articles.
2 papers · 0 benchmarks
Talk2Nav is a large-scale dataset with verbal navigation instructions.
2 papers · 0 benchmarks
Tc1 Mouse cerebellum atlas (Tc1 Mouse cerebellum atlas with Purkinje layer segmentation)
This mouse cerebellar atlas can be used for mouse cerebellar morphometry.
2 papers · 0 benchmarks
The ULS23 test set contains 725 lesions from 284 patients of the Radboudumc and JBZ hospitals in the Netherlands.
2 papers · 1 benchmark
ThermoHands is the first benchmark dataset specifically designed for egocentric 3D hand pose estimation from thermal images.
2 papers · 0 benchmarks
Tiny ImageNet-R is a subset of the ImageNet-R dataset by Hendrycks et al.
2 papers · 0 benchmarks
U-DIADS-Bib is a proprietary dataset developed through the collaboration of computer scientists and humanities at the University of Udine.
2 papers · 1 benchmark
UBOFAB19 (SVBRDF Database Bonn)
A database of several hundred high quality fabric material measurements, provided as carefully calibrated rectified HDR images, together with SVBRDF fits.
2 papers · 0 benchmarks
40,764 images (11,659 protest images and hard negatives) with various annotations of visual attributes and sentiments.
2 papers · 0 benchmarks
This dataset contains 2,000 dial meter images obtained on-site by employees of the Energy Company of Paraná (Copel), which serves more than 4 million consuming units in the Brazilian state of Paraná.
2 papers · 1 benchmark
This dataset contains 2,000 images taken from inside a warehouse of the Energy Company of Paraná (Copel), which directly serves more than 4 million consuming units in the Brazilian state of Paraná.
2 papers · 1 benchmark
The UFPR-Eyeglasses dataset has 1,135 images of both eyes (2,270 cropped images of each eye) from 83 subjects (166 classes).
2 papers · 0 benchmarks
UI5k (Mobile App User Interface Dataset)
This dataset contains 54,987 UI screenshots and the metadata from 7,748 Android applications belonging to 25 application categories Download link: https://www.dropbox.com/sh/kfkhevxykzwputb/AAAhL6ipmOg4zZn4jULmyF0a?dl=0
2 papers · 0 benchmarks
UK Biobank Brain MRI (UK Biobank Data - Brain MRI)
UK Biobank participants have generously provided a very wide range of information about their health and well-being since recruitment began in 2006.
2 papers · 1 benchmark
UW Indoor Scenes (UW-IS) Occluded dataset is curated using commodity hardware (Intel RealSense D435) to reflect real world robotics scenarios.
2 papers · 0 benchmarks
UW-IS (UW Indoor Scenes)
UW-IS (UW Indoor Scenes) is a dataset for object recognition in indoor environments comprising scene images from two different environments, namely, a living room and a mock warehouse.
2 papers · 0 benchmarks
The Udacity dataset is mainly composed of video frames taken from urban roads.
2 papers · 1 benchmark
Underwater Trash Detection Dataset Overview The Underwater Trash Detection Dataset is a custom-annotated dataset designed to address the challenges of underwater trash detection caused by varying environmental features.
2 papers · 0 benchmarks
UofTPed50 is an object detection and tracking dataset which uses GPS to ground truth the position and velocity of a pedestrian.
2 papers · 0 benchmarks
V2VBench is a comprehensive benchmark designed to evaluate video editing methods.
2 papers · 0 benchmarks
325 word images intended for font recognition, whose fonts are included in [VFR-447] (and [VFR-2420]).
2 papers · 1 benchmark
The dataset, VIST-Edit, includes 14,905 human-edited versions of 2,981 machine-generated visual stories.
2 papers · 0 benchmarks
VQA 360° is a dataset for visual question answering on 360° images containing around 17,000 real-world image-question-answer triplets for a variety of question types.
2 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.