Home › Datasets › modality › Images

Images datasets

archive 2025-07-28

3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 38 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Images datasets 1777–1824 of 3,239

TLL (Totally-Looks-Like)
Contains 6016 image-pairs from the wild, shedding light upon a rich and diverse set of criteria employed by human beings.
3 papers · 0 benchmarks
Tiny ImageNetv2 is a subset of the ImageNetV2 (matched frequency) dataset by Recht et al.
3 papers · 0 benchmarks
Tobacco800 is a public subset of the complex document image processing (CDIP) test collection constructed by Illinois Institute of Technology, assembled from 42 million pages of documents (in 7 million multi-page TIFF images) released by…
3 papers · 0 benchmarks
Dataset of restaurant reviews from TripAdvisor that includes images and texts uploaded in reviews by users.
3 papers · 0 benchmarks
Tsinghua-Tencent 100K (Traffic-Sign Detection and Classification in the Wild)
Although promising results have been achieved in the areas of traffic-sign detection and classification, few works have provided simultaneous solutions to these two tasks for realistic real world images.
3 papers · 1 benchmark
TuSimple Lane is an extension of the TuSimple dataset with 14,336 lane boundaries annotations.
3 papers · 0 benchmarks
Detecting out-of-context media, such as "mis-captioned" images on Twitter, is a relevant problem, especially in domains of high public significance.
3 papers · 0 benchmarks
The archive contains original images from U2OS cells stained with Hoechst 33342 as PNG files.
3 papers · 0 benchmarks
UASOL (A large-scale high-resolution outdoor stereo dataset)
The UASOL an RGB-D stereo dataset, that contains 160902 frames, filmed at 33 different scenes, each with between 2 k and 10 k frames.
3 papers · 1 benchmark
UAV-GESTURE is a dataset for UAV control and gesture recognition.
3 papers · 0 benchmarks
UAVA (UAV Assistant)
The UAVA,UAV-Assistant, dataset is specifically designed for fostering applications which consider UAVs and humans as cooperative agents.
3 papers · 0 benchmarks
The UFPR-Periocular dataset has 16,830 images of both eyes (33,660 cropped images of each eye) from 1,122 subjects (2,244 classes).
3 papers · 0 benchmarks
UIIS10K (General Underwater Image Instance Segmentation dataset 10K)
We propose a large-scale underwater instance segmentation dataset, UIIS10K, which includes 10,048 images with pixel-level annotations for 10 categories.
3 papers · 0 benchmarks
The Sheffield (previously UMIST) Face Database consists of 564 images of 20 individuals (mixed race/gender/appearance).
3 papers · 1 benchmark
UNDD (Urban Night Driving Dataset)
UNDD consists of 7125 unlabelled day and night images; additionally, it has 75 night images with pixel-level annotations having classes equivalent to Cityscapes dataset.
3 papers · 0 benchmarks
The newly introduced UP-COUNT dataset includes drone footage captured with cameras from the DJI Mini 2 family UAV.
3 papers · 1 benchmark
The US-4 is a dataset of Ultrasound (US) images.
3 papers · 0 benchmarks
UZLF (Leuven-Haifa High-Resolution Fundus Image Dataset for Retinal Blood Vessel Segmentation and Glaucoma Diagnosis)
The Leuven-Haifa dataset contains 240 disc-centered fundus images of 224 unique patients (75 patients with normal tension glaucoma, 63 patients with high tension glaucoma, 30 patients with other eye diseases and 56 healthy controls) from…
3 papers · 2 benchmarks
Vehicle-Rear is a novel dataset for vehicle identification that contains more than three hours of high-resolution videos, with accurate information about the make, model, color and year of nearly 3,000 vehicles, in addition to the position…
3 papers · 0 benchmarks
a vessel dataset using 85 videos.
3 papers · 1 benchmark
Hugging Face Datasets (New!) | Website | Github Repository | arXiv e-Print The Visual Writing Prompts (VWP) dataset contains almost 2K selected sequences of movie shots, each including 5-10 images.
3 papers · 0 benchmarks
The Vocal Folds dataset is a dataset for automatic segmentation of laryngeal endoscopic images.
3 papers · 0 benchmarks
WFDD (Woven Fabric Defect Detection)
WFDD is a dataset for benchmarking anomaly detection methods with a focus on textile inspection.
3 papers · 1 benchmark
WHU-Specular is a large dataset of annotated specular highlight regions created from real-world images.
3 papers · 0 benchmarks
WLD (WildLife Documentary)
WildLife Documentary is an animal object detection dataset.
3 papers · 0 benchmarks
The WebVid-CoVR dataset is a collection of video-text-video triplets that can be used for the task of composed video retrieval (CoVR).
3 papers · 1 benchmark
The WikiScenes dataset consists of paired images and language descriptions capturing world landmarks and cultural sites, with associated 3D models and camera poses.
3 papers · 0 benchmarks
This work introduces Zambezi Voice, an open-source multilingual speech resource for Zambian languages.
3 papers · 0 benchmarks
fluocells (Fluorescent Neuronal Cells)
By releasing this dataset, we aim at providing a new testbed for computer vision techniques using Deep Learning.
3 papers · 0 benchmarks
iFakeFaceDB is a face image dataset for the study of synthetic face manipulation detection, comprising about 87,000 synthetic face images generated by the Style-GAN model and transformed with the GANprintR approach.
3 papers · 0 benchmarks
iWildCam 2021 is a dataset for counting the number of animals of each species that appear in sequences of images captured with camera traps.
3 papers · 0 benchmarks
This package provides utilities for generation, filtering, solving, visualizing, and processing of mazes for training ML systems.
3 papers · 0 benchmarks
360-SOD contains 500 high-resolution equirectangular images.
2 papers · 0 benchmarks
This work was undertaken by members of the Lincoln Centre for Autonomous Systems, University of Lincoln, UK.
2 papers · 0 benchmarks
3D-Point Cloud dataset of various geometrical terrains (3D-Point Cloud dataset of various geometrical terrains in urban environments recorded during human locomotion)
Depth vision has been recently used in many locomotion devices with the objective to ease the life of disabled people toward reaching more ecological lifestyle.
2 papers · 0 benchmarks
4D Light Field Dataset is a light field benchmark consisting of 24 carefully designed synthetic, densely sampled 4D light fields with highly accurate disparity ground truth.
2 papers · 1 benchmark
Description: 5,011 Images – Human Frontal face Data (Male).
2 papers · 0 benchmarks
ALGAD (Andy Lomas Generative Art Dataset)
Repository of a generative art dataset by computer artist Andy Lomas.
2 papers · 0 benchmarks
ALM-Bench (All Languages Matter Benchmark)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
ARC-AGI (Abstraction and Reasoning Corpus for Artificial General Intelligence)
The ARC-AGI benchmark is a significant measure in the field of artificial intelligence, focusing on an AI's general reasoning capabilities.
2 papers · 0 benchmarks
The ARKitFace dataset is established by this work in order to train and evaluate both 3D face shape and 6DoF in the setting of perspective projection.
2 papers · 1 benchmark
ASIRRA ((Animal Species Image Recognition for Restricting Access)
Web services are often protected with a challenge that's supposed to be easy for people to solve, but difficult for computers.
2 papers · 0 benchmarks
The AU-AIR is a multi-modal aerial dataset captured by a UAV.
2 papers · 0 benchmarks
AV Digits Database is an audiovisual database which contains normal, whispered and silent speech.
2 papers · 0 benchmarks
AdvNet is a dataset of traffic signs images.
2 papers · 0 benchmarks
Alsat-2B is a remote sensing dataset of low and high spatial resolution images (10m and 2.5m respectively) for the single-image super-resolution task.
2 papers · 0 benchmarks
Ambiguous-HOI is a challenging dataset containing ambiguous human-object interaction images for HOI detection based on HICO-DET.
2 papers · 0 benchmarks
We present a novel Animation CelebHeads dataset (AnimeCeleb) to address an animation head reenactment.
2 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.