Home › Datasets › modality › Images

Images datasets

archive 2025-07-28

3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 61 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Images datasets 2881–2928 of 3,239

The TwinSynths dataset is a novel benchmark designed to overcome common limitations found in earlier synthetic image datasets, such as low image quality, inadequate content preservation, and limited class diversity.
1 paper · 0 benchmarks
This dataset contains two subsets of flood images from Twitter: The Harz17 dataset comprises images from tweets containing flood-related keywords during the occurrence of a flood in the Harz region in Germany in July 2017.
1 paper · 0 benchmarks
Twitter MediaEval (MediaEval Benchmarking Initiative for Multimedia Evaluation)
The task addresses the problem of the appearance and propagation of posts that share misleading multimedia content (images or video).
1 paper · 0 benchmarks
This task aims to extract named entities and entity types while further predicting segmentation masks of visual objects.
1 paper · 1 benchmark
This dataset supports the research detailed in the pre-print "Virtual Imaging Trials Improved the Transparency and Reliability of AI Systems in COVID-19 Imaging." The study employs both clinical and simulated CT data to evaluate AI models…
1 paper · 1 benchmark
U2-BENCH is the first large-scale benchmark for evaluating Large Vision-Language Models (LVLMs) on ultrasound imaging understanding.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
UAVBillboards (UAV Billboards)
Mapping urban large-area advertising structures using drone imagery and deep learning-based spatial data analysis.
1 paper · 1 benchmark
UAVDB (Trajectory-Guided Adaptable Bounding Boxes for UAV Detection)
UAVDB is a high-resolution RGB video dataset meticulously designed for UAV detection tasks across diverse scales and complex backgrounds.
1 paper · 1 benchmark
The UAVVaste dataset consists to date of 772 images and 3716 annotations.
1 paper · 1 benchmark
https://github.com/zzr-idam/Under-Display-Camera-UAV
1 paper · 0 benchmarks
UDA-CH (Unsupervised Domain Adaptation on Cultural Heritage)
UDA-CH contains 16 objects that cover a variety of artworks which can be found in a museum like sculptures, paintings and books.
1 paper · 1 benchmark
We introduce a set of 425 panoramic X-rays with Human annotated Bounding Boxes and Polygons, the 425 images are a subset of UFBA-UESC Dental Dataset.
1 paper · 1 benchmark
The UFPR-ADMR-v2 dataset contains 5,000 dial meter images obtained on-site by employees of the Energy Company of Paraná (Copel), which serves more than 4M consuming units in the Brazilian state of Paraná.
1 paper · 0 benchmarks
The UFPR-VCR dataset contains 10,039 images of 9,502 distinct vehicles across various categories, including cars, vans, buses, and trucks.
1 paper · 0 benchmarks
UHCSDB (Ultrahigh Carbon Steel micrograph DataBase)
DeCost, Hecht, Francis, Webler, Picard, and Holm.
1 paper · 0 benchmarks
UICaption is a dataset of 114k UI images paired with descriptions of their functionality.
1 paper · 0 benchmarks
UIUC Scooping Dataset (Granular Materials Manipulation Dataset with Scooping/Digging/Excavation Action)
Overview: This dataset encompasses a compilation of 6,700 executed scoops (excavations), mapped across a vast spectrum of materials, terrain topography, and compositions.
1 paper · 0 benchmarks
ULI-RI (Unreal Labeled Images for Person Re-ID)
The ULI-RI dataset is generated using the Unreal Engine 4 to simulate various outdoor environments with 115 high-quality 3D human models.
1 paper · 0 benchmarks
In this dataset UR5 robot used 6 tools: metal-scissor, metal-whisk, plastic-knife, plastic-spoon, wooden-chopstick, and wooden-fork to perform 6 behaviors: look, stirring-slow, stirring-fast, stirring-twist, whisk, and poke.
1 paper · 0 benchmarks
USB (Universal-Scale Object Detection Benchmark)
The Universal-Scale object detection Benchmark (USB) is a benchmark for object detection that has variations in object scales and image domains by incorporating COCO with the recently proposed Waymo Open Dataset and Manga109-s dataset.
1 paper · 1 benchmark
We introduce USPTO-30K, a large-scale benchmark dataset of annotated molecule images, which overcomes these limitations.
1 paper · 0 benchmarks
The semantic segmentation of clothes is a challenging task due to the wide variety of clothing styles, layers and shapes.
1 paper · 1 benchmark
The UTRSet-Real dataset is a comprehensive, manually annotated dataset specifically curated for Printed Urdu OCR research.
1 paper · 0 benchmarks
The UTRSet-Synth dataset is introduced as a complementary training resource to the UTRSet-Real Dataset, specifically designed to enhance the effectiveness of Urdu OCR models.
1 paper · 0 benchmarks
UV6K (Urban Vehicle Segmentation Dataset)
UV6K is a high-resolution remote sensing urban vehicle segmentation dataset.
1 paper · 1 benchmark
UW IOM (University of Washington Indoor Object Manipulation)
Comprises twenty individuals picking up and placing objects of varying weights to and from cabinet and table locations at various heights.
1 paper · 0 benchmarks
Street-View images captured at different timestamps often undergo geometric transformations.
1 paper · 1 benchmark
This dataset extends the Semantic Segmentation of Underwater Imagery: Dataset and Benchmark, adding an uncertainty evaluation component.
1 paper · 0 benchmarks
Unpaired haze images (Unpaired haze images from google images)
Unpaired dataset: The dataset is built by ourselves, and there are all real haze images from websites.
1 paper · 0 benchmarks
Unsplash2K is high-resolution image dataset with 2K resolution.
1 paper · 0 benchmarks
Unsplash_1k (Unsplash_1k_crops)
Inpainting networks are typically benchmarked on samples from Places2 dataset.
1 paper · 0 benchmarks
The prospective upper body thermal images SARS-CoV2 association study was designed to test the hypothesis that thermal videos may aid in the early diagnosis of COVID-19.
1 paper · 0 benchmarks
The UrduDoc Dataset is a benchmark dataset for Urdu text line detection in scanned documents.
1 paper · 1 benchmark
V-MIND enhanced the MIND dataset with news pictures.
1 paper · 0 benchmarks
V-PCCD (simulated Visual Point Cloud Change Detection dataset)
A simulated dataset built in Unreal Engine 4 with AirSim.
1 paper · 0 benchmarks
V2AIX (A Multi-Modal Real-World Dataset of ETSI ITS V2X Messages in Public Road Traffic)
Connectivity is a main driver for the ongoing megatrend of automated mobility: future Cooperative Intelligent Transport Systems (C-ITS) will connect road vehicles, traffic signals, roadside infrastructure, and even vulnerable road users,…
1 paper · 0 benchmarks
This task stems from the observation that text embedded in images is intrinsically different from common visual elements and natural language due to the need to align the modalities of vision, text, and text embedded in images.
1 paper · 0 benchmarks
VD-Ref is a dataset with ground-truth mappings from both noun phrases and pronouns to image regions.
1 paper · 0 benchmarks
VDQG (Visual Discriminative Question Generation)
The Visual Discriminative Question Generation (VDQG) dataset contains 11202 ambiguous image pairs collected from Visual Genome.
1 paper · 0 benchmarks
VESSEL12 (VESsel SEgmentation in the Lung 2012)
1 paper · 0 benchmarks
VETRA is a dataset for vehicle tracking in aerial image sequences and presents unique challenges such as low frame rates, small and fast-moving objects, as well as high camera movement.
1 paper · 0 benchmarks
A synthetic dataset containing word images of 447 typefaces with font variations for each typeface, created for visual font recognition.
1 paper · 1 benchmark
A synthetic dataset containing 447 typefaces with only one font variation for each typeface, created for visual font recognition.
1 paper · 1 benchmark
The VGG Cell dataset (made up entirely of synthetic images) is the main public benchmark used to compare cell counting techniques.
1 paper · 0 benchmarks
A high-resolution version of VGGFace2 for academic face editing purposes.
1 paper · 0 benchmarks
The VIA dataset is a dataset for aiding the visually impaired.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.