Home › Datasets › modality › Images
Images datasets
archive 2025-07-28
3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 45 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Images datasets 2113–2160 of 3,239
VQDv1 (Visual Query Detection v1)
In Visual Query Detection (VQD), a system is given a query (prompt) natural language and an image, and then the system must produce 0 - N boxes that satisfy that query.
2 papers · 0 benchmarks
VTC (Videos, Titles and Comments)
VTC is a large-scale multimodal dataset containing video-caption pairs (~300k) alongside comments that can be used for multimodal representation learning.
2 papers · 0 benchmarks
A dataset for Visual Voice Activity Detection extracted from the LRS3 dataset.
2 papers · 0 benchmarks
Large-scale benchmark dataset of full-field digital mammography, called VinDr-Mammo, which consists of 5,000 four-view exams with breast-level assessment and finding annotations.
2 papers · 0 benchmarks
The Vistas-NP dataset is an out-of-distribution detection dataset based on the Mapillary Vistas dataset.
2 papers · 0 benchmarks
A large-scale multi-view RGBD visual affordance learning dataset, a benchmark of 47210 RGBD images from 37 object categories, annotated with 15 visual affordance categories and 35 cluttered/complex scenes with different objects and…
2 papers · 0 benchmarks
Visual Beliefs is a dataset of abstract scenes to study visual beliefs.
2 papers · 0 benchmarks
A large-scale dataset that links the assessment of image quality issues to two practical vision tasks: image captioning and visual question answering.
2 papers · 0 benchmarks
WeatherKITTI is currently the most realistic all-weather simulated enhancement of the KITTI dataset.
2 papers · 0 benchmarks
WikiTableSet is a large publicly available image-based table recognition dataset in three languages built from Wikipedia.
2 papers · 0 benchmarks
ccHarmony is a color checker (cc) based image harmonization dataset.
2 papers · 0 benchmarks
dacl10k (dacl10k: Dataset for Semantic Bridge Damage Segmentation)
dacl10k stands for damage classification 10k images and is a multi-label semantic segmentation dataset for 19 classes (13 damages and 6 objects) present on bridges.
2 papers · 0 benchmarks
This dataset comprises 1344 expert annotated images of muscle-tendon junctions recorded with 3 ultrasound imaging systems (Aixplorer V6, Esaote MyLab60, Telemed ArtUs), on 2 muscles (Lateral Gastrocnemius, Medial Gastrocnemius), and 2…
2 papers · 0 benchmarks
This dataset contains pre and post destruction images and also segmentation labels for test images.
2 papers · 0 benchmarks
ePillID is a benchmark for developing and evaluating computer vision models for pill identification.
2 papers · 0 benchmarks
iBugMask is an in-the-wild face parsing dataset that contains 1,000 challenging face images and manually annotated labels for 11 semantic classes: background, facial skin, left/right brow, left/right eye, nose, upper/lower lip, inner…
2 papers · 1 benchmark
iRodent (iRodent Animal Pose Estimation)
Description: The "iRodent" dataset contains rodent species observations obtained using the iNaturalist API, with a focus on Suborder Myomorpha (Taxon ID: 16).
2 papers · 1 benchmark
mini-ImageNet was proposed by Matching networks for one-shot learning for few-shot learning evaluation, in an attempt to have a dataset like ImageNet while requiring fewer resources.
2 papers · 1 benchmark
robo-vln (Robotics Vision-and-Language Navigation)
The Robo-VLN dataset is a continuous control formulation of the VLN-CE dataset by Krantz et al ported over from Room-to-Room (R2R) dataset created by Anderson et al.
2 papers · 1 benchmark
Unsustainable fishing practices worldwide pose a major threat to marine resources and ecosystems.
2 papers · 1 benchmark
Description: 23 Pairs of Identical Twins Face Image Data.
1 paper · 0 benchmarks
This paper constructs 7-digit product Supply-Use Tables (SUTs) and symmetric Input-Output Tables (IOTs) for the Indian economy using microdata from the Annual Survey of Industries (ASI) for the period 2016-2021.
1 paper · 0 benchmarks
A Ball-Collision Dataset (ABCD) serves as a comprehensive benchmark for investigating the interaction dynamics of moving objects within 3D environments.
1 paper · 1 benchmark
The dataset contains aerial agricultural images of a potato field with manual labels of healthy and stressed plant regions.
1 paper · 1 benchmark
This dataset contains 9 different seafood types collected from a supermarket in Izmir, Turkey for a university-industry collaboration project at Izmir University of Economics, and this work was published in ASYU 2020.
1 paper · 0 benchmarks
A View From Somewhere (AVFS)—a dataset of 638,180 face similarity judgments over 4,921 faces.
1 paper · 0 benchmarks
This dataset contains a collection of 131 X-ray CT scans of pieces of modeling clay (Play-Doh) with various numbers of stones inserted, retrieved in the FleX-ray lab at CWI.
1 paper · 0 benchmarks
This dataset contains a collection of 235800 X-ray projections of 131 pieces of modeling clay (Play-Doh) with various numbers of stones inserted.
1 paper · 0 benchmarks
ACCT Data Repository (ACCT is a fast and accessible automatic cell counting tool using machine learning for 2D image segmentation)
This dataset is a collection of fluorescent images from mice in order to test an automatic cell counting tool that we developed.
1 paper · 0 benchmarks
ACFR Orchard Fruit Dataset is an agricultural dataset containing images and annotations for different fruits, collected at different farms across Australia.
1 paper · 0 benchmarks
We introduce a new AI-ready computational pathology dataset containing restained and co-registered digitized images from eight head-and-neck squamous cell carcinoma patients.
1 paper · 0 benchmarks
AIDERV2 (Aerial Image Dataset for Emergency Response Applications (version 2))
The dataset contains aerial images containing three commonly occurring natural disasters earthquake/collapsed buildings, flood, wildfire/fire, and a normal class; do not reflect any disaster.
1 paper · 1 benchmark
AIROGS (Rotterdam EyePACS AIROGS)
The Rotterdam EyePACS AIROGS dataset (in full, so including train and test) contains 113,893 color fundus images from 60,357 subjects and approximately 500 different sites with a heterogeneous ethnicity.
1 paper · 0 benchmarks
ALLO (Anomaly Localization in Lunar Orbit)
ALLO is an anomaly detection and localization dataset for space stations in lunar orbit.
1 paper · 0 benchmarks
ANUBIS (Skeleton-Based Action Recognition Dataset)
ANUBIS is a large-scale human skeleton dataset containing 80 actions.
1 paper · 0 benchmarks
AODRaw (Adverse condition Object Detection with RAW images)
We introduce the AODRaw dataset, which offers 7,785 high-resolution real RAW images with 135,601 annotated instances spanning 62 categories, capturing a broad range of indoor and outdoor scenes under 9 distinct light and weather conditions.
1 paper · 1 benchmark
A database of 56 high quality fabric material measurements, provided as carefully calibrated rectified HDR images, together with SVBRDF fits.
1 paper · 0 benchmarks
ARC Ukiyo-e Faces is a large-scale (>10k paintings, >20k faces) Ukiyo-e dataset with coherent semantic labels and geometric annotations through augmenting and organizing existing datasets with automatic detection.
1 paper · 0 benchmarks
The ARC-100 dataset was collected as part of a prototype retail checkout system titled ARC (Automatic Retail Checkout).
1 paper · 0 benchmarks
ARIA (Automated Retinal Image Analysis (ARIA) Data Set)
This data set was collected in 2004 to 2006 in the United Kingdom.
1 paper · 0 benchmarks
ARKitTrack is a new RGB-D tracking dataset for both static and dynamic scenes captured by consumer-grade LiDAR scanners equipped on Apple's iPhone and iPad.
1 paper · 0 benchmarks
ARVSU (Addressee Recognition in Visual Scenes with Utterances)
ARVSU contains a vast body of image variations in visual scenes with an annotated utterance and a corresponding addressee for each scenario.
1 paper · 0 benchmarks
ASRD (Anime Style Recognition Dataset)
A well-labeled challenging dataset, to facilitate the research on style recognition on anime images by collecting images from 190 anime and cartoon works covering 93 years from 13 countries and regions, 2D and 3D work into consideration…
1 paper · 0 benchmarks
Multimodal object recognition is still an emerging field.
1 paper · 0 benchmarks
The existing multi-modality image fusion dataset lacks comprehensive coverage of adverse weather scenarios.
1 paper · 0 benchmarks
Accidental Turntables contains a challenging set of 41,212 images of cars in cluttered backgrounds, motion blur and illumination changes that serves as a benchmark for 3D pose estimation.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.