Home › Datasets › modality › Images
Images datasets
archive 2025-07-28
3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 27 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Images datasets 1249–1296 of 3,239
DocCVQA (Document Collection Visual Question Answering)
DocCVQA is a Document Visual Question Answering dataset, where the questions are posed over a whole collection of 14,362 scanned documents.
6 papers · 0 benchmarks
DroneSURF (DroneSURF: Benchmark Dataset for Drone-based Face Recognition)
Drone Surveillance of Faces, is a large-scale drone dataset intended to facilitate research for face recognition using drones.
6 papers · 1 benchmark
EarthVQA (A multi-modal multi-task VQA dataset for remote sensing)
Earth vision research typically focuses on extracting geospatial object locations and categories but neglects the exploration of relations between objects and comprehensive reasoning.
6 papers · 1 benchmark
The EuroCity Persons dataset provides a large number of highly diverse, accurate and detailed annotations of pedestrians, cyclists and other riders in urban traffic scenes.
6 papers · 0 benchmarks
Falling Things (FAT) is a dataset for advancing the state-of-the-art in object detection and 3D pose estimation in the context of robotics.
6 papers · 0 benchmarks
FFHQ-Aging is a Dataset of human faces designed for benchmarking age transformation algorithms as well as many other possible vision tasks.
6 papers · 0 benchmarks
The Food-101N dataset is introduced in "CleanNet: Transfer Learning for Scalable Image Training with Label Noise (CVPR'18).
6 papers · 1 benchmark
The Freiburg Forest dataset was collected using a Viona autonomous mobile robot platform equipped with cameras for capturing multi-spectral and multi-modal images.
6 papers · 2 benchmarks
GBCU (Gallbladder Cancer Ultrasound Dataset)
GBCU is the first public dataset for Gallbladder Cancer identification from Ultrasound images.
6 papers · 1 benchmark
Geo-Diverse Visual Commonsense Reasoning (GD-VCR) is a new dataset to test vision-and-language models' ability to understand cultural and geo-location-specific commonsense.
6 papers · 1 benchmark
This dataset is a new benchmark, grounded in real-world usages is developed to support more authentic and comprehensive evaluation of image editing models.
6 papers · 2 benchmarks
GFF (Global Flood Forecasting)
Floods are among the most common and devastating natural hazards, imposing immense costs on our society and economy due to their disastrous consequences.
6 papers · 1 benchmark
HRS-Bench (Holistic, Reliable, and Scalable Benchmark)
HRS-Bench is a concrete evaluation benchmark for T2I models that is Holistic, Reliable, and Scalable.
6 papers · 0 benchmarks
iCubWorld datasets are collections of images recording the visual experience of iCub while observing objects in its typical environment, a laboratory or an office.
6 papers · 0 benchmarks
IIIT-AR-13K is created by manually annotating the bounding boxes of graphical or page objects in publicly available annual reports.
6 papers · 0 benchmarks
IIIT-ILST is a dataset and benchmark for scene text recognition for three Indic scripts - Devanagari, Telugu and Malayalam.
6 papers · 0 benchmarks
The ISIC 2018 dataset was published by the International Skin Imaging Collaboration (ISIC) as a large-scale dataset of dermoscopy images.
6 papers · 0 benchmarks
Dataset of over 6 million GPS-tagged images from Flickr.
6 papers · 1 benchmark
OpenImage-O is built for the ID dataset ImageNet-1k.
6 papers · 1 benchmark
IndustReal (IndustReal Dataset of Egocentric Videos for Procedure Understanding)
IndustReal is an ego-centric, multi-modal dataset where 27 participants are challenged to perform assembly and maintenance procedures on a construction-toy car.
6 papers · 3 benchmarks
LIS (low-light instance segmentation)
To reveal and systematically investigate the effectiveness of the proposed method in the real world, a real low-light image dataset for instance segmentation is necessary and urgently needed.
6 papers · 0 benchmarks
The unsupervised Labeled Lane MArkerS dataset (LLAMAS) is a dataset for lane detection and segmentation.
6 papers · 1 benchmark
LSMI (Large Scale Multi-Illuminant dataet)
Large Scale Multi-Illuminant (LSMI) Dataset for Developing White Balance Algorithm under Mixed Illumination (ICCV 2021) Change Log LSMI Dataset Version : 1.1 1.0 : LSMI dataset released.
6 papers · 0 benchmarks
Lytro Illum is a new light field dataset using a Lytro Illum camera.
6 papers · 0 benchmarks
MARIDA (Marine Debris Archive)
MARIDA (Marine Debris Archive) is the first dataset based on the multispectral Sentinel-2 (S2) satellite data, which distinguishes Marine Debris from various marine features that co-exist, including Sargassum macroalgae, Ships, Natural…
6 papers · 1 benchmark
MCXFACE (Multi-Channel Heterogeneous Face Recognition dataset)
MCXFace is a heterogeneous face recognition dataset consisting of multi-channel image samples for 51 subjects.
6 papers · 0 benchmarks
Contains video clips shot with modern high-resolution mobile cameras, with strong projective distortions and with low lighting conditions.
6 papers · 0 benchmarks
MIT Traffic is a dataset for research on activity analysis and crowded scenes.
6 papers · 0 benchmarks
MLe2 is a dataset for the evaluation of scene text end-to-end reading systems and all intermediate stages such as text detection, script identification and text recognition.
6 papers · 0 benchmarks
MatterportLayout extends the Matterport3D dataset with general Manhattan layout annotations.
6 papers · 0 benchmarks
Modern Office-31 is a refurbished version of the commonly used Office-31 dataset.
6 papers · 0 benchmarks
NCD (Natural-Color Dataset)
The Natural-Color Dataset (NCD) is an image colorization dataset where images are true to their colors.
6 papers · 0 benchmarks
OOD-CV (Out Of Distribution Generalization in Computer Vision)
Enhancing the robustness of vision algorithms in real-world scenarios is challenging.
6 papers · 1 benchmark
OmniFlow is a synthetic omnidirectional human optical flow dataset.
6 papers · 0 benchmarks
Open Images is a computer vision dataset covering ~9 million images with labels spanning thousands of object categories.
6 papers · 0 benchmarks
The PieAPP dataset is a large-scale dataset used for training and testing perceptually-consistent image-error prediction algorithms.
6 papers · 0 benchmarks
The RAD-ChestCT dataset is a large medical imaging dataset developed by Duke MD/PhD Rachel Draelos during her Computer Science PhD supervised by Lawrence Carin.
6 papers · 0 benchmarks
This dataset arises from the READ project (Horizon 2020).
6 papers · 1 benchmark
Rad-ReStruct is a fine-grained structured reporting dataset for Chest X-Ray images.
6 papers · 0 benchmarks
RealCQA Scientific Chart Question Answering as a Test-bed for First-Order Logic check on huggingface : https://huggingface.co/datasets/sal4ahm/RealCQA
6 papers · 1 benchmark
The Relative Size dataset contains 486 object pairs between 41 physical objects.
6 papers · 0 benchmarks
RxRx1 is a biological dataset designed specifically for the systematic study of batch effect correction methods.
6 papers · 1 benchmark
SMART-101 (Simple Multimodal Algorithmic Reasoning Task Dataset)
Recent times have witnessed an increasing number of applications of deep neural networks towards solving tasks that require superior cognitive abilities, e.g., playing Go, generating art, ChatGPT, etc.
6 papers · 0 benchmarks
A multimodal dataset for sentiment analysis on internet memes.
6 papers · 0 benchmarks
ShapeTalk contains over half a million discriminative utterances produced by contrasting the shapes of common 3D objects for a variety of object classes and degrees of similarity.
6 papers · 0 benchmarks
ShipSG (Ship Segmentation and Georeferencing Dataset)
The ShipSG dataset is the first public dataset of its kind for ship segmentation and georeferencing.
6 papers · 0 benchmarks
TRECVID is a yearly set of competitions centered on video retrieval and indexing, hosting a variety of video data sets.
6 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.