Home › Datasets › modality › Images
Images datasets
archive 2025-07-28
3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 17 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Images datasets 769–816 of 3,239
Kennedy Space Center is a dataset for the classification of wetland vegetation at the Kennedy Space Center, Florida using hyperspectral imagery.
18 papers · 1 benchmark
The LIVE Public-Domain Subjective Image Quality Database is a resource developed by the Laboratory for Image and Video Engineering at the University of Texas at Austin.
18 papers · 7 benchmarks
LIVECell (Label-free In Vitro image Examples of Cells)
The LIVECell (Label-free In Vitro image Examples of Cells) dataset is a large-scale microscopic image dataset for instance-segmentation of individual cells in 2D cell cultures.
18 papers · 1 benchmark
The MMD (MultiModal Dialogs) dataset is a dataset for multimodal domain-aware conversations.
18 papers · 0 benchmarks
The dataset was created for video quality assessment problem.
18 papers · 2 benchmarks
PASTIS (Panoptic Segmentation of satellite image TImes Series)
PASTIS is a benchmark dataset for panoptic and semantic segmentation of agricultural parcels from satellite image time series.
18 papers · 2 benchmarks
The Replay-Mobile Database for face spoofing consists of 1190 video clips of photo and video attack attempts to 40 clients, under different lighting conditions.
18 papers · 0 benchmarks
SCICAP is a large-scale image captioning dataset that contains real-world scientific figures and captions.
18 papers · 1 benchmark
SeaDronesSee (SeaDronesSee: A Maritime Benchmark for Detecting Humans in Open Water)
SeaDronesSee is a large-scale data set aimed at helping develop systems for Search and Rescue (SAR) using Unmanned Aerial Vehicles (UAVs) in maritime scenarios.
18 papers · 3 benchmarks
TUM-VIE (TUM Stereo Visual-Inertial Event Dataset)
TUM-VIE is an event camera dataset for developing 3D perception and navigation algorithms.
18 papers · 0 benchmarks
VAST (VAried Stance Topics)
VAST consists of a large range of topics covering broad themes, such as politics (e.g., ‘a Palestinian state’), education (e.g., ‘charter schools’), and public health (e.g., ‘childhood vaccination’).
18 papers · 1 benchmark
Violin (VIdeO-and-Language INference)
Video-and-Language Inference is the task of joint multimodal understanding of video and text.
18 papers · 0 benchmarks
WTW (Wired Table in the Wild)
WTW (Wired Table in the Wild) is a large-scale dataset which includes well-annotated structure parsing of multiple style tables in several scenes like the photo, scanning files, web pages.
18 papers · 1 benchmark
ZInd (Zillow Indoor Dataset)
The Zillow Indoor Dataset (ZInD) provides extensive visual data that covers a real world distribution of unfurnished residential homes.
18 papers · 1 benchmark
bFFHQ (Gender-biased FFHQ dataset)
Gender-biased FFHQ dataset (bFFHQ) has age as a target label and gender as a correlated bias, and the images are from the FFHQ dataset.
18 papers · 1 benchmark
ApolloCar3DT is a dataset that contains 5,277 driving images and over 60K car instances, where each car is fitted with an industry-grade 3D CAD model with absolute model size and semantically labelled keypoints.
17 papers · 14 benchmarks
CASIA-HWDB is a dataset for handwritten Chinese character recognition.
17 papers · 0 benchmarks
CDTB (Color-and-Depth Tracking)
Source: https://www.vicos.si/Projects/CDTB 4.2 State-of-the-art Comparison A TH CTB (color-and-depth visual object tracking) dataset is recorded by several passive and active RGB-D setups and contains indoor as well as outdoor sequences…
17 papers · 0 benchmarks
CLEVR-Ref+ is a synthetic diagnostic dataset for referring expression comprehension.
17 papers · 1 benchmark
COCO-Noisy (Microsoft Common Objects in Context with 20% of Noisy Correspondence and 1K test data)
This dataset is based on MS COCO that have 20% of data randomly shuffled to simulate noisy correspondence.
17 papers · 1 benchmark
The D-HAZY dataset is generated from NYU depth indoor image collection.
17 papers · 0 benchmarks
FMD (Fluorescence Microscopy Denoising)
The Fluorescence Microscopy Denoising (FMD) dataset is dedicated to Poisson-Gaussian denoising.
17 papers · 2 benchmarks
The Kumar dataset contains 30 1,000×1,000 image tiles from seven organs (6 breast, 6 liver, 6 kidney, 6 prostate, 2 bladder, 2 colon and 2 stomach) of The Cancer Genome Atlas (TCGA) database acquired at 40× magnification.
17 papers · 1 benchmark
This dataset contains 2100+ high resolution indoor panoramas, captured using a Canon 5D Mark III and a robotic panoramic tripod head.
17 papers · 0 benchmarks
MLRSNet is a a multi-label high spatial resolution remote sensing dataset for semantic scene understanding.
17 papers · 2 benchmarks
MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities
17 papers · 1 benchmark
Memorability dataset with 10000 3-second videos.
17 papers · 0 benchmarks
The dataset for this challenge was obtained by carefully annotating tissue images of several patients with tumors of different organs and who were diagnosed at multiple hospitals.
17 papers · 2 benchmarks
NINCO (No ImageNet Class Objects)
The NINCO (No ImageNet Class Objects) dataset is introduced in the ICML 2023 paper In or Out?
17 papers · 0 benchmarks
The New College Data is a freely available dataset collected from a robot completing several loops outdoors around the New College campus in Oxford.
17 papers · 0 benchmarks
Recent accelerations in multi-modal applications have been made possible with the plethora of image and text data available online.
17 papers · 0 benchmarks
RPC (Retail Product Checkout)
RPC is a large-scale retail product checkout dataset and collects 200 retail SKUs.
17 papers · 0 benchmarks
SECOND (SEmantic Change detectiON Dataset)
SECOND is a well-annotated semantic change detection dataset.
17 papers · 1 benchmark
SceneNet is a dataset of labelled synthetic indoor scenes.
17 papers · 0 benchmarks
The shiny folder contains 8 scenes with challenging view-dependent effects used in our paper.
17 papers · 0 benchmarks
TID2013 is a dataset for image quality assessment that contains 25 reference images and 3000 distorted images (25 reference images x 24 types of distortions x 5 levels of distortions).
17 papers · 1 benchmark
VQA-E is a dataset for Visual Question Answering with Explanation, where the models are required to generate and explanation with the predicted answer.
17 papers · 0 benchmarks
ViViD++ (Vision for Visibility Dataset)
A dataset capturing diverse visual data formats that target varying luminance conditions, and was recorded from alternative vision sensors, by handheld or mounted on a car, repeatedly in the same space but in different conditions.
17 papers · 0 benchmarks
The York Urban Line Segment Database is a compilation of 102 images (45 indoor, 57 outdoor) of urban environments consisting mostly of scenes from the campus of York University and downtown Toronto, Canada.
17 papers · 2 benchmarks
AgeDB contains 16, 488 images of various famous people, such as actors/actresses, writers, scientists, politicians, etc.
16 papers · 3 benchmarks
BG-20k (Background Dataset - 20k)
BG-20k contains 20,000 high-resolution background images excluded salient objects, which can be used to help generate high quality synthetic data.
16 papers · 0 benchmarks
The CLEVR-Hans data set is a novel confounded visual scene data set, which captures complex compositions of different objects.
16 papers · 0 benchmarks
Casual Conversations dataset is designed to help researchers evaluate their computer vision and audio models for accuracy across a diverse set of age, genders, apparent skin tones and ambient lighting conditions.
16 papers · 0 benchmarks
CholecT45 is a subset of CholecT50 consisting of 45 videos from the Cholec80 dataset.
16 papers · 1 benchmark
CryoNuSeg is a fully annotated FS-derived cryosectioned and H&E-stained nuclei instance segmentation dataset.
16 papers · 0 benchmarks
DADA-2000 is a large-scale benchmark with 2000 video sequences (named as DADA-2000) is contributed with laborious annotation for driver attention (fixation, saccade, focusing time), accident objects/intervals, as well as the accident…
16 papers · 0 benchmarks
This dataset, based on Flickr30K, is introduced in Learning with Noisy Correspondence for Cross-modal Matching.
16 papers · 1 benchmark
GRIT (General Robust Image Task Benchmark)
The General Robust Image Task (GRIT) Benchmark is an evaluation-only benchmark for evaluating the performance and robustness of vision systems across multiple image prediction tasks, concepts, and data sources.
16 papers · 5 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.