Home › Datasets › modality › Images
Images datasets
archive 2025-07-28
3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 13 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Images datasets 577–624 of 3,239
We build a large-scale, comprehensive, and high-quality synthetic dataset for city-scale neural rendering researches.
33 papers · 0 benchmarks
SEVIR (Storm EVent ImagRy)
SEVIR is an annotated, curated and spatio-temporally aligned dataset containing over 10,000 weather events that each consist of 384 km x 384 km image sequences spanning 4 hours of time.
33 papers · 2 benchmarks
WMCA (Wide Multi Channel Presentation Attack)
The Wide Multi Channel Presentation Attack (WMCA) database consists of 1941 short video recordings of both bonafide and presentation attacks from 72 different identities.
33 papers · 1 benchmark
🤖 Robo3D - The nuScenes-C Benchmark nuScenes-C is an evaluation benchmark heading toward robust and reliable 3D perception in autonomous driving.
33 papers · 2 benchmarks
CholecT50 is a dataset of endoscopic videos of laparoscopic cholecystectomy surgery introduced to enable research on fine-grained action recognition in laparoscopic surgery.
32 papers · 5 benchmarks
Dress Code is a new dataset for image-based virtual try-on composed of image pairs coming from different catalogs of YOOX NET-A-PORTER.
32 papers · 1 benchmark
EMDB contains in-the-wild videos of human activity recorded with a hand-held iPhone.
32 papers · 2 benchmarks
MS-CXR (Making the Most of Text Semantics to Improve Biomedical Vision-Language Processing)
The MS-CXR dataset provides 1162 image–sentence pairs of bounding boxes and corresponding phrases, collected across eight different cardiopulmonary radiological findings, with an approximately equal number of pairs for each finding.
32 papers · 0 benchmarks
PACO (Parts and Attributes of Common Objects)
Parts and Attributes of Common Objects (PACO) is a detection dataset that goes beyond traditional object boxes and masks and provides richer annotations such as part masks and attributes.
32 papers · 0 benchmarks
PIRM (Perceptual Image Restoration and Manipulation)
The PIRM dataset consists of 200 images, which are divided into two equal sets for validation and testing.
32 papers · 1 benchmark
We introduce our new dataset, Spaces, to provide a more challenging shared dataset for future view synthesis research.
32 papers · 0 benchmarks
TinyFace is a large scale face recognition benchmark to facilitate the investigation of natively LRFR (Low Resolution Face Recognition) at large scales (large gallery population sizes) in deep learning.
32 papers · 0 benchmarks
UCF-CC-50 is a dataset for crowd counting and consists of images of extremely dense crowds.
32 papers · 0 benchmarks
ArtEmis is a large-scale dataset aimed at providing a detailed understanding of the interplay between visual content, its emotional effect, and explanations for the latter in language.
31 papers · 0 benchmarks
The DQN Replay Dataset was collected as follows: We first train a [DQN][naturedqn] agent, on all 60 [Atari 2600 games][ale] with [sticky actions][stochasticale] enabled for 200 million frames (standard protocol) and save all of the…
31 papers · 0 benchmarks
IGLUE (Image-Grounded Language Understanding Evaluation)
The Image-Grounded Language Understanding Evaluation (IGLUE) benchmark brings together—by both aggregating pre-existing datasets and creating new ones—visual question answering, cross-modal retrieval, grounded reasoning, and grounded…
31 papers · 0 benchmarks
The Image Paragraph Captioning dataset allows researchers to benchmark their progress in generating paragraphs that tell a story about an image.
31 papers · 1 benchmark
The IMAGE-CHAT dataset is a large collection of (image, style trait for speaker A, style trait for speaker B, dialogue between A & B) tuples that we collected using crowd-workers, Each dialogue consists of consecutive turns by speaker A…
31 papers · 2 benchmarks
KVQA (Knowledge-aware VQA)
It contains manually verified 183K question-answer pairs about more than 18K persons and 24K images.
31 papers · 0 benchmarks
LC25000 (Lung And Colon Histopathological Image Dataset)
The LC25000 dataset contains 25,000 color images with 5 classes of 5,000 images each.
31 papers · 0 benchmarks
Multimodal C4 (MMC4) is an augmentation of the popular text-only c4 corpus with images interleaved.
31 papers · 0 benchmarks
MSRC-12 (MSRC-12 Kinect Gesture Dataset)
The Microsoft Research Cambridge-12 Kinect gesture data set consists of sequences of human movements, represented as body-part locations, and the associated gesture to be recognized by the system.
31 papers · 2 benchmarks
OpenEDS (Open Eye Dataset) is a large scale data set of eye-images captured using a virtual-reality (VR) head mounted display mounted with two synchronized eyefacing cameras at a frame rate of 200 Hz under controlled illumination.
31 papers · 1 benchmark
PhraseCut is a dataset consisting of 77,262 images and 345,486 phrase-region pairs.
31 papers · 1 benchmark
SOC (Salient Objects in Clutter)
SOC (Salient Objects in Clutter) is a dataset for Salient Object Detection (SOD).
31 papers · 1 benchmark
When glancing at a magazine, or browsing the Internet, we are continuously being exposed to photographs.
31 papers · 4 benchmarks
BRACS (BReAst Carcinoma Subtyping)
BReAst Carcinoma Subtyping (BRACS) dataset, a large cohort of annotated Hematoxylin & Eosin (H&E)-stained images to facilitate the characterization of breast lesions.
30 papers · 0 benchmarks
DeepFashion2 is a versatile benchmark of four tasks including clothes detection, pose estimation, segmentation, and retrieval.
30 papers · 0 benchmarks
The General-100 dataset is a dataset for image super-resolution.
30 papers · 0 benchmarks
InteriorNet is a RGB-D for large scale interior scene understanding and mapping.
30 papers · 0 benchmarks
🤖 Robo3D - The KITTI-C Benchmark KITTI-C is an evaluation benchmark heading toward robust and reliable 3D object detection in autonomous driving.
30 papers · 1 benchmark
ShoeV2 is a dataset of 2,000 photos and 6648 sketches of shoes.
30 papers · 0 benchmarks
AVD (Active Vision Dataset)
AVD focuses on simulating robotic vision tasks in everyday indoor environments using real imagery.
29 papers · 1 benchmark
Comic2k is a dataset used for cross-domain object detection which contains 2k comic images with image and instance-level annotations.
29 papers · 4 benchmarks
MaRVL (Multicultural Reasoning over Vision and Language)
Multicultural Reasoning over Vision and Language (MaRVL) is a dataset based on an ImageNet-style hierarchy representative of many languages and cultures (Indonesian, Mandarin Chinese, Swahili, Tamil, and Turkish).
29 papers · 1 benchmark
The NVGesture dataset focuses on touchless driver controlling.
29 papers · 1 benchmark
Omni-Realm Benchmark (OmniBenchmark) is a diverse (21 semantic realm-wise datasets) and concise (realm-wise datasets have no concepts overlapping) benchmark for evaluating pre-trained model generalization across semantic…
29 papers · 1 benchmark
PartImageNet is a large, high-quality dataset with part segmentation annotations.
29 papers · 0 benchmarks
Spring (Spring: A High-Resolution High-Detail Dataset and Benchmark for Scene Flow, Optical Flow and Stereo)
Spring is a large, high-resolution and high-detail, computer-generated benchmark for scene flow, optical flow, and stereo.
29 papers · 3 benchmarks
Semi-Supervised Object Detection on COCO 10% labeled data
28 papers · 2 benchmarks
CelebA-Spoof is a large-scale face anti-spoofing dataset with the following properties: 1.
28 papers · 0 benchmarks
The Gaming 3D Dataset (G3D) focuses on real-time action recognition in a gaming scenario.
28 papers · 2 benchmarks
KITTI MOTS (KITTI Multi-Object Tracking and Segmentation (MOTS) Evaluation)
The Multi-Object and Segmentation (MOTS) benchmark [2] consists of 21 training sequences and 29 test sequences.
28 papers · 1 benchmark
MHIST (Minimalist Histopathology image analysis dataset)
The minimalist histopathology image analysis dataset (MHIST) is a binary classification dataset of 3,152 fixed-size images of colorectal polyps, each with a gold-standard label determined by the majority vote of seven board-certified…
28 papers · 1 benchmark
The MIT-Adobe FiveK dataset consists of 5,000 photographs taken with SLR cameras by a set of different photographers.
28 papers · 4 benchmarks
MVSEC (Multi Vehicle Stereo Event Camera)
The Multi Vehicle Stereo Event Camera (MVSEC) dataset is a collection of data designed for the development of novel 3D perception algorithms for event based cameras.
28 papers · 2 benchmarks
The MannequinChallenge Dataset (MQC) provides in-the-wild videos of people in static poses while a hand-held camera pans around the scene.
28 papers · 0 benchmarks
MedICaT is a dataset of medical images, captions, subfigure-subcaption annotations, and inline textual references.
28 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.