Home › Datasets › modality › Images

Images datasets

archive 2025-07-28

3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 4 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Images datasets 145–192 of 3,239

YouTube-VOS 2018 (Youtube Video Object Segmentation)
Youtube-VOS is a Video Object Segmentation dataset that contains 4,453 videos - 3,471 for training, 474 for validation, and 508 for testing.
203 papers · 10 benchmarks
LSP (Leeds Sports Pose)
The Leeds Sports Pose (LSP) dataset is widely used as the benchmark for human pose estimation.
202 papers · 1 benchmark
The HELEN dataset is composed of 2330 face images of 400×400 pixels with labeled facial components generated through manually-annotated contours along eyes, eyebrows, nose, lips and jawline.
201 papers · 1 benchmark
MPI (Max Planck Institute) Sintel is a dataset for optical flow evaluation that has 1064 synthesized stereo images and ground truth data for disparity.
198 papers · 5 benchmarks
PASCAL VOC (PASCAL Visual Object Classes Challenge)
The PASCAL Visual Object Classes (VOC) 2012 dataset contains 20 object categories including vehicles, household, animals, and other: aeroplane, bicycle, boat, bus, car, motorbike, train, bottle, chair, dining table, potted plant, sofa,…
198 papers · 18 benchmarks
CINIC-10 is a dataset for image classification.
197 papers · 3 benchmarks
The Moving MNIST dataset contains 10,000 video sequences, each consisting of 20 frames.
194 papers · 1 benchmark
MNIST-M is created by combining MNIST digits with the patches randomly extracted from color photos of BSDS500 as their background.
193 papers · 1 benchmark
The MOTChallenge datasets are designed for the task of multiple object tracking.
192 papers · 0 benchmarks
CMU Multimodal Opinion Sentiment and Emotion Intensity (CMU-MOSEI) is the largest dataset of sentence-level sentiment analysis and emotion recognition in online videos.
190 papers · 3 benchmarks
Celeb-DF is a large-scale challenging dataset for deepfake forensics.
187 papers · 0 benchmarks
RESISC45 dataset is a dataset for Remote Sensing Image Scene Classification (RESISC).
187 papers · 3 benchmarks
SUNCG is a large-scale dataset of synthetic 3D scenes with dense volumetric annotations.
186 papers · 0 benchmarks
The Extended Yale B database contains 2414 frontal-face images with size 192×168 over 38 subjects and about 64 images per subject.
185 papers · 1 benchmark
SA-1B consists of 11M diverse, high resolution, licensed, and privacy protecting images and 1.1B high-quality segmentation masks.
183 papers · 1 benchmark
OTB-2015, also referred as Visual Tracker Benchmark, is a visual tracking dataset.
182 papers · 1 benchmark
MORPH is a facial age estimation dataset, which contains 55,134 facial images of 13,617 subjects ranging from 16 to 77 years old.
180 papers · 8 benchmarks
FUNSD (Form Understanding in Noisy Scanned Documents)
Form Understanding in Noisy Scanned Documents (FUNSD) comprises 199 real, fully annotated, scanned forms.
179 papers · 3 benchmarks
VCR (Visual Commonsense Reasoning)
Visual Commonsense Reasoning (VCR) is a large-scale dataset for cognition-level visual understanding.
179 papers · 13 benchmarks
The WebVision dataset is designed to facilitate the research on learning visual representation from noisy web data.
179 papers · 4 benchmarks
LabelMe database is a large collection of images with ground truth labels for object detection and recognition.
178 papers · 1 benchmark
The Hateful Memes data set is a multimodal dataset for hateful meme detection (image + text) that contains 10,000+ new multimodal examples created by Facebook AI.
177 papers · 3 benchmarks
PASCAL-5i is a dataset used to evaluate few-shot segmentation.
177 papers · 1 benchmark
The UCF-QNRF dataset is a crowd counting dataset and it contains large diversity both in scenes, as well as in background types.
176 papers · 1 benchmark
R2R (Room-to-Room)
R2R is a dataset for visually-grounded natural language navigation in real buildings.
174 papers · 2 benchmarks
CAMELYON16 (Cancer Metastases in Lymph Nodes Challenge 2016)
The dataset consists of 400 whole-slide images (WSIs) of lymph node sections stained with hematoxylin and eosin (H&E), collected from two medical centers in the Netherlands.
172 papers · 1 benchmark
RAF-DB (Real-world Affective Faces)
The Real-world Affective Faces Database (RAF-DB) is a dataset for facial expression.
172 papers · 3 benchmarks
The UCY dataset consist of real pedestrian trajectories with rich multi-human interaction scenarios captured at 2.5 Hz (Δt=0.4s).
170 papers · 1 benchmark
LAION-400M is a dataset with CLIP-filtered 400 million image-text pairs, their CLIP embeddings and kNN indices that allow efficient similarity search.
169 papers · 1 benchmark
FER2013 (Facial Expression Recognition 2013 Dataset)
Fer2013 contains approximately 30,000 facial RGB images of different expressions with size restricted to 48×48, and the main labels of it can be divided into 7 types: 0=Angry, 1=Disgust, 2=Fear, 3=Happy, 4=Sad, 5=Surprise, 6=Neutral.
168 papers · 5 benchmarks
COD10K (Camouflaged/Concealed Object Detection)
Sensory ecologists have found that this s background matching camouflage strategy works by deceiving the visual perceptual system of the observer.
166 papers · 2 benchmarks
FDDB (Face Detection Dataset and Benchmark)
The Face Detection Dataset and Benchmark (FDDB) dataset is a collection of labeled faces from Faces in the Wild dataset.
165 papers · 1 benchmark
CelebAMask-HQ is a large-scale face image dataset that has 30,000 high-resolution face images selected from the CelebA dataset by following CelebA-HQ.
164 papers · 5 benchmarks
The YCB-Video dataset is a large-scale video dataset for 6D object pose estimation.
164 papers · 5 benchmarks
IJB-B (IARPA Janus Benchmark-B)
The IJB-B dataset is a template-based face dataset that contains 1845 subjects with 11,754 images, 55,025 frames and 7,011 videos where a template consists of a varying number of still images and video frames from different sources.
163 papers · 5 benchmarks
CrowdHuman is a large and rich-annotated human detection dataset, which contains 15,000, 4,370 and 5,000 images collected from the Internet for training, validation and testing respectively.
161 papers · 2 benchmarks
Objects365 is a large-scale object detection dataset, Objects365, which has 365 object categories over 600K training images.
161 papers · 2 benchmarks
SALICON (Salicency in Context)
The SALIency in CONtext (SALICON) dataset contains 10,000 training images, 5,000 validation images and 5,000 test images for saliency prediction.
161 papers · 5 benchmarks
V-COCO (Verbs in COCO)
Verbs in COCO (V-COCO) is a dataset that builds off COCO for human-object interaction detection.
159 papers · 1 benchmark
VisDial (Visual Dialog)
Visual Dialog (VisDial) dataset contains human annotated questions based on images of MS COCO dataset.
159 papers · 6 benchmarks
IJB-A (IARPA Janus Benchmark A)
The IARPA Janus Benchmark A (IJB-A) database is developed with the aim to augment more challenges to the face recognition task by collecting facial images with a wide variations in pose, illumination, expression, resolution and occlusion.
156 papers · 2 benchmarks
Total-Text is a text detection dataset that consists of 1,555 images with a variety of text types including horizontal, multi-oriented, and curved text instances.
156 papers · 2 benchmarks
ObjectNet is a test set of images collected directly using crowd-sourcing.
155 papers · 4 benchmarks
SID (See-in-the-Dark)
The See-in-the-Dark (SID) dataset contains 5094 raw short-exposure images, each with a corresponding long-exposure reference image.
155 papers · 3 benchmarks
A-OKVQA is crowdsourced visual question answering dataset composed of a diverse set of about 25K questions requiring a broad base of commonsense and world knowledge to answer.
154 papers · 1 benchmark
AFLW (Annotated Facial Landmarks in the Wild)
The Annotated Facial Landmarks in the Wild (AFLW) is a large-scale collection of annotated face images gathered from Flickr, exhibiting a large variety in appearance (e.g., pose, expression, ethnicity, age, gender) as well as general…
154 papers · 11 benchmarks
AFW (Annotated Faces in the Wild)
AFW (Annotated Faces in the Wild) is a face detection dataset that contains 205 images with 468 faces.
154 papers · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.