Home › Datasets › modality › Images

Images datasets

archive 2025-07-28

3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 6 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Images datasets 241–288 of 3,239

NCLT (North Campus Long-Term Vision and LiDAR)
The NCLT dataset is a large scale, long-term autonomy dataset for robotics research collected on the University of Michigan’s North Campus.
124 papers · 0 benchmarks
JFT-300M is an internal Google dataset used for training image classification models.
123 papers · 1 benchmark
PubLayNet is a dataset for document layout analysis by automatically matching the XML representations and the content of over 1 million PDF articles that are publicly available on PubMed Central.
123 papers · 1 benchmark
ETHD is a multi-view stereo benchmark / 3D reconstruction benchmark that covers a variety of indoor and outdoor scenes.
121 papers · 3 benchmarks
Permuted MNIST is an MNIST variant that consists of 70,000 images of handwritten digits from 0 to 9, where 60,000 images are used for training, and 10,000 images for test.
121 papers · 1 benchmark
GTEA (Georgia Tech Egocentric Activity)
The Georgia Tech Egocentric Activities (GTEA) dataset contains seven types of daily activities such as making sandwich, tea, or coffee.
120 papers · 2 benchmarks
Raindrop is a set of image pairs, where each pair contains exactly the same background scene, yet one is degraded by raindrops and the other one is free from raindrops.
120 papers · 1 benchmark
SHAPES (Swarm Heuristics based Adaptive and Penalized Estimation of Splines)
SHAPES is a dataset of synthetic images designed to benchmark systems for understanding of spatial and logical relations among multiple objects.
120 papers · 1 benchmark
The 'shape bias' dataset was introduced in Geirhos et al.
120 papers · 1 benchmark
AFLW2000-3D is a dataset of 2000 images that have been annotated with image-level 68-point 3D facial landmarks.
117 papers · 8 benchmarks
Kvasir (The Kvasir Dataset)
The KVASIR Dataset was released as part of the medical multimedia challenge presented by MediaEval.
117 papers · 1 benchmark
The 20BN-SOMETHING-SOMETHING dataset is a large collection of labeled video clips that show humans performing pre-defined basic actions with everyday objects.
117 papers · 3 benchmarks
VGGFace2 Dataset (Vggface2: A dataset for recognising faces across pose and age)
VGGFace2 is a large-scale face recognition dataset.
117 papers · 0 benchmarks
LLVIP (A Visible-infrared Paired Dataset for Low-light Vision)
Visible-infrared Paired Dataset for Low-light Vision 30976 images (15488 pairs) 24 dark scenes, 2 daytime scenes Support for image-to-image translation (visible to infrared, or infrared to visible), visible and infrared image fusion,…
116 papers · 6 benchmarks
PadChest is a labeled large-scale, high resolution chest x-ray dataset for the automated exploration of medical images along with their associated reports.
116 papers · 0 benchmarks
CSIQ (Categorical Subjective Image Quality)
The CSIQ database consists of 30 original images, each is distorted using six different types of distortions at four to five different levels of distortion.
115 papers · 1 benchmark
COFW (Caltech Occluded Faces in the Wild)
The Caltech Occluded Faces in the Wild (COFW) dataset is designed to present faces in real-world conditions.
114 papers · 5 benchmarks
The NLPR dataset for salient object detection consists of 1,000 image pairs captured by a standard Microsoft Kinect with a resolution of 640×480.
113 papers · 1 benchmark
Imagenet32 is a huge dataset made up of small images called the down-sampled version of Imagenet.
112 papers · 4 benchmarks
Visual7W is a large-scale visual question answering (QA) dataset, with object-level groundings and multimodal answers.
112 papers · 1 benchmark
The smallNORB dataset is a datset for 3D object recognition from shape.
112 papers · 1 benchmark
N-Caltech 101 (Neuromorphic-Caltech101)
The Neuromorphic-Caltech101 (N-Caltech101) dataset is a spiking version of the original frame-based Caltech101 dataset.
110 papers · 3 benchmarks
OTB2013 is the previous version of the current OTB2015 Visual Tracker Benchmark.
110 papers · 2 benchmarks
PCam (PatchCamelyon)
PatchCamelyon is an image classification dataset.
110 papers · 4 benchmarks
The Penn Action Dataset contains 2326 video sequences of 15 different actions and human joint annotations for each sequence.
110 papers · 4 benchmarks
Cambridge Landmarks, a large scale outdoor visual relocalisation dataset taken around Cambridge University.
109 papers · 0 benchmarks
The Chairs dataset contains rendered images of around 1000 different three-dimensional chair models.
109 papers · 1 benchmark
For many fundamental scene understanding tasks, it is difficult or impossible to obtain per-pixel ground truth labels from real images.
108 papers · 4 benchmarks
WFLW (Wider Facial Landmarks in the Wild)
The Wider Facial Landmarks in the Wild or WFLW database contains 10000 faces (7500 for training and 2500 for testing) with 98 annotated landmarks.
108 papers · 4 benchmarks
iSUN is a ground truth of gaze traces on images from the SUN dataset.
108 papers · 0 benchmarks
Rationale and objectives: Computer-aided detection and diagnosis (CAD) systems have been developed in the past two decades to assist radiologists in the detection and diagnosis of lesions seen on breast imaging exams, thus providing a…
107 papers · 3 benchmarks
The RVL-CDIP dataset consists of scanned document images belonging to 16 classes such as letter, form, email, resume, memo, etc.
107 papers · 2 benchmarks
SegTrack v2 is a video segmentation dataset with full pixel-level annotations on multiple objects at each frame within each video.
107 papers · 5 benchmarks
VIST (Visual Storytelling)
The Visual Storytelling Dataset (VIST) consists of 210,819 unique photos and 50,000 stories.
107 papers · 2 benchmarks
The Stylized-ImageNet dataset is created by removing local texture cues in ImageNet while retaining global shape information on natural images via AdaIN style transfer.
106 papers · 1 benchmark
The BP4D-Spontaneous dataset is a 3D video database of spontaneous facial expressions in a diverse group of young adults.
104 papers · 3 benchmarks
The PoseTrack dataset is a large-scale benchmark for multi-person pose estimation and tracking in videos.
103 papers · 5 benchmarks
The image dataset TinyImages contains 80 million images of size 32×32 collected from the Internet, crawling the words in WordNet.
103 papers · 0 benchmarks
FRGC (Face Recognition Grand Challenge)
The data for FRGC consists of 50,000 recordings divided into training and validation partitions.
102 papers · 1 benchmark
The Synthetic Rain Datasets consists of 13,712 clean-rain image pairs gathered from multiple datasets (Rain14000, Rain1800, Rain800, Rain12).
102 papers · 6 benchmarks
Mapillary Vistas Dataset is a diverse street-level imagery dataset with pixel‑accurate and instance‑specific human annotations for understanding street scenes around the world.
101 papers · 3 benchmarks
The McMaster dataset is a dataset for color demosaicing, which contains 18 cropped images of size 500×500.
101 papers · 5 benchmarks
CORD (Consolidated Receipt Dataset for Post-OCR Parsing)
OCR is inevitably linked to NLP since its final output is in text.
100 papers · 1 benchmark
CUHK-SYSU (CUHK-SYSU Person Search Dataset)
The CUKL-SYSY dataset is a large scale benchmark for person search, containing 18,184 images and 8,432 identities.
100 papers · 2 benchmarks
Recent breakthroughs in diffusion models, multimodal pretraining, and efficient finetuning have led to an explosion of text-to-image generative models.
100 papers · 1 benchmark
The dataset consists of 4,738 pairs of images of 232 different scenes including reference pairs.
100 papers · 6 benchmarks
The CrowdPose dataset contains about 20,000 images and a total of 80,000 human poses with 14 labeled keypoints.
99 papers · 2 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.