Home › Datasets › modality › Images

Images datasets

archive 2025-07-28

3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 8 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Images datasets 337–384 of 3,239

iSAID contains 655,451 object instances for 15 categories across 2,806 high-resolution images.
81 papers · 4 benchmarks
LFSD (Light Field Saliency Database)
The Light Field Saliency Database (LFSD) contains 100 light fields with 360×360 spatial resolution.
80 papers · 1 benchmark
Oulu-CASIA (Oulu-CASIA NIR&VIS facial expression database)
The Oulu-CASIA NIR&VIS facial expression database consists of six expressions (surprise, happiness, sadness, anger, fear and disgust) from 80 people between 23 and 58 years old.
80 papers · 4 benchmarks
Places-LT has an imbalanced training set with 62,500 images for 365 classes from Places-2.
80 papers · 1 benchmark
The ReferIt dataset contains 130,525 expressions for referring to 96,654 objects in 19,894 images of natural scenes.
80 papers · 0 benchmarks
VQG (Visual Question Generation)
VQG is a collection of datasets for visual question generation.
80 papers · 1 benchmark
Volleyball is a video action recognition dataset.
80 papers · 3 benchmarks
VeRi-776 is a vehicle re-identification dataset which contains 49,357 images of 776 vehicles from 20 cameras.
79 papers · 1 benchmark
WIT (Wikipedia-based Image Text)
Wikipedia-based Image Text (WIT) Dataset is a large multimodal multilingual dataset.
79 papers · 1 benchmark
MPIIGaze is a dataset for appearance-based gaze estimation in the wild.
78 papers · 2 benchmarks
The BRATS2017 dataset.
77 papers · 1 benchmark
The TuSimple dataset consists of 6,408 road images on US highways.
76 papers · 1 benchmark
ApolloScape is a large dataset consisting of over 140,000 video frames (73 street scene videos) from various locations in China under varying weather conditions.
74 papers · 4 benchmarks
PETA (Pedestrian Attribute)
The PEdesTrian Attribute dataset (PETA) is a dataset fore recognizing pedestrian attributes, such as gender and clothing style, at a far distance.
74 papers · 1 benchmark
DDAD (Dense Depth for Autonomous Driving)
DDAD is a new autonomous driving benchmark from TRI (Toyota Research Institute) for long range (up to 250m) and dense depth estimation in challenging and diverse urban conditions.
73 papers · 1 benchmark
To the best of our knowledge this is the largest publicly available dataset of face images with gender and age labels for training.
73 papers · 0 benchmarks
misc @inproceedings{RITAC18, author = {Radenovi\'{c}, F.
73 papers · 3 benchmarks
VisDrone is a large-scale benchmark with carefully annotated ground-truth for various important computer vision tasks, to make vision meet drones.
73 papers · 1 benchmark
Birdsnap is a large bird dataset consisting of 49,829 images from 500 bird species with 47,386 images used for training and 2,443 images used for testing.
72 papers · 2 benchmarks
GRAB is a dataset of full-body motions interacting and grasping 3D objects.
72 papers · 1 benchmark
CARPK (car parking lot dataset)
The Car Parking Lot Dataset (CARPK) contains nearly 90,000 cars from 4 different parking lots collected by means of drone (PHANTOM 3 PROFESSIONAL).
71 papers · 1 benchmark
SPair-71k contains 70,958 image pairs with diverse variations in viewpoint and scale.
71 papers · 2 benchmarks
CompCars (Comprehensive Cars)
The Comprehensive Cars (CompCars) dataset contains data from two scenarios, including images from web-nature and surveillance-nature.
70 papers · 1 benchmark
DiffusionDB is a large-scale text-to-image prompt dataset.
70 papers · 1 benchmark
MSP-IMPROV (MSP-IMPROV: An Acted Corpus of Dyadic Interactions to Study Emotion Perception)
We present the MSP-IMPROV corpus, a multimodal emotional database, where the goal is to have control over lexical content and emotion while also promoting naturalness in the recordings.
70 papers · 1 benchmark
MuPoTS-3D (Multiperson Pose Test Set in 3DMulti-person Pose estimation Test Set in 3D)
MuPoTs-3D (Multi-person Pose estimation Test Set in 3D) is a dataset for pose estimation composed of more than 8,000 frames from 20 real-world scenes with up to three subjects.
70 papers · 3 benchmarks
OmniObject3D is a large vocabulary 3D object dataset with massive high-quality real-scanned 3D objects.
70 papers · 0 benchmarks
RSICD (Remote Sensing Image Captioning Dataset)
70 papers · 3 benchmarks
UT-Kinect (UTKinect-Action3D Dataset)
The UT-Kinect dataset is a dataset for action recognition from depth sequences.
70 papers · 2 benchmarks
The BraTS 2015 dataset is a dataset for brain tumor image segmentation.
69 papers · 1 benchmark
CACD (Cross-Age Celebrity Dataset)
The Cross-Age Celebrity Dataset (CACD) contains 163,446 images from 2,000 celebrities collected from the Internet.
69 papers · 1 benchmark
UBFC-rPPG (Univ. Bourgogne Franche-Comté Remote PhotoPlethysmoGraphy)
We introduce here a new database called UBFC-rPPG (stands for Univ.
69 papers · 1 benchmark
VSR (Visual Spatial Reasoning)
The Visual Spatial Reasoning (VSR) corpus is a collection of caption-image pairs with true/false labels.
69 papers · 1 benchmark
AGORA is a synthetic human dataset with high realism and accurate ground truth.
68 papers · 4 benchmarks
CoNSeP (Colorectal Nuclear Segmentation and Phenotypes)
The colorectal nuclear segmentation and phenotypes (CoNSeP) dataset consists of 41 H&E stained image tiles, each of size 1,000×1,000 pixels at 40× objective magnification.
68 papers · 2 benchmarks
FSOD (Few-Shot Object Detection Dataset)
Few-Shot Object Detection Dataset (FSOD) is a high-diverse dataset specifically designed for few-shot object detection and intrinsically designed to evaluate thegenerality of a model on novel categories.
68 papers · 0 benchmarks
The PlantVillage dataset consists of 54303 healthy and unhealthy leaf images divided into 38 categories by species and disease.
68 papers · 1 benchmark
Recipe1M+ is a dataset which contains one million structured cooking recipes with 13M associated images.
68 papers · 3 benchmarks
InLoc is a dataset with reference 6DoF poses for large-scale indoor localization.
67 papers · 1 benchmark
T2I-CompBench is a comprehensive benchmark for open-world compositional text-to-image generation, consisting of 6,000 compositional textual prompts from 3 categories (attribute binding, object relationships, and complex compositions) and 6…
67 papers · 1 benchmark
This dataset focuses on heavily occluded human with comprehensive annotations including bounding-box, humans pose and instance mask.
66 papers · 6 benchmarks
Semantic3D is a point cloud dataset of scanned outdoor scenes with over 3 billion points.
66 papers · 1 benchmark
The CAD-60 and CAD-120 data sets comprise of RGB-D video sequences of humans performing activities which are recording using the Microsoft Kinect sensor.
65 papers · 1 benchmark
MMI (MMI Facial Expression Database)
The MMI Facial Expression Database consists of over 2900 videos and high-resolution still images of 75 subjects.
65 papers · 1 benchmark
Occluded REID is an occluded person dataset captured by mobile cameras, consisting of 2,000 images of 200 occluded persons (see Fig.
65 papers · 1 benchmark
PASCAL-Part is a set of additional annotations for PASCAL VOC 2010.
65 papers · 4 benchmarks
The Places365 dataset is a scene recognition dataset.
65 papers · 7 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.