Home › Datasets › modality › Images
Images datasets
archive 2025-07-28
3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 9 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Images datasets 385–432 of 3,239
MagicBrush is a manually-annotated instruction-guided image editing dataset covering diverse scenarios single-turn, multi-turn, mask-provided, and mask-free editing.
64 papers · 0 benchmarks
The Mall is a dataset for crowd counting and profiling research.
64 papers · 1 benchmark
We propose Localized Narratives, a new form of multimodal image annotations connecting vision and language.
63 papers · 3 benchmarks
SLAKE is an English-Chinese bilingual dataset consisting of 642 images and 14,028 question-answer pairs for training and testing Med-VQA systems.
63 papers · 0 benchmarks
SceneNN is an RGB-D scene dataset consisting of more than 100 indoor scenes.
63 papers · 1 benchmark
COCO-QA is a dataset for visual question answering.
62 papers · 0 benchmarks
FaceScape dataset provides 3D face models, parametric models and multi-view images in large-scale and high-quality.
62 papers · 1 benchmark
This dataset contains 21,889 outfits from polyvore.com, in which 17,316 are for training, 1,497 for validation and 3,076 for testing.
62 papers · 3 benchmarks
TaxiBJ consists of trajectory data from taxicab GPS data and meteorology data in Beijing from four time intervals: 1st Jul.
62 papers · 3 benchmarks
The UTD-MHAD dataset consists of 27 different actions performed by 8 subjects.
62 papers · 2 benchmarks
BTAD (beanTech Anomaly Detection)
The BTAD ( beanTech Anomaly Detection) dataset is a real-world industrial anomaly dataset.
61 papers · 2 benchmarks
CIRR (Compose Image Retrieval on Real-life images)
Composed Image Retrieval (or, Image Retreival conditioned on Language Feedback) is a relatively new retrieval task, where an input query consists of an image and short textual description of how to modify the image.
61 papers · 3 benchmarks
FigureQA is a visual reasoning corpus of over one million question-answer pairs grounded in over 100,000 images.
61 papers · 1 benchmark
The LIP (Look into Person) dataset is a large-scale dataset focusing on semantic understanding of a person.
61 papers · 1 benchmark
PanNuke is a semi automatically generated nuclei instance segmentation and classification dataset with exhaustive nuclei labels across 19 different tissue types.
61 papers · 4 benchmarks
SFEW (Static Facial Expression in the Wild)
The Static Facial Expressions in the Wild (SFEW) dataset is a dataset for facial expression recognition.
61 papers · 1 benchmark
The Salient Person dataset (SIP) contains 929 salient person samples with different poses and illumination conditions.
61 papers · 1 benchmark
CHASEDB1 is a dataset for retinal vessel segmentation which contains 28 color retina images with the size of 999×960 pixels which are collected from both left and right eyes of 14 school children.
59 papers · 2 benchmarks
The Middlebury 2014 dataset contains a set of 23 high resolution stereo pairs for which known camera calibration parameters and ground truth disparity maps obtained with a structured light scanner are available.
59 papers · 2 benchmarks
Set11 is a dataset of 11 grayscale images.
59 papers · 1 benchmark
CASIA-MFSD is a dataset for face anti-spoofing.
58 papers · 1 benchmark
ETH is a dataset for pedestrian detection.
58 papers · 3 benchmarks
ExDark (Exclusively Dark Image Dataset)
The Exclusively Dark (ExDARK) dataset is a collection of 7,363 low-light images from very low-light environments to twilight (i.e 10 different conditions) with 12 object classes (similar to PASCAL VOC) annotated on both image class level…
58 papers · 2 benchmarks
We introduce a dataset of 147 object categories containing over 6000 images that are suitable for the few-shot counting task.
58 papers · 4 benchmarks
SQA3D (Situated Question Answering in 3D Scenes)
SQA3D is a dataset for embodied scene understanding, where an agent needs to comprehend the scene it situates from an first person's perspective and answer questions.
58 papers · 3 benchmarks
Includes 4000 images; 200 from each of 20 categories covering different types of scenes such as Cartoons, Art, Objects, Low resolution images, Indoor, Outdoor, Jumbled, Random, and Line drawings.
57 papers · 2 benchmarks
Composition-1K is a large-scale image matting dataset including 49300 training images and 1000 testing images.
57 papers · 1 benchmark
Dark Zurich is an image dataset containing a total of 8779 images captured at nighttime, twilight, and daytime, along with the respective GPS coordinates of the camera for each image.
57 papers · 3 benchmarks
The EYEDIAP dataset is a dataset for gaze estimation from remote RGB, and RGB-D (standard vision and depth), cameras.
57 papers · 3 benchmarks
PGM (Procedurally Generated Matrices (PGM))
PGM dataset serves as a tool for studying both abstract reasoning and generalisation in models.
57 papers · 0 benchmarks
RELLIS-3D is a multi-modal dataset for off-road robotics.
57 papers · 3 benchmarks
The Stanford Dogs dataset contains 20,580 images of 120 classes of dogs from around the world, which are divided into 12,000 images for training and 8,580 images for testing.
57 papers · 6 benchmarks
The TotalCapture dataset consists of 5 subjects performing several activities such as walking, acting, a range of motion sequence (ROM) and freestyle motions, which are recorded using 8 calibrated, static HD RGB cameras and 13 IMUs…
57 papers · 2 benchmarks
The Wireframe dataset consists of 5,462 images (5,000 for training, 462 for test) of indoor and outdoor man-made scenes.
57 papers · 2 benchmarks
Consists of over one million high-resolution images of varying gaze under extreme head poses.
56 papers · 1 benchmark
MINC (Materials in Context Database)
MINC is a large-scale, open dataset of materials in the wild.
56 papers · 0 benchmarks
PA-100K is a recent-proposed large pedestrian attribute dataset, with 100,000 images in total collected from outdoor surveillance cameras.
56 papers · 1 benchmark
This dataset contains images of unusual dangers which can be encountered by a vehicle on the road – animals, rocks, traffic cones and other obstacles.
56 papers · 1 benchmark
PMC-VQA is a large-scale medical visual question-answering dataset that contains 227k VQA pairs of 149k images that cover various modalities or diseases.
55 papers · 2 benchmarks
Tiny ImageNet-C is an open-source data set comprising algorithmically generated corruptions applied to the Tiny ImageNet (ImageNet-200) test set comprising 200 classes following the concept of ImageNet-C.
55 papers · 0 benchmarks
Fisheye cameras are commonly employed for obtaining a large field of view in surveillance, augmented reality and in particular automotive applications.
55 papers · 1 benchmark
CVC-ClinicDB is an open-access dataset of 612 images with a resolution of 384×288 from 31 colonoscopy sequences.It is used for medical image segmentation, in particular polyp detection in colonoscopy videos.
54 papers · 1 benchmark
DAQUAR (DAtaset for QUestion Answering on Real-world images) is a dataset of human question answer pairs about images.
54 papers · 0 benchmarks
FGNet is a dataset for age estimation and face recognition across ages.
54 papers · 2 benchmarks
Common corruptions dataset for MNIST.
54 papers · 0 benchmarks
The SEMAINE videos dataset contains spontaneous data capturing the audiovisual interaction between a human and an operator undertaking the role of an avatar with four personalities: Poppy (happy), Obadiah (gloomy), Spike (angry) and…
54 papers · 1 benchmark
BEHAVE is a full body human-object interaction dataset with multi-view RGBD frames and corresponding 3D SMPL and object fits along with the annotated contacts between them.
53 papers · 3 benchmarks
CUHK02 (CUHK Person Re-identification Dataset)
CUHK02 is a dataset for person re-identification.
53 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.