Home › Datasets › modality › Images

Images datasets

archive 2025-07-28

3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 11 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Images datasets 481–528 of 3,239

The ScanNet200 benchmark studies 200-class 3D semantic segmentation - an order of magnitude more class categories than previous 3D scene understanding benchmarks.
45 papers · 3 benchmarks
The Talk2Car dataset finds itself at the intersection of various research domains, promoting the development of cross-disciplinary solutions for improving the state-of-the-art in grounding natural language into visual space.
45 papers · 0 benchmarks
The dataset contains over 15K images of 20 people (6 females and 14 males - 4 people were recorded twice).
44 papers · 1 benchmark
COCO-O(ut-of-distribution) contains 6 domains (sketch, cartoon, painting, weather, handmake, tattoo) of COCO objects which are hard to be detected by most existing detectors.
44 papers · 1 benchmark
Occ3D is a dataset for 3D occupancy prediction, which aims to estimate the detailed occupancy and semantics of objects from multi-view images.
44 papers · 1 benchmark
Oxford105k is the combination of the Oxford5k dataset and 99782 negative images crawled from Flickr using 145 most popular tags.
44 papers · 0 benchmarks
The PanoContext dataset contains 500 annotated cuboid layouts of indoor environments such as bedrooms and living rooms.
44 papers · 1 benchmark
The SCUT-CTW1500 dataset contains 1,500 images: 1,000 for training and 500 for testing.
44 papers · 3 benchmarks
For understanding multimodal language used in expressing humor.
44 papers · 0 benchmarks
DensePASS - a novel densely annotated dataset for panoramic segmentation under cross-domain conditions, specifically built to study the Pinhole-to-Panoramic transfer and accompanied with pinhole camera training examples obtained from…
43 papers · 1 benchmark
ImageNet-S (ImageNet Semantic Segmentation)
Powered by the ImageNet dataset, unsupervised learning on large-scale data has made significant advances for classification tasks.
43 papers · 6 benchmarks
Occluded-DukeMTMC contains 15,618 training images, 17,661 gallery images, and 2,210 occluded query images.
43 papers · 1 benchmark
Accurate modeling of priors over 3D human pose is fundamental to many problems in computer vision.
43 papers · 0 benchmarks
The Stacked MNIST dataset is derived from the standard MNIST dataset with an increased number of discrete modes.
43 papers · 1 benchmark
WebQA, is a new benchmark for multimodal multihop reasoning in which systems are presented with the same style of data as humans when searching the web: Snippets and Images.
43 papers · 0 benchmarks
gRefCOCO is the first large-scale Generalized Referring Expression Segmentation dataset that contains multi-target, no-target, and single-target expressions.
43 papers · 2 benchmarks
C-GQA (Compositional GQA)
We propose a split built on top of Stanford GQA dataset originally proposed for VQA and name it Compositional GQA (C-GQA) dataset (see supplementary for the details).
42 papers · 0 benchmarks
CASIA-FASD is a small face anti-spoofing dataset containing 50 subjects.
42 papers · 0 benchmarks
We release E-commerce Dialogue Corpus, comprising a training data set, a development set and a test set for retrieval based chatbot.
42 papers · 1 benchmark
IconQA (Icon Question Answering)
Current visual question answering (VQA) tasks mainly consider answering human-annotated questions for natural images in the daily-life context.
42 papers · 1 benchmark
The Oxford RobotCar Dataset contains over 100 repetitions of a consistent route through Oxford, UK, captured over a period of over a year.
42 papers · 3 benchmarks
ExpW (Expression in-the-Wild)
The Expression in-the-Wild (ExpW) dataset is for facial expression recognition and contains 91,793 faces manually labeled with expressions.
41 papers · 1 benchmark
The I-Haze dataset contains 25 indoor hazy images (size 2833×4657 pixels) training.
41 papers · 1 benchmark
JFT-3B is an internal Google dataset and a larger version of the JFT-300M dataset.
41 papers · 0 benchmarks
KITTI Road is road and lane estimation benchmark that consists of 289 training and 290 test images.
41 papers · 0 benchmarks
Million-AID is a large-scale benchmark dataset containing a million instances for RS scene classification.
41 papers · 0 benchmarks
NT-VOT211 consists of 211 diverse videos, offering 211,000 well-annotated frames with 8 attributes including camera motion, deformation, fast motion, motion blur, tiny target, distractors, occlusion and out-of-view.
41 papers · 1 benchmark
Partial REID is a specially designed partial person reidentification dataset that includes 600 images from 60 people, with 5 full-body images and 5 occluded images per person.
41 papers · 1 benchmark
AID (Aerial Image Dataset)
AID is a new large-scale aerial image dataset, by collecting sample images from Google Earth imagery.
40 papers · 2 benchmarks
Florence (Florence 3D Faces)
The Florence 3D faces dataset consists of: High-resolution 3D scans of human faces from many subjects.
40 papers · 1 benchmark
RedCaps is a large-scale dataset of 12M image-text pairs collected from Reddit.
40 papers · 0 benchmarks
Samanantar is the largest publicly available parallel corpora collection for Indic languages: Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Oriya, Punjabi, Tamil, Telugu.
40 papers · 0 benchmarks
We propose the first question-answering dataset driven by STEM theorems.
40 papers · 1 benchmark
Veri-Wild is the largest vehicle re-identification dataset (as of CVPR 2019).
40 papers · 3 benchmarks
Visual Wake Words represents a common microcontroller vision use-case of identifying whether a person is present in the image or not, and provides a realistic benchmark for tiny vision models.
40 papers · 1 benchmark
AP-10K is the first large-scale benchmark for general animal pose estimation, to facilitate the research in animal pose estimation.
39 papers · 2 benchmarks
BUFF (Bodies Under Flowing Fashion)
BUFF consists of 5 subjects, 3 male and 2 female wearing 2 clothing styles: a) t-shirt and long pants and b) a soccer outfit.
39 papers · 1 benchmark
The Campus and Shelf datasets were presented in the paper 3D Pictorial Structures for Multiple Human Pose Estimation.
39 papers · 2 benchmarks
DIS5K (Dichotomous Image Segmentation (DIS) Dataset)
To build the highly accurate Dichotomous Image Segmentation dataset (DIS5K), we first manually collected over 12,000 images from Flickr1 based on our pre-designed keywords.
39 papers · 5 benchmarks
The largest and cleanest face recognition dataset Glint360K, which contains 17,091,657 images of 360,232 individuals, baseline models trained on Glint360K can easily achieve state-of-the-art performance.
39 papers · 0 benchmarks
RoadTracer is a dataset for extraction of road networks from aerial images.
39 papers · 0 benchmarks
The SIXray dataset is constructed by the Pattern Recognition and Intelligent System Development Laboratory, University of Chinese Academy of Sciences.
39 papers · 1 benchmark
TDIUC (Task Directed Image Understanding Challenge)
Task Directed Image Understanding Challenge (TDIUC) dataset is a Visual Question Answering dataset which consists of 1.6M questions and 170K images sourced from MS COCO and the Visual Genome Dataset.
39 papers · 1 benchmark
VerSe (Large Scale Vertebrae Segmentation Challenge)
Spine or vertebral segmentation is a crucial step in all applications regarding automated quantification of spinal morphology and pathology.
39 papers · 0 benchmarks
EMOTIC (EMOTIons in Context)
The EMOTIC dataset, named after EMOTions In Context, is a database of images with people in real environments, annotated with their apparent emotions.
38 papers · 2 benchmarks
GazeFollow is a large-scale dataset annotated with the location of where people in images are looking.
38 papers · 1 benchmark
Imagenette is a subset of 10 easily classified classes from Imagenet (bench, English springer, cassette player, chain saw, church, French horn, garbage truck, gas pump, golf ball, parachute).
38 papers · 1 benchmark
The SUN Attribute dataset consists of 14,340 images from 717 scene categories, and each category is annotated with a taxonomy of 102 discriminate attributes.
38 papers · 2 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.