Home › Datasets › modality › Images

Images datasets

archive 2025-07-28

3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 16 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Images datasets 721–768 of 3,239

CDD-11 (Composite Degradation Dataset 11)
An image restoration dataset
21 papers · 1 benchmark
COCO-CN is a bilingual image description dataset enriching MS-COCO with manually written Chinese sentences and tags.
21 papers · 1 benchmark
ContactDB is a dataset of contact maps for household objects that captures the rich hand-object contact that occurs during grasping, enabled by use of a thermal camera.
21 papers · 1 benchmark
A team of researchers from Qatar University, Doha, Qatar, and the University of Dhaka, Bangladesh along with their collaborators from Pakistan and Malaysia in collaboration with medical doctors have created a database of chest X-ray images…
21 papers · 0 benchmarks
DigiFace-1M is a synthetic dataset for face recognition, obtained by rendering digital faces using a computer graphics pipeline.
21 papers · 0 benchmarks
Digits (Optical Recognition of Handwritten Digits)
The DIGITS dataset consists of 1797 8×8 grayscale images (1439 for training and 360 for testing) of handwritten digits.
21 papers · 3 benchmarks
EPHOIE (phtnsantader@gmail.com)
EPHOIE is a fully-annotated dataset which is the first Chinese benchmark for both text spotting and visual information extraction.
21 papers · 2 benchmarks
FIW (Families In The Wild)
FIW is a large and comprehensive database available for kinship recognition.
21 papers · 0 benchmarks
The INRIA Aerial Image Labeling dataset is comprised of 360 RGB tiles of 5000×5000px with a spatial resolution of 30cm/px on 10 cities across the globe.
21 papers · 1 benchmark
A benchmark dataset for out-of-distribution detection.
21 papers · 1 benchmark
SICAPv2 is a database containing prostate histology whole slide images with both annotations of global Gleason scores and path-level Gleason grades.
21 papers · 0 benchmarks
Common Objects in 3D is a large-scale dataset with real multi-view images of object categories annotated with camera poses and ground truth 3D point clouds.
20 papers · 1 benchmark
EORSSD (Extended Optical Remote Sensing Saliency Detection)
The Extended Optical Remote Sensing Saliency Detection (EORSSD) dataset is an extension of the ORSSD dataset.
20 papers · 0 benchmarks
This dataset has 1,842 images with pixel-level DR-related lesion annotations, and 1,000 images with image-level labels graded by six board-certified ophthalmologists with intra-rater consistency.
20 papers · 0 benchmarks
FMB Dataset (Full-time Multi-modality Benchmark Dataset)
FMB contains 1500 well-registered infrared and visible image pairs with 14 annotated pixel-level categories.
20 papers · 1 benchmark
The George Washington dataset contains 20 pages of letters written by George Washington and his associates in 1755 and thereby categorized into historical collection.
20 papers · 0 benchmarks
MSU NR VQA Database (MSU No-Reference Video Quality Assessment Database)
The dataset was created for video quality assessment problem.
20 papers · 2 benchmarks
PASCAL VOC 2011 is an image segmentation dataset.
20 papers · 2 benchmarks
PhotoChat, the first dataset that casts light on the photo sharing behavior in online messaging.
20 papers · 2 benchmarks
he RSSCN7 dataset contains satellite images acquired from Google Earth, which is originally collected for remote sensing scene classification.
20 papers · 1 benchmark
20 real low-resolution images selected from existing datasets or downloaded from internet
20 papers · 0 benchmarks
StreetHazards is a synthetic dataset for anomaly detection, created by inserting a diverse array of foreign objects into driving scenes and re-render the scenes with these novel objects.
20 papers · 1 benchmark
VIDIT (Virtual Image Dataset for Illumination Transfer)
VIDIT is a reference evaluation benchmark and to push forward the development of illumination manipulation methods.
20 papers · 1 benchmark
WebFace260M is a million-scale face benchmark, which is constructed for the research community towards closing the data gap behind the industry.
20 papers · 0 benchmarks
The iLIDS-VID dataset is a person re-identification dataset which involves 300 different pedestrians observed across two disjoint camera views in public open space.
20 papers · 2 benchmarks
Argoverse-HD is a dataset built for streaming object detection, which encompasses real-time object detection, video object detection, tracking, and short-term forecasting.
19 papers · 4 benchmarks
BCI (Breast Cancer Immunohistochemical Image Generation)
The evaluation of human epidermal growth factor receptor 2 (HER2) expression is essential to formulate a precise treatment for breast cancer.
19 papers · 1 benchmark
Chaoyang dataset contains 1111 normal, 842 serrated, 1404 adenocarcinoma, 664 adenoma, and 705 normal, 321 serrated, 840 adenocarcinoma, 273 adenoma samples for training and testing, respectively.
19 papers · 2 benchmarks
DeepFish as a benchmark suite with a large-scale dataset to train and test methods for several computer vision tasks.
19 papers · 1 benchmark
FBMS-59 (Freiburg-Berkeley Motion Segmentation)
The Freiburg-Berkeley Motion Segmentation Dataset (FBMS-59) is a dataset for motion segmentation, which extends the BMS-26 dataset with 33 additional video sequences.
19 papers · 3 benchmarks
FoodSeg103 (lewisnjue)
FoodSeg103 is a new food image dataset containing 7,118 images.
19 papers · 1 benchmark
A benchmark dataset for out-of-distribution detection.
19 papers · 1 benchmark
InfiMM-Eval (Complex Open-ended Reasoning Evaluation for Multi-Modal Language Models)
Multi-modal Large Language Models (MLLMs) are increasingly prominent in the field of artificial intelligence.
19 papers · 1 benchmark
This data set provides Light Detection and Ranging (LiDAR) data and stereo image with various position sensors targeting a highly complex urban environment.
19 papers · 0 benchmarks
MCubeS (Multimodal Material Segmentation Dataset)
Multimodal material segmentation (MCubeS) dataset contains 500 sets of images from 42 street scenes.
19 papers · 1 benchmark
ModaNet is a street fashion images dataset consisting of annotations related to RGB images.
19 papers · 1 benchmark
OVEN (Open-domain Visual Entity Recognition)
In this project, we formally present the task of Open-domain Visual Entity recognitioN (OVEN), where a model need to link an image onto a Wikipedia entity with respect to a text query.
19 papers · 1 benchmark
PKLot (A Robust Dataset for Parking Lot Classification)
The PKLot dataset contains 12,417 images of parking lots and 695,899 images of parking spaces segmented from them, which were manually checked and labeled.
19 papers · 1 benchmark
SSP-3D (Sports Shape and Pose 3D)
SSP-3D is an evaluation dataset consisting of 311 images of sportspersons in tight-fitted clothes, with a variety of body shapes and poses.
19 papers · 1 benchmark
Contains data from three platforms, i.e., synthetic drones, satellites and ground cameras of 1,652 university buildings around the world.
19 papers · 2 benchmarks
WanJuan is a large-scale training corpus that includes multiple modalities.
19 papers · 0 benchmarks
AI-TOD (Tiny Object Detection in Aerial Images)
AI-TOD comes with 700,621 object instances for eight categories across 28,036 aerial images.
18 papers · 2 benchmarks
Amazon Beauty (Amazon Beauty 5-core)
This dataset includes reviews (ratings, text, helpfulness votes), product metadata (descriptions, category information, price, brand, and image features), and links (also viewed/also bought graphs).
18 papers · 3 benchmarks
CASIA-B is a large multiview gait database, which is created in January 2005.
18 papers · 1 benchmark
We collect a new dataset of human-posed free-form natural language questions about CLEVR images.
18 papers · 1 benchmark
The Endomapper dataset is the first collection of complete endoscopy sequences acquired during regular medical practice, including slow and careful screening explorations, making secondary use of medical data.
18 papers · 0 benchmarks
The dataset collected at the University of Florence during 2012, has been captured using a Kinect camera.
18 papers · 1 benchmark
ICDAR2017 is a dataset for scene text detection.
18 papers · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.