Home › Datasets › modality › Images
Images datasets
archive 2025-07-28
3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 64 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Images datasets 3025–3072 of 3,239
This dataset is an extremely challenging set of over 8000+ original Fire and Smoke images captured and crowdsourced from over 1200+ urban and rural areas, where each image is manually reviewed and verified by computer vision professionals…
0 papers · 0 benchmarks
BMS-26 (Berkeley Motion Segmentation)
The Berkeley Motion Segmentation Dataset (BMS-26) is a dataset for motion segmentation, which consists of 26 video sequences with pixel-accurate segmentation annotation of moving objects.
0 papers · 0 benchmarks
BSTLD (Bosch Small Traffic Lights Dataset)
This dataset contains 13427 camera images at a resolution of 1280x720 pixels and contains about 24000 annotated traffic lights.
0 papers · 0 benchmarks
Reflectance measurements of Bidirectional Texture Functions (BTFs) Database contains both flat samples: as well as 3D geometry with texture mapped BTFs: furthermore, there are some multispectral BTFs:
0 papers · 0 benchmarks
A Filipino multi-modal language dataset for text+visual tasks.
0 papers · 0 benchmarks
This dataset consists of odometer or speedometer images of bike and car vehicles.
0 papers · 0 benchmarks
Annotated and original images of billboards in Japanese street scapes
0 papers · 0 benchmarks
This dataset comprises fractured and non-fractured X-ray images covering all anatomical body regions, including lower limb, upper limb, lumbar, hips, knees, etc.
0 papers · 0 benchmarks
This dataset consists of both fractured and non-fractured X-ray images encompassing various anatomical regions of the body, such as the lower limb, upper limb, lumbar region, hips, knees, and more.
0 papers · 0 benchmarks
This dataset consists of images of bottles and cups.
0 papers · 0 benchmarks
Boxy (Boxy Vehicles Dataset)
A large vehicle detection dataset with almost two million annotated vehicles for training and evaluating object detection methods for self-driving cars on freeways.
0 papers · 0 benchmarks
The Burmese Handwritten Digit Dataset (BHDD) is a dataset project specifically created for recognizing handwritten Burmese digits.
0 papers · 0 benchmarks
Medical report generation (MRG), which aims to automatically generate a textual description of a specific medical image (e.g., a chest X-ray), has recently received increasing research interest.
0 papers · 0 benchmarks
CBLPRD-330k (China-Balanced-License-Plate-Recognition-Dataset-330k)
A high-quality, balanced dataset of 330,000 images featuring various types of Chinese license plates.
0 papers · 0 benchmarks
COCO-Facet is a benchmark for attribute-focused text-to-image retrieval, comprising 9,112 queries with 100 candidate images for each.
0 papers · 0 benchmarks
Applications of unmanned aerial vehicle (UAV) in logistics, agricultural automation, urban management, and emergency response are highly dependent on oriented object detection (OOD) to enhance visual perception.
0 papers · 0 benchmarks
CUHK occlusion dataset includes 1,063 images with occluded pedestrians.
0 papers · 0 benchmarks
CVGL Camera Calibration Dataset consists of 49 camera configurations with town 1 having 25 configurations while town 2 having 24 configurations.
0 papers · 0 benchmarks
ChaBuD (Change detection for Burned area Delineation)
The dataset comprises patches of size 512x512 pixels collected from Sentinel-2 L2A satellite mission.
0 papers · 0 benchmarks
ChaLearn Pose is a subset of the ChaLearn 2013 Multi-modal gesture dataset from Escalera et al.
0 papers · 0 benchmarks
The photo fixation of cherry fruitlets was done in the LatHort orchard in Dobele, at the development of fruit (BBCH stage 72).
0 papers · 0 benchmarks
The photo fixation of cherry fruits was done in the LatHort orchard in Dobele, at the beginning of fruit coloration (BBCH stage 81).
0 papers · 0 benchmarks
The Cityscapes-Motion dataset is a supplement to the semantic annotations provided by the Cityscapes dataset, containing 2975 training images and 500 validation images.
0 papers · 0 benchmarks
CloudSEN12 is a LARGE dataset (~1 TB) for cloud semantic understanding that consists of 49,400 image patches (IP) that are evenly spread throughout all continents except Antarctica.
0 papers · 0 benchmarks
Key Points - Purpose: Captures crossroad navigation under cloudy weather conditions.
0 papers · 0 benchmarks
CodeSCAN (ScreenCast ANalysis for Video Programming Tutorials)
CodeSCAN is the first large-scale and diverse dataset of coding screenshots with pixel-perfect annotations.
0 papers · 0 benchmarks
Colorectal-Liver-Metastases (Colorectal-Liver-Metastases | Preoperative CT and Survival Data for Patients Undergoing Resection of Colorectal Liver Metastases)
This collection consists of DICOM images and DICOM Segmentation Objects (DSOs) for 197 patients with Colorectal Liver Metastases (CRLM).
0 papers · 0 benchmarks
A fundamental component of human vision is our ability to parse complex visual scenes and judge the relations between their constituent objects.
0 papers · 0 benchmarks
This dataset is an extremely challenging set of over 20,000+ original Construction vehicle images captured and crowdsourced from over 600+ urban and rural areas, where each image is manually reviewed and verified by computer vision…
0 papers · 0 benchmarks
This dataset is the images of corn seeds considering the top and bottom view independently (two images for one corn seed: top and bottom).
0 papers · 0 benchmarks
This dataset consists of images of Cracked screen like cracked mobile screen.
0 papers · 0 benchmarks
Crowd 11 (A Dataset for Fine Grained Crowd Behaviour Analysis)
This dataset defines a total of 11 crowd motion patterns and it is composed of over 6000 video sequences with an average length of 100 frames per sequence.
0 papers · 0 benchmarks
This dataset is an extremely challenging set of over 3000+ original Crowd images captured and crowdsourced from over 300+ urban and rural areas, where each image is manually reviewed and verified by computer vision professionals at…
0 papers · 0 benchmarks
The database consists of 89 colour fundus images of which 84 contain at least mild non-proliferative signs (Microaneurysms) of the diabetic retinopathy, and 5 are considered as normal which do not contain any signs of the diabetic…
0 papers · 0 benchmarks
This is an image splicing dataset including different types of preprocessing and postprocessing techniques.
0 papers · 0 benchmarks
DMS (Dense Material Segmentation Dataset)
The Dense Material Segmentation Dataset (DMS) consists of 3 million polygon labels of material categories (metal, wood, glass, etc) for 44 thousand RGB images.
0 papers · 0 benchmarks
DR HAGIS (Diabetic Retinopathy, Hypertension, Age-related macular degeneration and Glacuoma ImageS)
The DR HAGIS database has been created to aid the development of vessel extraction algorithms suitable for retinal screening programmes.
0 papers · 0 benchmarks
Dataset for the DREAMING - Diminished Reality for Emerging Applications in Medicine through Inpainting Challenge!
0 papers · 0 benchmarks
DUS (Daimler Urban Segmentation)
The Daimler Urban Segmentation Dataset is a dataset for semantic segmentation.
0 papers · 0 benchmarks
The deliberate manipulation of public opinion, especially through altered images, poses a significant danger to society.
0 papers · 0 benchmarks
This dataset is collected by Datacluster Labs.
0 papers · 0 benchmarks
a large video dataset captured with UAVs in different complex real-world scenes, with multiple representations, suitable for multi-task learning.
0 papers · 0 benchmarks
Context As mentioned in the reference paper: Dust storms are considered a severe meteorological disaster, especially in arid and semi-arid regions, which is characterized by dust aerosol-filled air and strong winds across an extensive area.
0 papers · 0 benchmarks
The MICCAI 2020 EMIDEC dataset is a dataset for classifying normal and pathological cases from the clinical information with or without DE-MRI, and secondly to automatically detect the different relevant areas (the myocardial contours, the…
0 papers · 0 benchmarks
ESP dataset (Evaluation for Styled Prompt dataset) is a new benchmark for zero-shot domain-conditional caption generation.
0 papers · 0 benchmarks
The Edge Milling Heads data set comprises 144 images of an edge profile cutting head of a milling machine.
0 papers · 0 benchmarks
This dataset is an extremely challenging set of over 5000+ original Electronic Items images captured and crowdsourced from over 1000+ urban and rural areas, where each image is manually reviewed and verified by computer vision…
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.