Home › Datasets › modality › Images
Images datasets
archive 2025-07-28
3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 56 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Images datasets 2641–2688 of 3,239
Onchocerciasis is causing blindness in over half a million people in the world today.
1 paper · 0 benchmarks
The AASL-Clear dataset is a collection of RGB images featuring Arabic alphabet sign Language gestures with backgrounds removed.
1 paper · 1 benchmark
Noise of Web (NoW) is a challenging noisy correspondence learning (NCL) benchmark for robust image-text matching/retrieval models.
1 paper · 0 benchmarks
Number of images: 1,657 images during or after the fire If you use the dataset, please cite the following works: > Padilha, Rafael and Andaló, Fernanda A.
1 paper · 0 benchmarks
Authors of the Dataset: - Pratik Bhowal (B.E., Dept of Electronics and Instrumentation Engineering, Jadavpur University Kolkata, India)…
1 paper · 1 benchmark
Dynamic occupancy grids generated from NuScenes dataset.
1 paper · 0 benchmarks
ODSI-DB (ODSI-DB – Oral and Dental Spectral Image Database)
ODSI-DB is an image database of oral and dental reflectance spectral images of human test subjects.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
OLID I (An Open Leaf Image Dataset of Bangladesh's Major Crops)
The success of any AI-driven system relies heavily on vast amounts of training data.
1 paper · 0 benchmarks
OSLD (Open Set Logo Detection Dataset)
Open Set Logo Detection Dataset (OSLD Dataset) is a dataset of eCommerce product images with associated brand logo images.
1 paper · 0 benchmarks
This is the paper “DF-RAP: A Robust Adversarial Perturbation for Defending against Deepfakes in Real-world Social Network Scenarios" OSN-transmission CelebA sampling dataset collected by manual upload and download.
1 paper · 0 benchmarks
Due to the free-form nature of the open vocabulary image classification task, special annotations are required for image sets used for evaluation purposes.
1 paper · 4 benchmarks
The ObMan-Ego is a large-scale synthetic hand dataset with egocentric scenes in which the simulated hands are provided by ObMan.
1 paper · 0 benchmarks
An object-centric version of Stylized COCO to benchmark texture bias and out-of-distribution robustness of vision models.
1 paper · 0 benchmarks
This dataset is a collection of memes from various existing datasets, online forums, and freshly scrapped contents.
1 paper · 0 benchmarks
The dataset is for research on the label distribution shift between multiple domain adaptations.
1 paper · 0 benchmarks
The Swiss Drone data set was recorded around Cheseaux-sur-Lausanne in Switzerland using a senseFly eBee Classic in 2014 (SenseFly, 2020).
1 paper · 1 benchmark
Omni-Image is built as a challenging but tractable dataset for continual learning and few-shot learning.
1 paper · 0 benchmarks
In order to evaluate the effectiveness of NToP in real-world scenarios, we collect a new dataset OmniLab with a top-view omnidirectional camera, mounted on the ceiling of two different rooms (bedroom, living room) at 2.5 m height.
1 paper · 0 benchmarks
To effectively evaluate OmniCount across open-vocabulary, supervised, and few-shot counting tasks, a dataset catering to a broad spectrum of visual categories and instances featuring various visual categories with multiple instances and…
1 paper · 2 benchmarks
Open MIC (Open Museum Identification Challenge)
Open MIC (Open Museum Identification Challenge) contains photos of exhibits captured in 10 distinct exhibition spaces of several museums which showcase paintings, timepieces, sculptures, glassware, relics, science exhibits, natural history…
1 paper · 0 benchmarks
We create the first open-source large-scale S2V generation dataset OpenS2V-5M, which consists of five million high-quality 720P subject-text-video triples.
1 paper · 1 benchmark
A high-resolution multi-sensor remote sensing scene classification dataset, appropriate for training and evaluating image classification models in the remote sensing domain.
1 paper · 0 benchmarks
Image corruptions modelling primary optical aberrations.
1 paper · 0 benchmarks
Orchid2024 is a fine-grained classification dataset specifically designed for Chinese Cymbidium orchid cultivars.
1 paper · 0 benchmarks
The Oxford Road Boundaries is a dataset designed for training and testing machine-learning-based road-boundary detection and inference approaches.
1 paper · 0 benchmarks
Overview PASSION derm is a pioneering initiative dedicated to closing the diversity gap in dermatology datasets.
1 paper · 0 benchmarks
We compiled a new dataset (the PERO layout dataset) that contains 683 images from various sources and historical periods with complete manual text block, text line polygon and baseline annotations.
1 paper · 0 benchmarks
PEnG (Pose-Enhanced Geo-Localisation)
This dataset builds upon the SpaGBOL dataset - a graph-based dataset covering numerous cities across the globe for the purpose of structured city-scale Cross-View Geo-Localisation (CVGL).
1 paper · 0 benchmarks
PFN-VT (PFN Visuo-Tactile Dataset)
PFN-VT is a dataset for the estimation of tactile properties from vision, such as slipperiness or roughness.
1 paper · 0 benchmarks
PLAD (Point Line and Depth dataset)
PLAD is a dataset where sparse depth is provided by line-based visual SLAM to verify StructMDC.
1 paper · 1 benchmark
Object detection dataset featuring people walking on grass captured aboard a UAV.
1 paper · 0 benchmarks
POIE (Products for OCR and Information Extraction)
Products for OCR and Information Extraction (POIE) dataset derives from camera images of various products in the real world.
1 paper · 0 benchmarks
The Prima head pose dataset consists of 2790 images of 15 persons recorded twice.
1 paper · 1 benchmark
PSU NRTDB (PSU Near-Regular Texture Database)
The PSU Near-Regular Texture Database is a texture dataset.
1 paper · 0 benchmarks
PTCGA200 (Patch TCGA in 200 microns by 512 px)
PTCGA200 is a public pathological H&E image datasets from Patch TCGA in 200 microns by 512 px.
1 paper · 0 benchmarks
PTVD is a plot-oriented multimodal dataset in the TV domain.
1 paper · 0 benchmarks
PVDN (Provident Vehicle Detection at Night)
PVDN is a dataset of vehicle detection at night, using light reflections caused by their headlamps.
1 paper · 0 benchmarks
PWISeg (PWISeg Surgical Instruments Dataset)
Overview The Surgical Instruments Recognition Dataset is a groundbreaking collection of high-resolution images (1280x960 pixels) specifically designed for the recognition and categorization of surgical instruments.
1 paper · 0 benchmarks
Pan+ChiPhoto dataset is a Chinese character dataset.
1 paper · 0 benchmarks
Panoramic Video Panoptic Segmentation Dataset is a large-scale dataset that offers high-quality panoptic segmentation labels for autonomous driving.
1 paper · 0 benchmarks
Despite recent advances in vision-and-language tasks, most progress is still focused on resource-rich languages such as English.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
PatternCom is a composed image retrieval benchmark based on PatternNet.
1 paper · 1 benchmark
The Peripheral Blood Cell} (PBC) dataset consists of 17,092 images.
1 paper · 0 benchmarks
Persian Font Recognition (PFR) A dataset in order to solve font recognition for the Persian language.
1 paper · 1 benchmark
Persian Text Image Segmentation (PTI SEG) This dataset is part of a paper titled "Persis: A Persian Font Recognition Pipeline Using Convolutional Neural Networks".
1 paper · 1 benchmark
Pesteh-Set is made of two parts.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.