Home › Datasets › modality › Images
Images datasets
archive 2025-07-28
3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 43 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Images datasets 2017–2064 of 3,239
PAD Dataset (Pose-agnostic/Multi-pose Anomaly Detection Dataset)
Multi-pose Anomaly Detection (MAD) dataset, which represents the first attempt to evaluate the performance of pose-agnostic anomaly detection.
2 papers · 1 benchmark
PAL4Inpaint is a dataset consisting of 4,795 inpainting results with per-pixel perceptual artifacts annotations designed for image inpainting tasks.
2 papers · 0 benchmarks
Appearance-based gaze estimation systems have shown great progress recently, yet the performance of these techniques depend on the datasets used for training.
2 papers · 0 benchmarks
PHSPD (Polarization Human Shape and Pose Dataset)
PHSPD is a home-grown polarization image dataset of various human shapes and poses.
2 papers · 0 benchmarks
The PKU dataset has almost 4,000 images categorized into five groups (G1-G5) that show different situations.
2 papers · 0 benchmarks
Automated leaf segmentation is a challenging area in computer vision.
2 papers · 0 benchmarks
PUMaVOS (Partial and Unusual Masks for Video Object Segmentation)
PUMaVOS is a dataset of challenging and practical use cases inspired by the movie production industry.
2 papers · 0 benchmarks
Dataset Card for The Cancer Genome Atlas (TCGA) Multimodal Dataset The Cancer Genome Atlas (TCGA) Multimodal Dataset is a comprehensive collection of clinical data, pathology reports, molecular, and slide images for cancer patients.
2 papers · 0 benchmarks
Pano3D is a new benchmark for depth estimation from spherical panoramas.
2 papers · 0 benchmarks
Paper2Fig100k is a dataset with over 100k images of figures and texts from research papers.
2 papers · 0 benchmarks
Synthetic dataset of over 13,000 images of damaged and intact parcels with full 2D and 3D annotations in the COCO format.
2 papers · 0 benchmarks
The data includes all movement trajectories extracted from the videos of Parkinson's assessments using Convolutional Pose Machines (CPM) as well as the confidence values from CPM.
2 papers · 0 benchmarks
Perseus is a dataset for Cross-Lingual Summarization (CLS) which collects about 94K Chinese scientific documents paired with English summaries.
2 papers · 0 benchmarks
PersonPath22 is a large-scale multi-person tracking dataset containing 236 videos captured mostly from static-mounted cameras, collected from sources where we were given the rights to redistribute the content and participants have given…
2 papers · 1 benchmark
A set of 221 stereo videos captured by the SOCRATES stereo camera trap in a wildlife park in Bonn, Germany between February and July of 2022.
2 papers · 0 benchmarks
The Poser dataset is a dataset for pose estimation which consists of 1927 training and 418 test images.
2 papers · 0 benchmarks
PubFig (Public Figures Face Database)
The PubFig database is a large, real-world face dataset consisting of 58,797 images of 200 people collected from the internet.
2 papers · 0 benchmarks
RHM (Rhm: Robot house multi-view human activity recognition dataset)
The Robot House Multi-View dataset (RHM) contains four views: Front, Back, Ceiling, and Robot Views.
2 papers · 1 benchmark
RISE is a large-scale video dataset for Recognizing Industrial Smoke Emissions.
2 papers · 0 benchmarks
RISEdb (Robust Indoor Localization in Complex Scenarios (RISE) database)
The RISE (Robust Indoor Localization in Complex Scenarios) dataset is meant to train and evaluate visual indoor place recognizers.
2 papers · 0 benchmarks
RL Unplugged is suite of benchmarks for offline reinforcement learning.
2 papers · 0 benchmarks
ROPE (Recognition-based Object Probing Evaluation)
We introduce Recognition-based Object Probing Evaluation (ROPE), an automated evaluation protocol that considers the distribution of object classes within a single image during testing and uses visual referring prompts to eliminate…
2 papers · 0 benchmarks
RSOC (Remote Sensing Object Counting)
RSOC is a large-scale object counting dataset with remote sensing images, which contains four important geographic objects: buildings, crowded ships in harbors, large-vehicles and small-vehicles in parking lots.
2 papers · 0 benchmarks
RUGD (RUGD: Robot Unstructured Ground Driving)
A Video Dataset for Visual Perception and Autonomous Navigation in Unstructured Environments.
2 papers · 1 benchmark
RaidaR (RaidaR: A Rich Annotated Image Dataset of Rainy Street Scenes)
RaidaR is a rich annotated image dataset of rainy street scenes.
2 papers · 0 benchmarks
RaindropClarity (A Dual-Focused Dataset for Day and Night Raindrop Removal)
Existing raindrop removal datasets have two shortcomings.
2 papers · 0 benchmarks
Our dataset consists of over 1000 fractured frescoes.
2 papers · 0 benchmarks
ReactionGIF is an affective dataset of 30K tweets which can be used for tasks like induced sentiment prediction and multilabel classification of induced emotions.
2 papers · 0 benchmarks
Real-CE is a real-world Chinese-English benchmark dataset for the task of STISR with the emphasis on restoring structurally complex Chinese characters.
2 papers · 0 benchmarks
The dataset contains patches of facial reflectance as described in the paper, namely the diffuse albedo, diffuse normals, specular albedo, specular normals, as well as the shape in UV space.
2 papers · 0 benchmarks
Rendered Handpose Dataset contains 41258 training and 2728 testing samples.
2 papers · 0 benchmarks
The Rendered SST2 dataset is a dataset released by OpenAI, that measures the optical character recognition capability of visual representations.
2 papers · 1 benchmark
Rent3D++ is an extension of the Rent3D floorplans + photos dataset.
2 papers · 1 benchmark
The Retina Benchmark is a set of real-world tasks that accurately reflect such complexities and are designed to assess the reliability of predictive models in safety-critical scenarios.
2 papers · 0 benchmarks
The Retinal Microsurgery dataset is a dataset for surgical instrument tracking.
2 papers · 0 benchmarks
We collect a dataset of Rich Human Feedback on 18K images (RichHF-18K), which contains (i) point annotations on the image that highlight regions of implausibility/artifacts, and text-image misalignment; (ii) labeled words on the prompts…
2 papers · 0 benchmarks
The S2-100K dataset is a dataset of 100,000 multi-spectral satellite images and their corresponding locations (latitude / longitude coordinates of the image centroid) sampled from Sentinel-2 via the Microsoft Planetary Computer.
2 papers · 0 benchmarks
Our proposed Synthetic-to-Real benchmark for more practical visual DA (termed S2RDA) includes two challenging transfer tasks of S2RDA-49 and S2RDA-MS-39.
2 papers · 0 benchmarks
SA-Det-100k is a large-scale class-agnostic object detection dataset for Research Purposes only.
2 papers · 1 benchmark
SB20 (Sugar Beet 2020 University of Bonn)
Video sequences captured at a field on Campus Kleinaltendorf (CKA), University of Bonn, captured by BonBot-I, an autonomous weeding robot.
2 papers · 0 benchmarks
SEPE 8K dataset is made of 40 different 8K (8192 x 4320) video sequences and 40 variant 8K (8192 x 5464) images.
2 papers · 1 benchmark
SI-SCORE is a synthetic dataset for the analysis of robustness to object location, rotation and size.
2 papers · 0 benchmarks
SILVR (A Synthetic Immersive Large-Volume Plenoptic Dataset)
We present SILVR, a dataset of light field images for six-degrees-of-freedom navigation in large fully-immersive volumes.
2 papers · 0 benchmarks
SIMARA (SIMARA: a database for key-value information extraction from full-page handwritten documents)
Description We propose a new database for information extraction from historical handwritten documents.
2 papers · 2 benchmarks
SK-VG is a dataset for Scene Knowledge-guided Visual Grounding, where the image content and referring expressions are not sufficient to ground the target objects, forcing the models to have a reasoning ability on the long-form scene…
2 papers · 0 benchmarks
SKILL-102 (SKILL 102 Lifelong Learning Dataset)
SKILL-102 consists of 102 image classification datasets.
2 papers · 0 benchmarks
SKSF-A consists of seven distinct styles drawn by professional artists.
2 papers · 1 benchmark
SPARF is a large-scale ShapeNet-based synthetic dataset for novel view synthesis consisting of ~17 million images rendered from nearly 40,000 shapes at high resolution (400×400 pixels).
2 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.