Home › Datasets › modality › Images
Images datasets
archive 2025-07-28
3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 24 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Images datasets 1105–1152 of 3,239
FloorPlanCAD is a large-scale real-world CAD drawing dataset containing over 15,000 floor plans, ranging from residential to commercial buildings.
8 papers · 0 benchmarks
GPA (Geometric Pose Affordance)
multi-view imagery of people interacting with a variety of rich 3D environments
8 papers · 2 benchmarks
A GQA-based dataset with 1,040,830 multi-modal explanations of visual reasoning processes.
8 papers · 1 benchmark
GTA (A Benchmark for General Tool Agents)
A benchmark to evaluate the tool-use capabilities of LLM-based agents in real-world scenarios.
8 papers · 0 benchmarks
GVFC (Gun Violence Frame Corpus)
This is a new dataset of news headlines and their frames related to the issue of gun violence in the United States.
8 papers · 0 benchmarks
H3WB (Human 3.6M 3D WholeBody)
Human3.6M 3D WholeBody (H3WB) is a large scale dataset with 133 whole-body keypoint annotations on 100K images, made possible by a new multi-view pipeline.
8 papers · 3 benchmarks
HANDAL (HANDAL: A Dataset of Real-World Manipulable Object Categories with Pose Annotations, Affordances, and Reconstructions)
We present the HANDAL dataset for category-level object pose estimation and affordance prediction.
8 papers · 0 benchmarks
The HInt dataset is frequently used as a generalizability benchmark for 3D Hand Reconstruction.
8 papers · 1 benchmark
Contains 446,684 images annotated by humans that cover 43 incidents across a variety of scenes.
8 papers · 0 benchmarks
JSRT (Japanese Society of Radiological Technology Database)
The standard digital image database with and without chest lung nodules (JSRT database) was created(1) by the Japanese Society of Radiological Technology (JSRT) in cooperation with the Japanese Radiological Society (JRS) in 1998.
8 papers · 0 benchmarks
LPW (Labeled Pedestrian in the Wild)
Labeled Pedestrian in the Wild (LPW) is a pedestrian detection dataset that contains 2,731 pedestrians in three different scenes where each annotated identity is captured by from 2 to 4 cameras.
8 papers · 0 benchmarks
LReID is a benchmark for lifelong person reidentification.
8 papers · 0 benchmarks
MeGlass is an eyeglass dataset originally designed for eyeglass face recognition evaluation.
8 papers · 0 benchmarks
The NCT-CRC-HE-100K dataset is a set of 100,000 non-overlapping image patches extracted from 86 H&E stained human cancer tissue slides and normal tissue from the NCT biobank (National Center for Tumor Diseases) and the UMM pathology…
8 papers · 2 benchmarks
NIND (Natural Image Noise Dataset)
An open dataset of real photographs with real noise, from identical scenes captured with varying ISO values.
8 papers · 0 benchmarks
OPRA (Online Product Reviews for Affordances)
The OPRA Dataset was introduced in Demo2Vec: Reasoning Object Affordances From Online Videos (CVPR'18) for reasoning object affordances from online demonstration videos.
8 papers · 2 benchmarks
The Oulu-NPU face presentation attack detection database consists of 4950 real access and attack videos.
8 papers · 1 benchmark
OpenForensics is a large-scale dataset posing a high level of challenges that is designed with face-wise rich annotations explicitly for face forgery detection and segmentation.
8 papers · 0 benchmarks
OpenS2V-Eval introduces 180 prompts from seven major categories of S2V, which incorporate both real and synthetic test data.
8 papers · 1 benchmark
OpenViDial is a large-scale open-domain dialogue dataset with visual contexts.
8 papers · 0 benchmarks
PAD (Purpose-driven Affordance Dataset)
PAD (Purpose-driven Affordance Dataset) is a dataset for affordance detection, which refers to identifying the potential action possibilities of objects in an image, which is an important ability for robot perception and manipulation.
8 papers · 0 benchmarks
RFUND (Revised FUNSD and XFUND)
RFUND is a relabeled version of FUNSD and XFUND datasets, tackling the following issues in their original annotations: 1.
8 papers · 0 benchmarks
RIMES (Reconnaissance & Indexation de données Manuscrites et de fac similÉS / Recognition & Indexing of handwritten documents & faxes)
The RIMES database (Reconnaissance et Indexation de données Manuscrites et de fac similÉS / Recognition and Indexing of handwritten documents and faxes) was created to evaluate automatic systems of recognition and indexing of handwritten…
8 papers · 0 benchmarks
The RIT-18 dataset was built for the semantic segmentation of remote sensing imagery.
8 papers · 0 benchmarks
ROBUST-MIS (Robust Medical Instrument Segmentation Challenge 2019)
The ROBUST-MIS dataset was made available to support the Robust Medical Instrument Segmentation (ROBUST-MIS) Challenge 2019, part of the Endoscopic Vision Challenge associated with MICCAI.
8 papers · 1 benchmark
Over 1.5K images selected from the public Kaggle DR Detection dataset; Five DR grades (DR0 / DR1 / DR2 / DR3 / DR4), re-labeled by a panel of 45 experienced ophthalmologists; Eight retinal lesion classes, including microaneurysm,…
8 papers · 0 benchmarks
This dataset, called RodoSol-ALPR dataset, contains 20,000 images captured by static cameras located at pay tolls owned by the Rodovia do Sol (RodoSol) concessionaire, which operates 67.5 kilometers of a highway (ES-060) in the Brazilian…
8 papers · 0 benchmarks
SI-HDR (Single-image high dynamic range dataset)
The dataset consists of 181 HDR images.
8 papers · 0 benchmarks
SKU110K-R is a dataset relabeled with oriented bounding boxes based on SKU110K.
8 papers · 0 benchmarks
This dataset aims at evaluating the License Plate Character Segmentation (LPCS) problem.
8 papers · 1 benchmark
SciGraphQA is a large-scale, open-domain dataset focused on generating multi-turn conversational question-answering dialogues centered around understanding and describing scientific graphs and figures.
8 papers · 0 benchmarks
Scribble is a new outline dataset consisting of 200 images (150 train, 50 test) for each of 10 classes – basketball, chicken, cookie, cupcake, moon, orange, soccer, strawberry, watermelon and pineapple.
8 papers · 1 benchmark
SenseReID is a person re-identification dataset for evaluating ReID models.
8 papers · 1 benchmark
SoundingEarth consists of co-located aerial imagery and audio samples all around the world.
8 papers · 1 benchmark
Autonomous trucking is a promising technology that can greatly impact modern logistics and the environment.
8 papers · 1 benchmark
Histopathological characterization of colorectal polyps allows to tailor patients' management and follow up with the ultimate aim of avoiding or promptly detecting an invasive carcinoma.
8 papers · 0 benchmarks
The dataset is maintained by VISION AND IMAGE PROCESSING LAB, University of Waterloo.
8 papers · 3 benchmarks
Similar to CVUSA and CVACT, the VIGOR dataset contains satellites and street imagery to match them to each other to find the location of the street imagery.
8 papers · 3 benchmarks
Vision-language generative reward models (VL-GenRMs) play a crucial role in aligning and evaluating multimodal AI systems, yet their own evaluation remains under-explored.
8 papers · 0 benchmarks
VRAI (Vehicle Re-identification for Aerial Image)
VRAI is a large-scale vehicle ReID dataset for UAV-based intelligent applications.
8 papers · 2 benchmarks
The WebUI dataset contains 400K web UIs captured over a period of 3 months and cost about $500 to crawl.
8 papers · 0 benchmarks
Who's Waldo is a dataset of 270K image–caption pairs, depicting interactions of people, that is automatically mined from Wikimedia Commons.
8 papers · 1 benchmark
WildReceipt is a collection of receipts.
8 papers · 0 benchmarks
WorldStrat (The WorldStrat Dataset: Open High-Resolution Satellite Imagery With Paired Multi-Temporal Low-Resolution)
Nearly 10,000 km² of free high-resolution and paired multi-temporal low-resolution satellite imagery of unique locations which ensure stratified representation of all types of land-use across the world: from agriculture to ice caps, from…
8 papers · 0 benchmarks
ZEB (Zero-shot Evaluation Benchmark)
A evaluation benchmark ZEB for image matching by merging 8 real-world datasets and 4 simulated datasets with diverse image resolutions, scene conditions and view points.
8 papers · 1 benchmark
Research on semantic segmentation of traffic scenes using color and polarization information (including training and testing sets).
8 papers · 1 benchmark
3DOH50K is the first real 3D human dataset for the problem of human reconstruction and pose estimation in occlusion scenarios.
7 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.