Home › Datasets › modality › Images

Images datasets

archive 2025-07-28

3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 26 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Images datasets 1201–1248 of 3,239

Building footprints are useful for a range of important applications, from population estimation, urban planning and humanitarian response, to environmental and climate science.
7 papers · 0 benchmarks
OpenEDS2020 is a dataset of eye-image sequences captured at a frame rate of 100 Hz under controlled illumination, using a virtual-reality head-mounted display mounted with two synchronized eye-facing cameras.
7 papers · 0 benchmarks
The Oxford-Affine dataset is a small dataset containing 8 scenes with sequence of 6 images per scene.
7 papers · 0 benchmarks
PFN-PIC (PFN Picking Instructions for Commodities Dataset)
This dataset is a collection of spoken language instructions for a robotic system to pick and place common objects.
7 papers · 0 benchmarks
PIDray is a large-scale dataset which covers various cases in real-world scenarios for prohibited item detection, especially for deliberately hidden items.
7 papers · 0 benchmarks
PROBA-V (PROBA-V Super-Resolution dataset)
The PROBA-V Super-Resolution dataset is the official dataset of ESA's Kelvins competition for "PROBA-V Super Resolution".
7 papers · 1 benchmark
The PS-Battles dataset is gathered from a large community of image manipulation enthusiasts and provides a basis for media derivation and manipulation detection in the visual domain.
7 papers · 0 benchmarks
Panoptic nuScenes is a benchmark dataset that extends the popular nuScenes dataset with point-wise groundtruth annotations for semantic segmentation, panoptic segmentation, and panoptic tracking tasks.
7 papers · 0 benchmarks
RCooper (Roadside Cooperative Perception Dataset)
The first real-world, large-scale Roadside Cooperative Perception Dataset, RCooper, is released to bloom research on roadside cooperative perception for practical applications.
7 papers · 0 benchmarks
Understanding spatial relations (e.g., “laptop on table”) in visual input is important for both humans and robots.
7 papers · 1 benchmark
The Retrieval-SFM dataset is used for instance image retrieval.
7 papers · 0 benchmarks
S-COCO (Synthetic COCO)
Synthetic COCO (S-COCO) is a synthetically created dataset for homography estimation learning.
7 papers · 1 benchmark
English subset of the SLAKE dataset, comprising 642 images and more than 7,000 question–answer pairs.
7 papers · 0 benchmarks
SME (Standard Multimodal Explanation)
SME is a new dataset for Multi-modal Explanation for Visual Question Answering comprising 1,028,230 samples, with 1,656 visual objects requiring detection in explanations.
7 papers · 1 benchmark
SNARE, short for ShapeNet Annotated with Referring Expressions, is a benchmark requires a model to choose which of two objects is being referenced by a natural language description.
7 papers · 0 benchmarks
SODA-A is a large-scale benchmark specialized for small object detection task under aerial scenes, which has 800203 instances with oriented rectangle box annotation across 9 classes.
7 papers · 0 benchmarks
A multimodal agent benchmark on professional data science and engineering.
7 papers · 0 benchmarks
The Stanford Light Field Archive is a collection of several light fields for research in computer graphics and vision.
7 papers · 0 benchmarks
The Sunnybrook Cardiac Data (SCD), also known as the 2009 Cardiac MR Left Ventricle Segmentation Challenge data, consist of 45 cine-MRI images from a mixed of patients and pathologies: healthy, hypertrophy, heart failure with infarction…
7 papers · 0 benchmarks
The SynthHands dataset is a dataset for hand pose estimation which consists of real captured hand motion retargeted to a virtual hand with natural backgrounds and interactions with different objects.
7 papers · 0 benchmarks
THEODORE (Learning from THEODORE)
Recent work about synthetic indoor datasets from perspective views has shown significant improvements of object detection results with Convolutional Neural Networks(CNNs).
7 papers · 0 benchmarks
The TUM Kitchen dataset is an action recognition dataset that contains 20 video sequences captured by 4 cameras with overlapping views.
7 papers · 0 benchmarks
UBI-Fights (Abnormal Event Detection Dataset)
UBI-Fights - Concerning a specific anomaly detection and still providing a wide diversity in fighting scenarios, the UBI-Fights dataset is a unique new large-scale dataset of 80 hours of video fully annotated at the frame level.
7 papers · 2 benchmarks
UPAR (Unified Pedestrian Attribute Recognition)
The Task: The challenge will use an extension of the UPAR Dataset [1], which consists of images of pedestrians annotated for 40 binary attributes.
7 papers · 1 benchmark
UPIQ (Unified Photometric Image Quality)
Contains over 4,000 images created by realigning and merging existing HDR and standard-dynamic-range (SDR) datasets.
7 papers · 0 benchmarks
VNHSGE (VietNamese High School Graduation Examination Dataset for Large Language Models)
The VNHSGE (VietNamese High School Graduation Examination) dataset, developed exclusively for evaluating large language models (LLMs), is introduced in this article.
7 papers · 9 benchmarks
VQA-CE (VQA Counterexamples)
This dataset provides a new split of VQA v2 (similarly to VQA-CP v2), which is built of questions that are hard to answer for biased models.
7 papers · 1 benchmark
WALT (Watch and Learn TimeLapse Images)
We introduce a new dataset, Watch and Learn Time-lapse (WALT), consisting of multiple (4K and 1080p) cameras capturing urban environments over a year.
7 papers · 1 benchmark
WebFG-496 is a dataset for fine-grained recognition that contains 200 subcategories of the "Bird" (Web-bird), 100 subcategories of the Aircraft" (Web-aircraft), and 196 subcategories of the "Car" (Web-car).
7 papers · 0 benchmarks
YCB-Slide (YCB-Slide: A tactile interaction dataset)
The YCB-Slide dataset comprises of DIGIT sliding interactions on YCB objects.
7 papers · 0 benchmarks
YUD+ (Additional Vanishing Point Labels for the York Urban Database)
YUD+ is a dataset containing additional Vanishing Point Labels for the York Urban Database.
7 papers · 0 benchmarks
iPhone dataset is a challenging benchmarks for dynamic reconstruction.
7 papers · 1 benchmark
The m2cai16-tool-locations dataset contains spatial tool annotations for 2,532 frames across the first 10 videos in the m2cai16-tool dataset, which includes 15 videos in total.
7 papers · 0 benchmarks
A Modular Simulation Framework and Benchmark for Robot Learning.
7 papers · 0 benchmarks
2D HeLa is a dataset of fluorescence microscopy images of HeLa cells stained with various organelle-specific fluorescent dyes.
6 papers · 0 benchmarks
ADVANCE (AuDio Visual Aerial sceNe reCognition datasEt)
The AuDio Visual Aerial sceNe reCognition datasEt (ADVANCE) is a brand-new multimodal learning dataset, which aims to explore the contribution of both audio and conventional visual messages to scene recognition.
6 papers · 0 benchmarks
AH36M (Ambiguous Human3.6M)
Since H36M is captured in a controlled environment, it rarely depicts challenging real-world scenarios such as body occlusions that are the main source of ambiguity in the single-view 3D shape estimation problem.
6 papers · 1 benchmark
AO-CLEVr is a new synthetic-images dataset containing images of "easy" Attribute-Object categories, based on the CLEVr.
6 papers · 0 benchmarks
ARCADE (Automatic Region-based Coronary Artery Disease diagnostics using x-ray angiography imagEs Dataset)
ARCADE: Automatic Region-based Coronary Artery Disease diagnostics using x-ray angiography imagEs Dataset Phase 2 consist of two folders with 300 images in each of them as well as annotations.
6 papers · 0 benchmarks
AcinoSet is a dataset of free-running cheetahs in the wild that contains 119,490 frames of multi-view synchronized high-speed video footage, camera calibration files and 7,588 human-annotated frames.
6 papers · 0 benchmarks
BAAI-VANJEE is a dataset for benchmarking and training various computer vision tasks such as 2D/3D object detection and multi-sensor fusion.
6 papers · 0 benchmarks
CMD is a publicly available collection of hundreds of thousands 2D maps and 3D grids containing different properties of the gas, dark matter, and stars from more than 2,000 different universes.
6 papers · 0 benchmarks
CITE is a crowd-sourced resource for multimodal discourse: this resource characterises inferences in image-text contexts in the domain of cooking recipes in the form of coherence relations.
6 papers · 1 benchmark
CeyMo is a novel benchmark dataset for road marking detection which covers a wide variety of challenging urban, sub-urban and rural road scenarios.
6 papers · 1 benchmark
ChineseFoodNet aims to automatically recognizing pictured Chinese dishes.
6 papers · 0 benchmarks
This is a synthetic dataset for defect detection on textured surfaces.
6 papers · 1 benchmark
DRTiD is a benchmark dataset for DR grading, consisting of 3,100 two-field fundus images.
6 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.