Home › Datasets › modality › Images
Images datasets
archive 2025-07-28
3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 37 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Images datasets 1729–1776 of 3,239
PDFVQA: A New Dataset for Real-World VQA on PDF Documents
3 papers · 0 benchmarks
PETRAW (PEg TRAnsfer Workflow recognition by different modalities)
PETRAW data set was composed of 150 sequences of peg transfer training sessions.
3 papers · 6 benchmarks
PIC (Person In Context 2021)
The Person In Context (PIC) dataset is a dataset for human-centric relation segmentation (HRS), which contains 17,122 high-resolution images and densely annotated entity segmentation and relations, including 141 object categories, 23…
3 papers · 0 benchmarks
A large-scale video portrait dataset that contains 291 videos from 23 conference scenes with 14K frames.
3 papers · 0 benchmarks
We collect a total of 13,380 images captured on 2,210 different scenes.
3 papers · 0 benchmarks
Real-world dataset of ~400 images of cuboid-shaped parcels with full 2D and 3D annotations in the COCO format.
3 papers · 0 benchmarks
Patzig contains handwritten texts written in modern German.
3 papers · 0 benchmarks
Pavia Centre is a hyperspectral dataset acquired by the ROSIS sensor during a flight campaign over Pavia, northern Italy.
3 papers · 0 benchmarks
Modeling what makes an advertisement persuasive, i.e., eliciting the desired response from consumer, is critical to the study of propaganda, social psychology, and marketing.
3 papers · 0 benchmarks
The PhotoSynth (PS) dataset for patch matching consists of a total of 30 scenes with 25 scenes for training and 5 scenes for validation.
3 papers · 0 benchmarks
PACS (Physical Audiovisual CommonSense) is the first audiovisual benchmark annotated for physical commonsense attributes.
3 papers · 1 benchmark
PoPArt (Poses of People in Art: A Data Set for Human Pose Estimation in Digital Art History)
Throughout the history of art, the pose—as the holistic abstraction of the human body's expression—has proven to be a constant in numerous studies.
3 papers · 1 benchmark
Polyps in the colon are widely known cancer precursors identified by colonoscopy.
3 papers · 1 benchmark
Large multimodal models extend the impressive capabilities of large language models by integrating multimodal understanding abilities.
3 papers · 0 benchmarks
QST contains 1,167 video clips that are cut out from 216 time-lapse 4K videos collected from YouTube, which can be used for a variety of tasks, such as (high-resolution) video generation, (high-resolution) video prediction,…
3 papers · 0 benchmarks
This paper introduces the RGB Arabic Alphabet Sign Language (AASL) dataset.
3 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
3 papers · 0 benchmarks
ROOR is a reading order prediction (ROP) benchmark which annotates layout reading order as ordering relations.
3 papers · 1 benchmark
A new large-scale retail product dataset for fine-grained image classification.
3 papers · 0 benchmarks
We manually labelled 3359 images from the RWTH-PHOENIX-Weather 2014 Development set.
3 papers · 1 benchmark
The RailEye3D dataset, a collection of train-platform scenarios for applications targeting passenger safety and automation of train dispatching, consists of 10 image sequences captured at 6 railway stations in Austria.
3 papers · 0 benchmarks
Dataset Overview This dataset contains individual-level data from a randomized controlled trial (RCT) conducted in northern Uganda, along with associated satellite imagery.
3 papers · 0 benchmarks
Ricordi contains handwritten texts written in Italian.
3 papers · 0 benchmarks
RobotPush is a dataset for object singulation – the task of separating cluttered objects through physical interaction.
3 papers · 0 benchmarks
S-VED (Sacrobosco Visual Element Dataset)
The Sacrobosco Visual Elements Dataset (S-VED) is derived from 359 Sphaera editions, centered on the Tractatus de sphaera by Johannes de Sacrobosco (—1256) and printed between 1472 and 1650.
3 papers · 0 benchmarks
SCapRepo (Google Play Screenshot Caption)
A screenshot-caption dataset containing 135k pairs of screenshots and captions extracted from Google Play.
3 papers · 0 benchmarks
> High-quality underwater coral detection dataset for machine learning and computer vision research.
3 papers · 1 benchmark
SDD dataset contains a variety of indoor and outdoor scenes, designed for Image Defocus Deblurring.
3 papers · 1 benchmark
SDN (Situated Dialogue Navigation)
Situated Dialogue Navigation (SDN) is a navigation benchmark of 183 trials with a total of 8415 utterances, around 18.7 hours of control streams, and 2.9 hours of trimmed audio.
3 papers · 0 benchmarks
SHIFT15M is a dataset that can be used to properly evaluate models in situations where the distribution of data changes between training and testing.
3 papers · 0 benchmarks
SIDD-Image (Segmented Intrusion Detection Dataset)
This is the first image-based network intrusion detection dataset.
3 papers · 1 benchmark
SMILE-UHURA (Small Vessel Segmentation at Mesoscopic Scale from Ultra-High Resolution 7T Magnetic Resonance Angiogram)
The human brain receives nutrients and oxygen through an intricate network of blood vessels.
3 papers · 0 benchmarks
This dataset is proposed in the ICLR 2024 paper: Measuring Vision-Language STEM Skills of Neural Models.
3 papers · 0 benchmarks
SYNTH-PEDES is a large-scale person dataset with image-text pairs by far, which contains 312,321 identities, 4,791,711 images, and 12,138,157 textual descriptions.
3 papers · 0 benchmarks
The San Francisco Landmark Dataset contains a database of 1.7 million images of buildings in San Francisco with ground truth labels, geotags, and calibration data, as well as a difficult query set of 803 cell phone images taken with a…
3 papers · 1 benchmark
Provides a set of stereo-rectified images and the associated groundtruthed disparities for 10 AOIs (Area of Interest) drawn from two sources: 8 AOIs from IARPA's MVS Challenge dataset and 2 AOIs from the CORE3D-Public dataset.
3 papers · 0 benchmarks
Schiller contains handwritten texts written in modern German.
3 papers · 0 benchmarks
Schwerin contains handwritten texts written in medieval German.
3 papers · 0 benchmarks
SeaTurtleID is a public large-scale, long-span dataset with sea turtle photographs captured in the wild.
3 papers · 0 benchmarks
Given the unavailability of real-world pharmaceutical inspection-domain datasets, we have created the Sensum Solid Oral Dosage Forms (SensumSODF) dataset intended for research and evaluation purposes.
3 papers · 0 benchmarks
SiW (Spoofing in the Wild) is a face anti-spoofing dataset recently introduced in [29] where images are extracted from short videos captured at high resolution and 30 frames per second.
3 papers · 1 benchmark
SketchHairSalon is a dataset for hair generation containing thousands of annotated hair sketch-image pairs and corresponding hair mattes.
3 papers · 0 benchmarks
SkinCon is a skin disease dataset densely annotated by dermatologists.
3 papers · 0 benchmarks
The Specs on Faces (SoF) dataset, a collection of 42,592 (2,662×16) images for 112 persons (66 males and 46 females) who wear glasses under different illumination conditions.
3 papers · 0 benchmarks
Social Relation Dataset is a dataset for social relation trait prediction from face images.
3 papers · 0 benchmarks
SpaGBOL (Spatial-Graph-Based Orientated Cross-View Geo-Localisation)
Cross-View Geo-Localisation within urban regions is challenging in part due to the lack of geo-spatial structuring within current datasets and techniques.
3 papers · 1 benchmark
Introduction This dataset supports Ye et al.
3 papers · 0 benchmarks
Synthehicle is a massive CARLA-based synthehic multi-vehicle multi-camera tracking dataset and includes ground truth for 2D detection and tracking, 3D detection and tracking, depth estimation, and semantic, instance and panoptic…
3 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.