Home › Datasets › modality › Images

Images datasets

archive 2025-07-28

3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 33 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Images datasets 1537–1584 of 3,239

For each dataset we provide a short description as well as some characterization metrics.
4 papers · 0 benchmarks
We introduce MultiScan, a scalable RGBD dataset construction pipeline leveraging commodity mobile devices to scan indoor scenes with articulated objects and web-based semantic annotation interfaces to efficiently annotate object and part…
4 papers · 1 benchmark
MultiSense is a dataset of 9,504 images annotated with an English verb and its translation in Spanish and German.
4 papers · 0 benchmarks
MultiSubs (MultiSubs: A Large-scale Multimodal and Multilingual Dataset)
MultiSubs is a dataset of multilingual subtitles gathered from the OPUS OpenSubtitles dataset, which in turn was sourced from opensubtitles.org.
4 papers · 5 benchmarks
N-Omniglot is a neuromorphic dataset for few-shot learning.
4 papers · 0 benchmarks
NHA12D (A New Pavement Crack Dataset)
NHA12D is an annotated pavement crack dataset that contains images with different viewpoints and pavements types.
4 papers · 0 benchmarks
This dataset is an OSN-transmitted (Online Social Network) version of the NIST dataset (https://www.nist.gov/itl/iad/mig/nimble-challenge-2017-evaluation).
4 papers · 1 benchmark
This dataset is an OSN-transmitted (Online Social Network) version of the NIST dataset (https://www.nist.gov/itl/iad/mig/nimble-challenge-2017-evaluation).
4 papers · 1 benchmark
This dataset is an OSN-transmitted (Online Social Network) version of the NIST dataset (https://www.nist.gov/itl/iad/mig/nimble-challenge-2017-evaluation).
4 papers · 1 benchmark
The nordland used in SALAD and BoQ (2760 queries, 27592 reference images, threshold: 1 frames).
4 papers · 1 benchmark
O3 (Odd-One-Out Dataset)
A set of realistic odd-one-out stimuli gathered "in the wild".
4 papers · 0 benchmarks
OCD (Out-of-Context Dataset)
OCD (Out-of-Context Dataset) is a synthetic dataset with fine-grained control over scene context.
4 papers · 0 benchmarks
OVDEval includes 9 sub-tasks and introduces evaluations on commonsense knowledge, attribute understanding, position understanding, object relation comprehension, and more.
4 papers · 0 benchmarks
OVQA contains 19,020 medical visual question and answer pairs generated from 2,001 medical images collected from 2,212 EMRs in Orthopedics.
4 papers · 0 benchmarks
OpenViVQA (Open-domain Visual Question Answering in Vietnamese)
In recent years, visual question answering (VQA) has attracted attention from the research community because of its highly potential applications (such as virtual assistance on intelligent cars, assistant devices for blind people, or…
4 papers · 0 benchmarks
Oracle-MNIST (Oracle-MNIST: a Realistic Image Dataset for Benchmarking Machine Learning Algorithms)
We introduce the Oracle-MNIST dataset, comprising of 2828 grayscale images of 30,222 ancient characters from 10 categories, for benchmarking pattern classification, with particular challenges on image noise and distortion.
4 papers · 1 benchmark
PADv2 (Purpose-driven Affordance Dataset v2)
With complex scenes and rich annotations, the PADv2 dataset can be used as a test bed to benchmark affordance detection methods and may also facilitate downstream vision tasks, such as scene understanding, action recognition, and robot…
4 papers · 0 benchmarks
PARIS Dataset (PARIS Two-Part Object Dataset)
From PARIS: Part-level Reconstruction and Motion Analysis for Articulated Objects: 5.1.
4 papers · 0 benchmarks
PDEBench provides a diverse and comprehensive set of benchmarks for scientific machine learning, including challenging and realistic physical problems.
4 papers · 0 benchmarks
PDS-COCO (Photometrically Distorted Synthetic COCO)
Photometrically Distorted Synthetic COCO (PDS-COCO) dataset is a synthetically created dataset for homography estimation learning.
4 papers · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
4 papers · 1 benchmark
PeopleSansPeople (PeopleSansPeople: A Synthetic Data Generator for Human-Centric Computer Vision)
In recent years, person detection and human pose estimation have made great strides, helped by large-scale labeled datasets.
4 papers · 0 benchmarks
Polaris (Polaris dataset)
The Polaris dataset offers a large-scale, diverse benchmark for evaluating metrics for image captioning, surpassing existing datasets in terms of size, caption diversity, number of human judgments, and granularity of the evaluations.
4 papers · 0 benchmarks
PsyMo (PsyMo: A Dataset for Estimating Self-Reported Psychological Traits from Gait)
Psychological trait estimation from external factors such as movement and appearance is a challenging and long-standing problem in psychology, and is principally based on the psychological theory of embodiment.
4 papers · 0 benchmarks
RDD-2020 (Road Damage Dataset 2020)
The Road Damage Dataset 2020 (RDD-2020) Secondly is a large-scale heterogeneous dataset comprising 26620 images collected from multiple countries using smartphones.
4 papers · 0 benchmarks
RED (Real Embodied Dataset)
The Real Embodied Dataset (RED) is a computer vision large-scale dataset for grasping in cluttered scenes.
4 papers · 0 benchmarks
RegDB-C is an evaluation set that consists of algorithmically generated corruptions applied to the RegDB test-set (color images).
4 papers · 0 benchmarks
Relative Human (RH) contains multi-person in-the-wild RGB images with rich human annotations, including: Depth layers: relative depth relationship/ordering between all people in the image.
4 papers · 2 benchmarks
RetVQA (Retrieval-Based Visual Question Answering)
The RetVQA dataset is a large-scale dataset designed for Retrieval-Based Visual Question Answering (RetVQA).
4 papers · 1 benchmark
RidgeBase (RidgeBase: A Cross-Sensor Multi-Finger Contactless Fingerprint Dataset)
Contactless fingerprint matching using smartphone cameras can alleviate major challenges of traditional fingerprint systems including hygienic acquisition, portability and presentation attacks.
4 papers · 0 benchmarks
S2TLD (SJTU Small Traffic Light Dataset)
S2TLD is a traffic light dataset, which contains 5,786 images of approximately 1,080 1,920 pixels and 720 1,280 pixels.
4 papers · 0 benchmarks
S3E is a novel large-scale multimodal dataset captured by a fleet of unmanned ground vehicles along four designed collaborative trajectory paradigms.
4 papers · 0 benchmarks
The Situated Corpus Of Understanding Transactions (SCOUT) is a multi-modal collection of human-robot dialogue in the task domain of collaborative exploration.
4 papers · 0 benchmarks
Machine-learning Data Set Prepared from NASA Solar Dynamics Observatory Mission data.
4 papers · 0 benchmarks
SOD4SB (Small Object Detection for Spotting Birds)
The Small Object Detection for Spotting Birds (SOD4SB) dataset is a dataset consisting of 39,070 images including 137,121 bird instances.
4 papers · 2 benchmarks
STAIR Captions is a large-scale dataset containing 820,310 Japanese captions.
4 papers · 0 benchmarks
SWORD ('Scenes with occluded regions' dataset)
The new dataset contains around 1,500 train videos and 290 test videos, with 50 frames per video on average.
4 papers · 1 benchmark
SYSU-MM01-C is an evaluation set that consists of algorithmically generated corruptions applied to the SYSU-MM01 test-set.
4 papers · 1 benchmark
ScenicOrNot (SoN) is a dataset of 185,548 images with associated natural beauty rating histograms.
4 papers · 0 benchmarks
Separated COCO is automatically generated subsets of COCO val dataset, collecting separated objects for a large variety of categories in real images in a scalable manner, where target object segmentation mask is separated into distinct…
4 papers · 1 benchmark
ShapenetRenderer is an extension of the ShapeNet Core dataset which has more variation in camera angles.
4 papers · 0 benchmarks
SportsPose (SportsPose - A Dynamic 3D sports pose dataset)
Accurate 3D human pose estimation is essential for sports analytics, coaching, and injury prevention.
4 papers · 0 benchmarks
Synthinel-1 is a collection of synthetic overhead imagery with full pixel-wise building segmentation labels.
4 papers · 0 benchmarks
TCG (Traffic Control Gesture)
The TCG dataset is used to evaluate Traffic Control Gesture recognition for autonomous driving.
4 papers · 1 benchmark
TexBiG (Text-Bild-Gefüge)
TexBiG (from the German Text-Bild-Gefüge, meaning Text-Image-Structure) is a document layout analysis dataset for historical documents in the late 19th and early 20th century.
4 papers · 4 benchmarks
The TimberSeg 1.0 dataset is composed of 220 images showing wood logs in various environments and conditions in Canada.
4 papers · 0 benchmarks
Topo-boundary is a new benchmark dataset, named \textit{Topo-boundary}, for off-line topological road-boundary detection.
4 papers · 0 benchmarks
USIS10K (Large-scale Underwater Salient Instance Segmentation Dataset)
We construct the first large-scale dataset, USIS10K, for the underwater salient instance segmentation task, which contains 10,632 images and pixel-level annotations of 7 categories.
4 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.