Home › Datasets › modality › Images
Images datasets
archive 2025-07-28
3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 22 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Images datasets 1009–1056 of 3,239
Description Detection Dataset (D³, /dikju:b/) is an attempt at creating a next-generation object detection dataset.
10 papers · 1 benchmark
FERET-Morphs is a dataset of morphed faces selected from the publicly available FERET dataset [1].
10 papers · 0 benchmarks
FM-IQA (Freestyle Multilingual Image Question Answering)
FM-IQA is a question-answering dataset containing over 150,000 images and 310,000 freestyle Chinese question-answer pairs and their English translations.
10 papers · 0 benchmarks
FunnyBirds is a synthetic vision dataset that is developed to automatically and quantitatively analyze XAI methods.
10 papers · 0 benchmarks
GRAZPEDWRI-DX is a public dataset of 20,327 pediatric wrist trauma X-ray images released by the University of Medicine of Graz.
10 papers · 4 benchmarks
The HO-3D v3 is the version 3 of the HO-3D dataset with more accurate hand-object poses.
10 papers · 1 benchmark
Dataset Introduction In this work, we introduce the In-Diagram Logic (InDL) dataset, an innovative resource crafted to rigorously evaluate the logic interpretation abilities of deep learning models.
10 papers · 1 benchmark
Kuzushiji-49 is an MNIST-like dataset that has 49 classes (28x28 grayscale, 270,912 images) from 48 Hiragana characters and one Hiragana iteration mark.
10 papers · 0 benchmarks
LAD (Large-scale Attribute Dataset)
LAD (Large-scale Attribute Dataset) has 78,017 images of 5 super-classes and 230 classes.
10 papers · 0 benchmarks
MP-DocVQA (Multipage Document Visual Question Answering)
The dataset is aimed to perform Visual Question Answering on multipage industry scanned documents.
10 papers · 0 benchmarks
A large-scale dataset that consists of 21,184 claims, where each claim is assigned a truthfulness label and ruling statement, with 58,523 pieces of evidence in the form of text and images.
10 papers · 0 benchmarks
NAF (National Archives Forms Dataset)
This dataset was created with images provided by the United States National Archive and FamilySearch.
10 papers · 0 benchmarks
Office-Caltech-10 a standard benchmark for domain adaptation, which consists of Office 10 and Caltech 10 datasets.
10 papers · 1 benchmark
PointQA is a set of datasets for Visual Question Datasets (VQA) that require a pointer to an object in the image to be answered correctly.
10 papers · 0 benchmarks
REFLACX (Reports and eye-tracking data for localization of abnormalities in chest x-rays)
The REFLACX dataset contains eye-tracking data for 3,032 readings of chest x-rays by five radiologists.
10 papers · 0 benchmarks
RVSD (Realistic Video DeSnowing Dataset)
Realistic Video DeSnowing Dataset (RVSD) contains a total of 110 pairs of videos.
10 papers · 0 benchmarks
ReaSCAN (ReaSCAN: Compositional Reasoning in Language Grounding)
ReaSCAN is a synthetic navigation task that requires models to reason about surroundings over syntactically difficult languages.
10 papers · 0 benchmarks
SK-LARGE is a benchmark dataset for object skeleton detection, built on the MS COCO dataset.
10 papers · 1 benchmark
SpaceNet 2: Building Detection v2 - is a dataset for building footprint detection in geographically diverse settings from very high resolution satellite images.
10 papers · 1 benchmark
A large-scale human image dataset with over 230K samples capturing diverse poses and textures.
10 papers · 0 benchmarks
TMED (Tufts Medical Echocardiogram Dataset)
TMED is a clinically-motivated benchmark dataset for computer vision and machine learning from limited labeled data.
10 papers · 0 benchmarks
USF (Human ID Gait Challenge Dataset)
The USF Human ID Gait Challenge Dataset is a dataset of videos for gait recognition.
10 papers · 0 benchmarks
ViP-Bench (Making Large Multimodal Models Understand Arbitrary Visual Prompts)
ViP-Bench is a comprehensive benchmark designed to assess the capability of multimodal models in understanding visual prompts across multiple dimensions.
10 papers · 1 benchmark
e-ViL is a benchmark for explainable vision-language tasks.
10 papers · 0 benchmarks
Annotated using images taken by a drone in 501 separate flights, totalling in over 62 hours of trajectory data.
10 papers · 0 benchmarks
xR-EgoPose is an egocentric synthetic dataset for egocentric 3D human pose estimation.
10 papers · 0 benchmarks
Our dataset which consists of multiple indoor and outdoor experiments for up to 30 m gNB-UE link.
9 papers · 0 benchmarks
Adaptiope is a domain adaptation dataset with 123 classes in the three domains synthetic, product and real life.
9 papers · 0 benchmarks
AmsterTime (AmsterTime: A Visual Place Recognition Benchmark Dataset for Severe Domain Shift)
AmsterTime dataset offers a collection of 2,500 well-curated images matching the same scene from a street view matched to historical archival image data from Amsterdam city.
9 papers · 3 benchmarks
Atari-HEAD is a dataset of human actions and eye movements recorded while playing Atari videos games.
9 papers · 0 benchmarks
A high-resolution semantic segmentation dataset with 50 validation and 100 test objects.
9 papers · 1 benchmark
BIMCV-COVID19+ dataset is a large dataset with chest X-ray images CXR (CR, DX) and computed tomography (CT) imaging of COVID-19 patients along with their radiographic findings, pathologies, polymerase chain reaction (PCR), immunoglobulin G…
9 papers · 0 benchmarks
This dataset contains 1200 images (1000 WLI images and 200 FICE images) with fine-grained segmentation annotations.
9 papers · 1 benchmark
BLVD is a large scale 5D semantics dataset collected by the Visual Cognitive Computing and Intelligent Vehicles Lab.
9 papers · 0 benchmarks
CUHK03-C is an evaluation set that consists of algorithmically generated corruptions applied to the CUHK03 test-set.
9 papers · 1 benchmark
CaDIS (Cataract Dataset for Image Segmentation)
CaDIS: a Cataract Dataset for Image Segmentation is a dataset for semantic segmentation created by Digital Surgery Ltd.
9 papers · 1 benchmark
ClueWeb22 is the newest iteration of the ClueWeb line of datasets, provides 10 billion web pages affiliated with rich information.
9 papers · 0 benchmarks
EgoCap is a dataest of 100,000 egocentric images of eight people in different clothing, with 75,000 images from six people used for training.
9 papers · 0 benchmarks
EgoHOS (Fine-Grained Egocentric Hand-Object Segmentation Dataset)
EgoHOS is a labeled dataset consisting of 11243 egocentric images with per-pixel segmentation labels of hands and objects being interacted with during a diverse array of daily activities.
9 papers · 0 benchmarks
The EntitySeg dataset contains 33,227 images with high-quality mask annotations.
9 papers · 0 benchmarks
A large publicly available retinal fundus image dataset for glaucoma classification called G1020.
9 papers · 0 benchmarks
Global WHEAT Dataset is the first large-scale dataset for wheat head detection from field optical images.
9 papers · 0 benchmarks
This is a gun detection dataset with 51K annotated gun images for gun detection and other 51K cropped gun chip images for gun classification collected from a few different sources.
9 papers · 6 benchmarks
HQ-WMCA (High-Quality Wide Multi-Channel Attack database)
The High-Quality Wide Multi-Channel Attack database (HQ-WMCA) database consists of 2904 short multi-modal video recordings of both bona-fide and presentation attacks.
9 papers · 0 benchmarks
HUMAN4D is a large and multimodal 4D dataset that contains a variety of human activities simultaneously captured by a professional marker-based MoCap, a volumetric capture and an audio recording system.
9 papers · 0 benchmarks
Imagewoof is a subset of 10 dog breed classes from Imagenet.
9 papers · 0 benchmarks
LoDoPaB-CT is a dataset of computed tomography images and simulated low-dose measurements.
9 papers · 1 benchmark
MEIR (Multimodal Entity Image Repurposing)
MEIR is a substantially challenging dataset over that which has been previously available to support research into image repurposing detection.
9 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.