Home › Datasets › modality › Images

Images datasets

archive 2025-07-28

3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 29 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Images datasets 1345–1392 of 3,239

GMVD (Generalized Multi-View Detection Dataset)
The GMVD dataset consists of synthetic scenes captured using the GTA-V and Unity graphics engines.
5 papers · 1 benchmark
GOO (Gaze on Objects)
GOO (Gaze-on-Objects) is a dataset for gaze object prediction, where the goal is to predict a bounding box for a person's gazed-at object.
5 papers · 0 benchmarks
HARPER (Exploring 3D Human Pose Estimation and Forecasting from the Robot’s Perspective: The HARPER Dataset)
We introduce HARPER, a novel dataset for 3D body pose estimation and forecast in dyadic interactions between users and \spot, the quadruped robot manufactured by Boston Dynamics.
5 papers · 3 benchmarks
HJDataset is a large dataset of Historical Japanese Documents with Complex Layouts.
5 papers · 0 benchmarks
The Helvipad dataset is a real-world stereo dataset designed for omnidirectional depth estimation.
5 papers · 1 benchmark
HowMany-Qa is a object counting dataset.
5 papers · 1 benchmark
IAM(line-level) (Line-level Handwritten Text Recognition on IAM)
The IAM database contains 13,353 images of handwritten lines of text created by 657 writers.
5 papers · 1 benchmark
IDDA is a large scale, synthetic dataset for semantic segmentation with more than 100 different source visual domains.
5 papers · 0 benchmarks
IG-1B-Targeted is an internal Facebook AI Research dataset that consists of 940 million public images with 1.5K hashtags matching with 1000 ImageNet1K synsets.
5 papers · 0 benchmarks
IIW (Intrinsic Images in the Wild)
Intrinsic Images in the Wild is a large scale, public dataset for intrinsic image decompositions of real-world scenes selected from the OpenSurfaces dataset.
5 papers · 0 benchmarks
The IPN Hand dataset is a benchmark video dataset with sufficient size, variation, and real-world elements able to train and evaluate deep neural networks for continuous Hand Gesture Recognition (HGR).
5 papers · 0 benchmarks
The ISIC 2017 dataset was published by the International Skin Imaging Collaboration (ISIC) as a large-scale dataset of dermoscopy images.
5 papers · 0 benchmarks
ImageNet-Hard is a new benchmark that comprises 10,980 images collected from various existing ImageNet-scale benchmarks (ImageNet, ImageNet-V2, ImageNet-Sketch, ImageNet-C, ImageNet-R, ImageNet-ReaL, ImageNet-A, and ObjectNet).
5 papers · 1 benchmark
ImageNet-Patch: A Dataset for Benchmarking Machine Learning Robustness against Adversarial Patches Adversarial patches are optimized contiguous pixel blocks in an input image that cause a machine-learning model to misclassify it.
5 papers · 0 benchmarks
We collect, organize and open-source the large-scale multimodal instruction dataset, Infinity-MM, consisting of tens of millions of samples.
5 papers · 0 benchmarks
InsPLAD (Inspection Power Line Asset Dataset)
InsPLAD is a Dataset for Power Line Asset Inspection containing 10,607 high-resolution Unmanned Aerial Vehicles colour images.
5 papers · 1 benchmark
A clickthrough prediction dataset, for more information please see the Kaggle page
5 papers · 1 benchmark
KiloGram is a resource for studying abstract visual reasoning in humans and machines.
5 papers · 0 benchmarks
Kvasir-Sessile dataset (Sessile polyps from Kvasir-SEG)
The Kvasir-SEG dataset includes 196 polyps smaller than 10 mm classified as Paris class 1 sessile or Paris class IIa.
5 papers · 0 benchmarks
LaRS (Lakes, Rivers and Seas Dataset)
LaRS is the largest and most diverse panoptic maritime obstacle detection dataset.
5 papers · 3 benchmarks
LabPics (LabPics Dataset for computer vision for autonomous chemistry labs and medical labs)
LabPics Chemistry Dataset Dataset for computer vision for materials segmentation and classification in chemistry labs, medical labs, and any setting where materials are handled inside containers.
5 papers · 0 benchmarks
Lesion Boundary Segmentation Dataset is a dataset for lesion segmentation from the ISIC2018 challenge.
5 papers · 0 benchmarks
LoLi-Phone is a large-scale low-light image and video dataset for Low-light image enhancement (LLIE).
5 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
5 papers · 0 benchmarks
MEDIC is a large social media image classification dataset for humanitarian response consisting of 71,198 images to address four different tasks in a multi-task learning setup.
5 papers · 0 benchmarks
MESSIDOR (MESSIDOR DATABASE)
The Messidor database has been established to facilitate studies on computer-assisted diagnoses of diabetic retinopathy.
5 papers · 0 benchmarks
- A large scale Chinese multi-modal dialogue corpus (120.84K dialogues and 198.82 K images).
5 papers · 0 benchmarks
MMFlood is remote sensing dataset derived from Sentinel-1 (VV-VH), MapZen (DEM) and OpenStreetMap (Hydrography).
5 papers · 1 benchmark
MSAW (Multi-Sensor All Weather Mapping)
Multi-Sensor All Weather Mapping (MSAW) is a dataset and challenge, which features two collection modalities (both SAR and optical).
5 papers · 1 benchmark
MSMT17-C is an evaluation set that consists of algorithmically generated corruptions applied to the MSMT17 test-set.
5 papers · 1 benchmark
MarKG (Multimodal analogical reasoning Knowledge Graph)
The MarKG dataset has 11,292 entities, 192 relations and 76,424 images, including 2,063 analogy entities and 27 analogy relations.
5 papers · 0 benchmarks
The Market1501-Attributes dataset is built from the Market1501 dataset.
5 papers · 1 benchmark
MatSynth MatSynth is a Physically Based Rendering (PBR) materials dataset designed for modern AI applications.
5 papers · 0 benchmarks
A large, realistic multimodal dataset consisting of real personal photos and crowd-sourced questions/answers.
5 papers · 1 benchmark
The Middlebury 2006 is a stereo dataset of indoor scenes with multiple handcrafted layouts.
5 papers · 0 benchmarks
MuMiN is a misinformation graph dataset containing rich social media data (tweets, replies, users, images, articles, hashtags), spanning 21 million tweets belonging to 26 thousand Twitter threads, each of which have been semantically…
5 papers · 0 benchmarks
N-Digit MNIST is a multi-digit MNIST-like dataset.
5 papers · 0 benchmarks
NERDS 360 (NeRF for Reconstruction, Decomposition and Scene Synthesis of 360° outdoor scenes)
We present a large-scale dataset for 3D urban scene understanding.
5 papers · 0 benchmarks
NIH-CXR-LT (Long-tailed (LT) NIH ChestXRay14)
NIH-CXR-LT.
5 papers · 1 benchmark
The NTIRE 2021 HDR was built for the first challenge on high-dynamic range (HDR) imaging that was part of the New Trends in Image Restoration and Enhancement (NTIRE) workshop, held in conjunction with CVPR 2021.
5 papers · 0 benchmarks
The NYU Symmetry database contains 176 single-symmetry and 63 multiple-symmetry images (.png files) with accompanying ground-truth annotations (.mat files).
5 papers · 0 benchmarks
NYU-VP is a new dataset for multi-model fitting, vanishing point (VP) estimation in this case.
5 papers · 0 benchmarks
The ObjectsRoom dataset is based on the MuJoCo environment used by the Generative Query Network [4] and is a multi-object extension of the 3d-shapes dataset.
5 papers · 2 benchmarks
OoDIS (Anomaly Instance Segmentation Benchmark)
OoDIS is a benchmark dataset for anomaly instance segmentation, crucial for autonomous vehicle safety.
5 papers · 2 benchmarks
Open6DOR V2 (Benchmarking Open-instruction 6-DoF Object Rearrangement and A VLM-based Approach)
We introduce a challenging and comprehensive benchmark for open-instruction 6-DoF object rearrangement tasks, termed Open6DOR.
5 papers · 1 benchmark
PGDP5K (Plane Geometry Diagram Parsing Dataset)
PGDP5K is a dataset consisting of 5000 diagram samples composed of 16 shapes, covering 5 positional relations, 22 symbol types and 6 text types, labeled with more fine-grained annotations at primitive level, including primitive classes,…
5 papers · 1 benchmark
Year after year, the demand for ever-better smartphone photos continues to grow, in particular in the domain of portrait photography.
5 papers · 1 benchmark
This dataset contains 114 individuals including 1824 images captured from two disjoint camera views.
5 papers · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.