Home › Datasets › modality › Images
Images datasets
archive 2025-07-28
3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 12 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Images datasets 529–576 of 3,239
WORD (Whole abdominal Organs Dataset)
WORD is a dataset for organ semantic segmentation that contains 150 abdominal CT volumes (30,495 slices) and each volume has 16 organs with fine pixel-level annotations and scribble-based sparse annotation, which may be the largest dataset…
38 papers · 0 benchmarks
The A3D dataset is a step forward to make autonomous driving safer for pedestrians and the public in the real world.
37 papers · 0 benchmarks
CV-Bench (Cambrian Vision-Centric Benchmark)
The Cambrian Vision-Centric Benchmark (CV-Bench) is designed to address the limitations of existing vision-centric benchmarks by providing a comprehensive evaluation framework for multimodal large language models (MLLMs).
37 papers · 0 benchmarks
The EgoGesture dataset contains 2,081 RGB-D videos, 24,161 gesture samples and 2,953,224 frames from 50 distinct subjects.
37 papers · 2 benchmarks
HumanAct12 is a new 3D human motion dataset adopted from the polar image and 3D pose dataset PHSPD, with proper temporal cropping and action annotating.
37 papers · 2 benchmarks
Learn2Reg is a dataset for medical image registration.
37 papers · 2 benchmarks
Open Images V4 offers large scale across several dimensions: 30.1M image-level labels for 19.8k concepts, 15.4M bounding boxes for 600 object classes, and 375k visual relationship annotations involving 57 classes.
37 papers · 1 benchmark
PIPAL (Perceptual Image Processing ALgorithms IQA Dataset)
PIPAL training set contains 200 reference images, 40 distortion types, 23k distortion images, and more than one million human ratings.
37 papers · 0 benchmarks
TextOCR is a dataset to benchmark text recognition on arbitrary shaped scene-text.
37 papers · 0 benchmarks
BRATS 2013 is a brain tumor segmentation dataset consists of synthetic and real images, where each of them is further divided into high-grade gliomas (HG) and low-grade gliomas (LG).
36 papers · 2 benchmarks
The color FERET database is a dataset for face recognition.
36 papers · 3 benchmarks
DRealSR (Diverse Real-world image Super-Resolution)
DRealSR establishes a Super Resolution (SR) benchmark with diverse real-world degradation processes, mitigating the limitations of conventional simulated image degradation.
36 papers · 1 benchmark
Fashion-Gen consists of 293,008 high definition (1360 x 1360 pixels) fashion images paired with item descriptions provided by professional stylists.
36 papers · 0 benchmarks
A hand-object interaction dataset with 3D pose annotations of hand and object.
36 papers · 2 benchmarks
In this project, we introduce InfoSeek, a visual question answering dataset tailored for information-seeking questions that cannot be answered with only common sense knowledge.
36 papers · 2 benchmarks
LSUI (Large Scale Underwater Image Dataset)
We released a large-scale underwater image (LSUI) dataset including 5004 image pairs, which involve richer underwater scenes (lighting conditions, water types and target categories) and better visual quality reference images than the…
36 papers · 1 benchmark
MAFL (Multi-Attribute Facial Landmark)
The MAFL dataset contains manually annotated facial landmark locations for 19,000 training and 1,000 test images.
36 papers · 1 benchmark
RAP (Richly Annotated Pedestrian)
The Richly Annotated Pedestrian (RAP) dataset is a dataset for pedestrian attribute recognition.
36 papers · 1 benchmark
Our project (STPLS3D) aims to provide a large-scale aerial photogrammetry dataset with synthetic and real annotated 3D point clouds for semantic and instance segmentation tasks.
36 papers · 3 benchmarks
SciTSR is a large-scale table structure recognition dataset, which contains 15,000 tables in PDF format and their corresponding structure labels obtained from LaTeX source files.
36 papers · 0 benchmarks
TEACh (Task-driven Embodied Agents that Chat)
Robots operating in human spaces must be able to engage in natural language interaction with people, both understanding and executing instructions, and using conversation to resolve ambiguity and recover from mistakes.
36 papers · 0 benchmarks
A large-scale V2X perception dataset using CARLA and OpenCDA
36 papers · 1 benchmark
VisualMRC (VisualMRC: Machine Reading Comprehension on Document Images)
VisualMRC is a visual machine reading comprehension dataset that proposes a task: given a question and a document image, a model produces an abstractive answer.
36 papers · 1 benchmark
Attribution, Relation, and Order (ARO) benchmark to systematically evaluate the ability of VLMs to understand different types of relationships, attributes, and order information.
35 papers · 0 benchmarks
description withheld: archive row vandalised before snapshot
35 papers · 6 benchmarks
The BirdSong dataset consists of audio recordings of bird songs at the H.
35 papers · 0 benchmarks
CIRCO (Composed Image Retrieval on Common Objects in context)
CIRCO (Composed Image Retrieval on Common Objects in context) is an open-domain benchmarking dataset for Composed Image Retrieval (CIR) based on real-world images from COCO 2017 unlabeled set.
35 papers · 1 benchmark
Contains hundreds of frontal view X-rays and is the largest public resource for COVID-19 image and prognostic data, making it a necessary resource to develop and evaluate tools to aid in the treatment of COVID-19.
35 papers · 1 benchmark
This dataset contains complex tables from the annual reports of S&P 500 companies with detailed table structure annotations to help table structure recognition and table data extraction.
35 papers · 0 benchmarks
This is the second version of the Google Landmarks dataset (GLDv2), which contains images annotated with labels representing human-made and natural landmarks.
35 papers · 4 benchmarks
MuCo-3DHP is a large scale training data set showing real images of sophisticated multi-person interactions and occlusions.
35 papers · 0 benchmarks
PST900 is a dataset of 894 synchronized and calibrated RGB and Thermal image pairs with per pixel human annotations across four distinct classes from the DARPA Subterranean Challenge.
35 papers · 1 benchmark
SVT (Street View Text Dataset)
The Street View Text (SVT) dataset was harvested from Google Street View.
35 papers · 1 benchmark
Spot-the-diff is a dataset consisting of 13,192 image pairs along with corresponding human provided text annotations stating the differences between the two images.
35 papers · 0 benchmarks
A dataset for robot grasp planning based on physics simulation.
34 papers · 0 benchmarks
CLEAR is a continual image classification benchmark dataset with a natural temporal evolution of visual concepts in the real world that spans a decade (2004-2014).
34 papers · 0 benchmarks
For the details of the work, the readers are refer to the paper "Feature Pyramid and Hierarchical Boosting Network for Pavement Crack Detection" (FPHB), T-ITS 2019.
34 papers · 0 benchmarks
This dataset focus on two blur types: camera motion blur and defocus blur.
34 papers · 0 benchmarks
ECSSD (Extended Complex Scene Saliency Dataset)
The Extended Complex Scene Saliency Dataset (ECSSD) is comprised of complex scenes, presenting textures and structures common to real-world images.
34 papers · 5 benchmarks
The EgoHands dataset contains 48 Google Glass videos of complex, first-person interactions between two people.
34 papers · 0 benchmarks
IBims-1 (Independent benchmark images and matched scans v1)
iBims-1 (independent Benchmark images and matched scans - version 1) is a new high-quality RGB-D dataset, especially designed for testing single-image depth estimation (SIDE) methods.
34 papers · 2 benchmarks
PolyU Dataset is a large dataset of real-world noisy images with reasonably obtained corresponding “ground truth” images.
34 papers · 1 benchmark
SUIM (Segmentation of Underwater IMagery)
The Segmentation of Underwater IMagery (SUIM) dataset contains over 1500 images with pixel annotations for eight object categories: fish (vertebrates), reefs (invertebrates), aquatic plants, wrecks/ruins, human divers, robots, and…
34 papers · 2 benchmarks
UMDFaces is a face dataset divided into two parts: Still Images - 367,888 face annotations for 8,277 subjects.
34 papers · 0 benchmarks
3dshapes is a dataset of 3D shapes procedurally generated from 6 ground truth independent latent factors.
33 papers · 0 benchmarks
CIHP (Crowd Instance-level Human Parsing)
The Crowd Instance-level Human Parsing (CIHP) dataset has 38,280 diverse human images.
33 papers · 1 benchmark
ClevrTex is a new benchmark designed as the next challenge to compare, evaluate and analyze algorithms for unsupervised multi-object segmentation.
33 papers · 1 benchmark
Cosal2015 is a large-scale dataset for co-saliency detection which consists of 2,015 images of 50 categories, and each group suffers from various challenging factors such as complex environments, occlusion issues, target appearance…
33 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.