Home › Datasets › modality › Images
Images datasets
archive 2025-07-28
3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 35 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Images datasets 1633–1680 of 3,239
CV-Cities comprises $223,736$ ground panoramic images and an equal number of satellite images all accompanied by high-precision GPS coordinates.
3 papers · 1 benchmark
The dataset has been generated using Town 1 and Town 2 of CARLA Simulator.
3 papers · 0 benchmarks
CelebA-Spoof is a large-scale face anti-spoofing dataset recently introduced in [53].
3 papers · 1 benchmark
CheXlocalize is a radiologist-annotated segmentation dataset on chest X-rays.
3 papers · 0 benchmarks
CoVA (CoVA dataset for Webpage Object Detection / Information Extraction)
We labeled 7,740 webpage screenshots spanning 408 domains (Amazon, Walmart, Target, etc.).
3 papers · 0 benchmarks
Country211 is a dataset released by OpenAI, designed to assess the geolocation capability of visual representations.
3 papers · 2 benchmarks
The 2021 SIGIR workshop on eCommerce is hosting the Coveo Data Challenge for "In-session prediction for purchase intent and recommendations".
3 papers · 1 benchmark
DAD-3DHeads dataset consists of 44,898 images collected from various sources (37,840 in the training set, 4,312 in the validation set, and 2,746 in the test set).
3 papers · 0 benchmarks
DEIC is a benchmark for measuring the data efficiency of models in the context of image classification.
3 papers · 3 benchmarks
DELAUNAY is a dataset of abstract paintings and non-figurative art objects labelled by the artists' names.
3 papers · 0 benchmarks
DOORS (Dataset fOr bOuldeRs Segmentation)
DOORS is a dataset designed for boulders recognition, centroid regression, segmentation, and navigation applications.
3 papers · 0 benchmarks
DeepSportradar is a benchmark suite of computer vision tasks, datasets and benchmarks for automated sport understanding.
3 papers · 0 benchmarks
Digital Peter is a dataset of Peter the Great's manuscripts annotated for segmentation and text recognition.
3 papers · 1 benchmark
DirtyMNIST is a concatenation of MNIST + AmbiguousMNIST, with 60k samples each in the training set.
3 papers · 0 benchmarks
E-NER is a publicly available legal Named Entity Recognition (NER) data set.
3 papers · 0 benchmarks
The automated recognition of different vehicle classes and their orientation on aerial images is an important task in the field of traffic research and also finds applications in disaster management, among other things.
3 papers · 0 benchmarks
EC-FUNSD is introduced in [[arXiv:2402.02379]](https://arxiv.org/abs/2402.02379) as a benchmark of semantic entity recognition (SER) and entity linking (EL), designed for the entity-centric robustness evaluation of pre-trained…
3 papers · 2 benchmarks
The ECUST Food Dataset is a food recognition dataset that contains 2978 images Source: https://github.com/Liang-yc/ECUSTFD-resized- Image Source: https://github.com/Liang-yc/ECUSTFD-resized-
3 papers · 0 benchmarks
From Grounded Human-Object Interaction Hotspots from Video (ICCV'19): We collect annotations for interaction keypoints on EPIC Kitchens in order to quantitatively evaluate our method in parallel to the OPRA dataset (where annotations are…
3 papers · 1 benchmark
The BIWI Walking Pedestrians dataset consists of walking pedestrians in busy scenarios from a birds eye view.
3 papers · 0 benchmarks
EasyPortrait (Face Parsing and Portrait Segmentation Dataset)
We introduce a large-scale image dataset EasyPortrait for portrait segmentation and face parsing.
3 papers · 0 benchmarks
ElBa (ElBa: Element Based Textures Dataset)
ElBa is composed of procedurally-generated realistic renderings, where we vary in a continuous way element shapes and colors and their distribution, to generate 30K texture images with different local symmetry, stationarity, and density of…
3 papers · 0 benchmarks
FEAFA+ is a dataset for Facial expression analysis and 3D Facial animation.
3 papers · 0 benchmarks
FFHQ-Text is a small-scale face image dataset with large-scale facial attributes, designed for text-to-face generation & manipulation, text-guided facial image manipulation, and other vision-related tasks.
3 papers · 0 benchmarks
FOD in Airports (FOD-A) is an image dataset of FOD, Foreign Object Degris, which consists of 31 object categories and over 30,000 annotation instances.
3 papers · 0 benchmarks
We introduce FUNSD-r and CORD-r in Token Path Prediction, the revised VrD-NER datasets to reflect the real-world scenarios of NER on scanned VrDs.
3 papers · 1 benchmark
Fetoscopic Placental Vessel Segmentation and Registration (FetReg) is a large-scale multi-centre dataset for the development of generalized and robust semantic segmentation and video mosaicking algorithms for the fetal environment with a…
3 papers · 0 benchmarks
FindingEmo is an image dataset containing annotations for 25k images, specifically tailored to Emotion Recognition.
3 papers · 0 benchmarks
Fishnet Open Images Database is a large dataset of EM imagery for fish detection and fine-grained categorisation onboard commercial fishing vessels.
3 papers · 0 benchmarks
Fitness-AQA (Fitness Action Quality Assessment [ECCV 2022])
Largest, first-of-its-kind, in-the-wild, fine-grained workout/exercise posture analysis dataset, covering three different exercises: BackSquat, Barbell Row, and Overhead Press.
3 papers · 0 benchmarks
The Five-Billion-Pixels dataset contains more than 5 billion labeled pixels of 150 high-resolution Gaofen-2 (4 m) satellite images, annotated in a 24-category system covering artificial-constructed, agricultural, and natural classes.
3 papers · 0 benchmarks
FixMyPose is a dataset for automated pose correction.
3 papers · 0 benchmarks
GFP-GOWT1 mouse stem cells Dr.
3 papers · 2 benchmarks
FoodLogoDet-1500 is a new large-scale publicly available food logo dataset, which has 1,500 categories, about 100,000 images and about 150,000 manually annotated food logo objects.
3 papers · 0 benchmarks
The Forms Dataset is a dataset for document structure extraction comprising of 5K forms.
3 papers · 0 benchmarks
We present a new large-scale photorealistic panoramic dataset named FutureHouse, which has the following characteristics.
3 papers · 0 benchmarks
GAS (Grasp Area Segmentation)
GAS (Grasp Area Segmentation) dataset consists of 10089 RGB images of cluttered scenes grouped into 1121 grasp-area segmentation tasks.
3 papers · 0 benchmarks
Most publications that aim to optimize neural networks for CBIR, train and test their models on domain specific datasets.
3 papers · 0 benchmarks
Goldfinch is a dataset for fine-grained recognition challenges.
3 papers · 0 benchmarks
HASY is a dataset of single symbols similar to MNIST.
3 papers · 0 benchmarks
An autnonomous driving dataset and benchmark for optical flow.
3 papers · 0 benchmarks
HDM05 is a MoCap (motion capture) dataset.
3 papers · 1 benchmark
HOPE-Image (Household Objects for Pose Estimation)
The NVIDIA HOPE datasets consist of RGBD images and video sequences with labeled 6-DoF poses for 28 toy grocery objects.
3 papers · 0 benchmarks
HSD (Honda Scenes Dataset)
An annotated dataset is released to enable dynamic scene classification that includes 80 hours of diverse high quality driving video data clips collected in the San Francisco Bay area.
3 papers · 0 benchmarks
HSPACE (Human-SPACE) is a large-scale photo-realistic dataset of animated humans placed in complex synthetic indoor and outdoor environments.
3 papers · 1 benchmark
Hazards&Robots (Hazards&Robots: A Dataset for Visual Anomaly Detection in Robotics)
We consider the problem of detecting, in the visual sensing data stream of an autonomous mobile robot, semantic patterns that are unusual (i.e., anomalous) with respect to the robot’s previous experience in similar environments.
3 papers · 0 benchmarks
Hephaestus (Hephaestus: A large scale multitask dataset towards InSAR understanding)
Hephaestus is the first large-scale InSAR dataset.
3 papers · 0 benchmarks
We present datasets containing urban traffic and rural road scenes recorded using hyperspectral snap-shot sensors mounted on a moving car.
3 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.