Home › Datasets › modality › Images

Images datasets

archive 2025-07-28

3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 15 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Images datasets 673–720 of 3,239

Description: 105,941 Images Natural Scenes OCR Data of 12 Languages.
24 papers · 0 benchmarks
AFLW-19 (The 19 landmark variant of AFLW.)
The original AFLW provides at most 21 points for each face, but excluding coordinates for invisible landmarks, causing difficulties for training most of the existing baseline approaches.
24 papers · 1 benchmark
CCPD (Chinese City Parking Dataset)
The Chinese City Parking Dataset (CCPD) is a dataset for license plate detection and recognition.
24 papers · 0 benchmarks
DADA-seg is a pixel-wise annotated accident dataset, which contains a variety of critical scenarios from traffic accidents.
24 papers · 1 benchmark
DHF1K is a video saliency dataset which contains a ground-truth map of binary pixel-wise gaze fixation points and a continuous map of the fixation points after being blurred by a gaussian filter.
24 papers · 1 benchmark
FlickrStyle10K is collected and built on Flickr30K image caption dataset.
24 papers · 2 benchmarks
The Google Landmarks dataset contains 1,060,709 images from 12,894 landmarks, and 111,036 additional query images.
24 papers · 0 benchmarks
HarMeme is a benchmark dataset for hateful meme classification containing 3, 544 memes related to COVID-19 collected from the Internet
24 papers · 1 benchmark
The INRIA Person dataset is a dataset of images of persons used for pedestrian detection.
24 papers · 0 benchmarks
InterHuman is a multimodal dataset, named InterHuman.
24 papers · 1 benchmark
The Segmenting and Tracking Every Pixel (STEP) benchmark consists of 21 training sequences and 29 test sequences.
24 papers · 2 benchmarks
MMSE-HR (Multimodal Spontaneous Expression-Heart Rate dataset)
The MMSE-HR benchmark consists of a dataset of 102 videos from 40 subjects recorded at 1040x1392 raw resolution at 25fps.
24 papers · 1 benchmark
RADIATE (RAdar Dataset In Adverse weaThEr)
RADIATE (RAdar Dataset In Adverse weaThEr) is new automotive dataset created by Heriot-Watt University which includes Radar, Lidar, Stereo Camera and GPS/IMU.
24 papers · 2 benchmarks
REALY (Region-aware benchmark based on the LYHM)
The REALY benchmark aims to introduce a region-aware evaluation pipeline to measure the fine-grained normalized mean square error (NMSE) of 3D face reconstruction methods from under-controlled image sets.
24 papers · 2 benchmarks
RecipeQA is a dataset for multimodal comprehension of cooking recipes.
24 papers · 1 benchmark
ARCTIC (Articulated Objects in Free-form Hand Interaction)
ARCTIC is a dataset of free-form interactions of hands and articulated objects.
23 papers · 0 benchmarks
CIFAR10-DVS is an event-stream dataset for object classification.
23 papers · 2 benchmarks
The DeepWeeds dataset consists of 17,509 images capturing eight different weed species native to Australia in situ with neighbouring flora.
23 papers · 0 benchmarks
DocUNet (Document Image Unwarping via a Stacked U-Net)
Various documents dataset.
23 papers · 3 benchmarks
HPS (Human POSEitioning System Dataset)
HPS Dataset is a collection of 3D humans interacting with large 3D scenes (300-1000 m², up to 2500 m²).
23 papers · 0 benchmarks
HRSOD (High-Resolution Salient Object Detection)
There exist several datasets for saliency detection, but none of them is specifically designed for high-resolution salient object detection.
23 papers · 1 benchmark
ITOP (Invariant-Top View Dataset)
The ITOP dataset consists of 40K training and 10K testing depth images for each of the front-view and top-view tracks.
23 papers · 3 benchmarks
IXI (IXI Brain Development Dataset)
IXI Dataset is a collection of 600 MR brain images from normal, healthy subjects.
23 papers · 4 benchmarks
ImageNet-W (ImageNet-Watermark)
ImageNet-W(atermark) is a test set to evaluate models’ reliance on the newly found watermark shortcut in ImageNet, which is used to predict the carton class.
23 papers · 0 benchmarks
OpenImages V6 is a large-scale dataset , consists of 9 million training images, 41,620 validation samples, and 125,456 test samples.
23 papers · 2 benchmarks
The PASCAL FACE dataset is a dataset for face detection and face recognition.
23 papers · 1 benchmark
PIE-Bench (Prompt-based Image Editing Benchmark)
PIE-Bench comprises 700 images featuring 10 distinct editing types.
23 papers · 1 benchmark
The PhysioNet Challenge 2012 dataset is publicly available and contains the de-identified records of 8000 patients in Intensive Care Units (ICU).
23 papers · 5 benchmarks
ShapeWorld is a new evaluation methodology and framework for multimodal deep learning models, with a focus on formal-semantic style generalization capabilities.
23 papers · 0 benchmarks
This is a 21 class land use image dataset meant for research purposes.
23 papers · 1 benchmark
Car CAD models from "3d object detection and viewpoint estimation with a deformable 3d cuboid model" were used to generate the dataset.
22 papers · 0 benchmarks
BAR (Biased Action Recognition)
Biased Action Recognition (BAR) dataset is a real-world image dataset categorized as six action classes which are biased to distinct places.
22 papers · 1 benchmark
COWC (Cars Overhead With Context)
The Cars Overhead With Context (COWC) data set is a large set of annotated cars from overhead.
22 papers · 0 benchmarks
CUFSF (CUHK Face Sketch FERET Database)
The CUHK Face Sketch FERET (CUFSF) is a dataset for research on face sketch synthesis and face sketch recognition.
22 papers · 1 benchmark
CxC (Crisscrossed Captions)
Crisscrossed Captions (CxC) contains 247,315 human-labeled annotations including positive and negative associations between image pairs, caption pairs and image-caption pairs.
22 papers · 1 benchmark
The Easy Communications (EasyCom) dataset is a world-first dataset designed to help mitigate the cocktail party effect from an augmented-reality (AR) -motivated multi-sensor egocentric world view.
22 papers · 4 benchmarks
IMDb-Face is large-scale noise-controlled dataset for face recognition research.
22 papers · 0 benchmarks
(JHU-CROWD) a crowd counting dataset that contains 4,250 images with 1.11 million annotations.
22 papers · 0 benchmarks
100 tasks from LIBERO-100 suite.
22 papers · 1 benchmark
The LOCATA dataset is a dataset for acoustic source localization.
22 papers · 0 benchmarks
Market-1501-C is an evaluation set that consists of algorithmically generated corruptions applied to the Market-1501 test-set.
22 papers · 1 benchmark
The PhotoShape dataset consists of photorealistic, relightable, 3D shapes produced by the work proposed in the work of Park et al.
22 papers · 1 benchmark
The Sku110k dataset provides 11,762 images with more than 1.7 million annotated bounding boxes captured in densely packed scenarios, including 8,233 images for training, 588 images for validation, and 2,941 images for testing.
22 papers · 1 benchmark
Satlas is a remote sensing dataset and benchmark that is large in both breadth, featuring all of the aforementioned applications and more, as well as scale, comprising 290M labels under 137 categories and 7 label modalities.
22 papers · 0 benchmarks
Our goal is to improve upon the status quo for designing image classification models trained in one domain that perform well on images from another domain.
22 papers · 3 benchmarks
WIDER (Web Image Dataset for Event Recognition)
WIDER is a dataset for complex event recognition from static images.
22 papers · 1 benchmark
WebSRC (WebSRC: A Dataset for Web-Based Structural Reading Comprehension)
WebSRC is a novel Web-based Structural Reading Comprehension dataset.
22 papers · 2 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.