Home › Datasets › modality › Images

Images datasets

archive 2025-07-28

3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 23 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Images datasets 1057–1104 of 3,239

MOD (Meme incorporated Open-domain Dialogue)
MOD is a large-scale open-domain multimodal dialogue dataset incorporating abundant Internet memes into utterances.
9 papers · 0 benchmarks
MVK (Marine Video Kit)
The dataset contains single-shot videos taken from moving cameras in underwater environments.
9 papers · 1 benchmark
Middlebury 2005 is a stereo dataset of indoor scenes.
9 papers · 0 benchmarks
The Pascal Panoptic Parts dataset consists of annotations for the part-aware panoptic segmentation task on the PASCAL VOC 2010 dataset.
9 papers · 2 benchmarks
PerSeg is a dataset for personalized segmentation.
9 papers · 1 benchmark
PhenoBench (PhenoBench — A Large Dataset and Benchmarks for Semantic Image Interpretation in the Agricultural Domain)
The PhenoBench dataset contains multiple image segmentation challenges from the agricultural domain.
9 papers · 0 benchmarks
Most existing MOT datasets are captured using pinhole cameras, which are characterized by a narrow-FoV and linear sensor motion.
9 papers · 1 benchmark
Consists of 330,000 sketches and 204,000 photos spanning across 110 categories.
9 papers · 0 benchmarks
ReDWeb-S is a large-scale challenging dataset for Salient Object Detection.
9 papers · 0 benchmarks
SAMRS is a remote sensing segmentation dataset which provides object category, location, and instance information that can be used for semantic segmentation, instance segmentation, and object detection, either individually or in…
9 papers · 0 benchmarks
SOBA (Shadow-OBject Association)
A new dataset called SOBA, named after Shadow-OBject Association, with 3,623 pairs of shadow and object instances in 1,000 photos, each with individual labeled masks.
9 papers · 1 benchmark
SUM is a new benchmark dataset of semantic urban meshes which covers about 4 km2 in Helsinki (Finland), with six classes: Ground, Vegetation, Building, Water, Vehicle, and Boat.
9 papers · 0 benchmarks
SketchyScene is a large-scale dataset of scene sketches to advance research on sketch understanding at both the object and scene level.
9 papers · 0 benchmarks
TUM-GAID (TUM Gait from Audio, Image and Depth) collects 305 subjects performing two walking trajectories in an indoor environment.
9 papers · 0 benchmarks
Social media are interactive platforms that facilitate the creation or sharing of information, ideas or other forms of expression among people.
9 papers · 1 benchmark
UDIVA is a new non-acted dataset of face-to-face dyadic interactions, where interlocutors perform competitive and collaborative tasks with different behavior elicitation and cognitive workload.
9 papers · 0 benchmarks
We introduce a novel Image Quality Assessment (IQA) dataset comprising 6073 UHD-1 (4K) images, annotated at a fixed width of 3840 pixels.
9 papers · 1 benchmark
UIIS (General Underwater Image Instance Segmentation dataset)
This is the first general Underwater Image Instance Segmentation (UIIS) dataset containing 4,628 images for 7 categories with pixel-level annotations for underwater instance segmentation task
9 papers · 1 benchmark
Have need seven multiple exposure ground truth images satisfying EV 0, ±1, ±2, ±3 for static scenes.
9 papers · 1 benchmark
VISUELLE is a repository build upon the data of a real fast fashion company, Nunalie, and is composed of 5577 new products and about 45M sales related to fashion seasons from 2016-2019.
9 papers · 1 benchmark
VLM2-Bench (VLM²-Bench)
VLM²-Bench: Benchmarking Vision-Language Models on Visual Cue Matching Description VLM²-Bench is the first comprehensive benchmark designed to evaluate vision-language models' (VLMs) ability to visually link matching cues across…
9 papers · 1 benchmark
VOT2020 is a Visual Object Tracking benchmark for short-term tracking in RGB.
9 papers · 1 benchmark
The Zenseact Open Dataset (ZOD) is a large-scale and diverse multi-modal autonomous driving (AD) dataset, created by researchers at Zenseact.
9 papers · 0 benchmarks
The iCartoonFace dataset is a large-scale dataset that can be used for two different tasks: cartoon face detection and cartoon face recognition.
9 papers · 1 benchmark
The iWildCam2020-WILDS dataset is a variant of the iWildCam 2020 dataset.
9 papers · 1 benchmark
3DFAW contains 23k images with 66 3D face keypoint annotations.
8 papers · 2 benchmarks
AOLP (Application-oriented License Plate)
The application-oriented license plate (AOLP) benchmark database has 2049 images of Taiwan license plates.
8 papers · 2 benchmarks
APRICOT is a collection of over 1,000 annotated photographs of printed adversarial patches in public locations.
8 papers · 0 benchmarks
ATRW (Amur Tiger Re-identification in the Wild)
The ATRW Dataset contains over 8,000 video clips from 92 Amur tigers, with bounding box, pose keypoint, and tiger identity annotations.
8 papers · 0 benchmarks
The Airport dataset is a dataset for person re-identification which consists of 39,902 images and 9,651 identities across six cameras.
8 papers · 0 benchmarks
Amazon Toys & Games (Amazon Toys & Games 5-core)
This dataset includes reviews (ratings, text, helpfulness votes), product metadata (descriptions, category information, price, brand, and image features), and links (also viewed/also bought graphs).
8 papers · 1 benchmark
CAMO++ is a dataset for camouflaged object segmentation.
8 papers · 0 benchmarks
CLEVR-Math is a multi-modal math word problems dataset consisting of simple math word problems involving addition/subtraction, represented partly by a textual description and partly by an image illustrating the scenario.
8 papers · 0 benchmarks
COCO-MIG (COCO-MIG benchmark)
The COCO-MIG benchmark (Common Objects in Context Multi-Instance Generation) is a benchmark used to evaluate the generation capability of generators on text containing multiple attributes of multi-instance objects.
8 papers · 1 benchmark
COUCH is a large human-chair interaction dataset with clean annotations.
8 papers · 0 benchmarks
Under a close collaboration with an expert radiologist team of the Hospital Universitario San Cecilio, the COVIDGR-1.0 dataset of patients' anonymized X-ray images has been built.
8 papers · 2 benchmarks
Collects shadow images for multiple scenarios and compiled a new dataset of 10,500 shadow images, each with labeled ground-truth mask, for supporting shadow detection in the complex world.
8 papers · 1 benchmark
The Cityscapes Panoptic Parts dataset introduces part-aware panoptic segmentation annotations for the Cityscapes dataset.
8 papers · 1 benchmark
Cops-Ref is a dataset for visual reasoning in context of referring expression comprehension with two main features.
8 papers · 0 benchmarks
DISC21 (Dataset for ISC 2021)
DISC21 is a benchmark for large-scale image similarity detection.
8 papers · 1 benchmark
DOTmark (Discrete Optimal Transport Benchmark)
DOTmark is a benchmark for discrete optimal transport, which is designed to serve as a neutral collection of problems, where discrete optimal transport methods can be tested, compared to one another, and brought to their limits on…
8 papers · 0 benchmarks
Are current 3D object tracking methods truely robust enough for low-fidelity depth sensors like the iPhone LiDAR?
8 papers · 2 benchmarks
DUO (Detecting Underwater Objects)
DUO is a dataset for Underwater object detection for robot picking.
8 papers · 1 benchmark
DeepLoc is a large-scale urban outdoor localization dataset.
8 papers · 0 benchmarks
Duke Breast Cancer MRI (Dynamic contrast-enhanced magnetic resonance images of breast cancer patients with tumor locations)
Breast MRI scans of 922 cancer patients from Duke University, with tumor bounding box annotations, clinical, imaging, and many other features, and more.
8 papers · 0 benchmarks
Echocardiography, or cardiac ultrasound, is the most widely used and readily available imaging modality to assess cardiac function and structure.
8 papers · 0 benchmarks
FERG (Facial Expression Research Group Database)
FERG is a database of cartoon characters with annotated facial expressions containing 55,769 annotated face images of six characters.
8 papers · 1 benchmark
FeTS2022 (Federated Tumor Segmentation Challenge 2022)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
8 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.