Home › Datasets › modality › Images
Images datasets
archive 2025-07-28
3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 42 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Images datasets 1969–2016 of 3,239
The LeukemiaAttri dataset is a large-scale, multi-domain collection of microscopy images derived from leukemia patient samples, enriched with detailed morphological information.
2 papers · 2 benchmarks
This data set contains 775 video sequences, captured in the wildlife park Lindenthal (Cologne, Germany) as part of the AMMOD project, using an Intel RealSense D435 stereo camera.
2 papers · 0 benchmarks
LoTE-Animal (LoTE-Animal: A Long Time-span Dataset for Endangered Animal Behavior Understanding)
Understanding and analyzing animal behavior is increasingly essential to protect endangered animal species.
2 papers · 1 benchmark
Usually, the information related to the crop types available in a given territory is annual information, that is, we only know the type of main crop grown over a year and we do not know any crops that have followed one another during the…
2 papers · 1 benchmark
LymphoMNIST is a comprehensive dataset designed for the nuanced classification of lymphocyte images.
2 papers · 0 benchmarks
MAMe (Museum Art Medium dataset)
The MAMe dataset contains images of high-resolution and variable shape of artworks from 3 different museums: - The Metropolitan Museum of Art of New York - The Los Angeles County Museum of Art - The Cleveland Museum of Art Source:…
2 papers · 1 benchmark
MAPS-KB is a million-scale probabilistic simile knowledge base, covering 4.3 million triplets over 0.4 million terms from 70 GB corpora.
2 papers · 0 benchmarks
Dataset page: https://github.com/mosamdabhi/MBW-Data MBW - Zoo is a challenging dataset consisting image frames of tail-end distribution categories (such as Fish, Colobus Monkeys, Chimpanzees, etc.) with their corresponding 2D, 3D, and…
2 papers · 0 benchmarks
MHSMA (The Modified Human Sperm Morphology Analysis)
The MHSMA dataset is a collection of human sperm images from 235 patients with male factor infertility.
2 papers · 0 benchmarks
MIAD contains more than 100K high-resolution color images in various outdoor industrial scenarios, designed for unsupervised anomaly detection.
2 papers · 0 benchmarks
This database is provided and maintained by Dr.
2 papers · 1 benchmark
MLP (Multimodal Lecture Presentations)
Multimodal Lecture Presentations (MLP) is a large-scale benchmark dataset for testing the capabilities of machine learning models in multimodal understanding of educational content.
2 papers · 0 benchmarks
MMSD2.0 (Towards a Reliable Multi-modal Sarcasm Detection System)
Multi-modal sarcasm detection has attracted much recent attention.
2 papers · 0 benchmarks
MMVax-Stance includes 113 Vaccine Hesitancy Framings found on Twitter about the COVID-19 vaccines.
2 papers · 0 benchmarks
MSDA (Multi-source domain adaptation dataset for text recognition)
5 domains: synthetic domain, document domain, street view domain, handwritten domain, and car license domain over five million images
2 papers · 2 benchmarks
MSRA10K (MSRA10K Salient Object Database)
MSRA10K is a dataset for salient object detection that contains 10,000 images with pixel-level saliency labeling for 10K images from the MSRA salient object detection dataset.
2 papers · 0 benchmarks
MTST (Mobile Turkish Scene Text)
The Mobile Turkish Scene Text (MTST 200) dataset consists of 200 indoor and outdoor Turkish scene text images.
2 papers · 0 benchmarks
MUStARD (Multimodal Sarcasm Detection Dataset)
We release the MUStARD dataset which is a multimodal video corpus for research in automated sarcasm discovery.
2 papers · 0 benchmarks
MVTec D2S (MVTec Densely Segmented Supermarket)
MVTec D2S is a benchmark for instance-aware semantic segmentation in an industrial domain.
2 papers · 0 benchmarks
MagicBathyNet is a benchmark dataset made up of image patches of Sentinel-2, SPOT-6 and aerial imagery, bathymetry in raster format and seabed classes annotations.
2 papers · 0 benchmarks
This dataset contains 1203 individuals captured from two disjoint camera views.
2 papers · 0 benchmarks
Our primary objective in creating this dataset is to support researchers in the advancement of algorithms for keypoints detection and the pretraining of large models on retinal images using a self-supervised approach.
2 papers · 0 benchmarks
MedMNIST-C is an open-source data set collection comprising algorithmically generated corruptions applied to the test sets of the MedMNIST collection following the concept of ImageNet-C.
2 papers · 0 benchmarks
Middlebury 2003 is a stereo dataset for indoor scenes.
2 papers · 0 benchmarks
Mila Simulated Floods Dataset is a 1.5 square km virtual world using the Unity3D game engine including urban, suburban and rural areas.
2 papers · 1 benchmark
MiniWob++ is a suite of web-browser based tasks introduced in Liu et al.
2 papers · 0 benchmarks
A large scale OCSR dataset, proposed in paper “MolParser: End-to-end Visual Recognition of Molecule Structures in the Wild“ MolParser-7M contains nearly 8 million paired image-SMILES data.
2 papers · 0 benchmarks
This dataset consists of blurred, noisy and defocused images.
2 papers · 0 benchmarks
MuCo-VQA consist of large-scale (3.7M) multilingual and code-mixed VQA datasets in multiple languages: Hindi (hi), Bengali (bn), Spanish (es), German (de), French (fr) and code-mixed language pairs: en-hi, en-bn, en-fr, en-de and en-es.
2 papers · 0 benchmarks
MuViHand is a dataset for 3D Hand Pose Estimation that consists of multi-view videos of the hand along with ground-truth 3D pose labels.
2 papers · 0 benchmarks
Multi Task Crowd is a new 100 image dataset fully annotated for crowd counting, violent behaviour detection and density level classification.
2 papers · 0 benchmarks
We applied our framework, dubbed as ”PreNeRF 360”, to enable the use of the Nutrition5k dataset in NeRF and introduce an updated version of this dataset, known as the N5k360 dataset.
2 papers · 0 benchmarks
The archive contains original images from NIH3T3 cells stained with Hoechst 33342 as PNG files.
2 papers · 0 benchmarks
A high-quality captured dataset for object relighting.
2 papers · 0 benchmarks
A high-quality synthetic dataset for object relighting.
2 papers · 0 benchmarks
This collection contains images from 422 non-small cell lung cancer (NSCLC) patients.
2 papers · 0 benchmarks
NVD (Naturalistic Variation Object Dataset)
Naturalistic Variation Object Dataset (NVD) is a large simulated dataset of 272k images of everyday objects with naturalistic variations such as object pose, scale, viewpoint, lighting and occlusions.
2 papers · 0 benchmarks
This dataset is recreated using offline augmentation from the original dataset.
2 papers · 2 benchmarks
Noise of Web (NoW) is a challenging noisy correspondence learning (NCL) benchmark for robust image-text matching/retrieval models.
2 papers · 0 benchmarks
Scene-focused, multi-modal, episodic data of the images and symbolic world-states seen by an agent completing a pogo-stick assembly task within a video game world.
2 papers · 0 benchmarks
The Online Action Detection Dataset (OAD) was captured using the Kinect V2 sensor, which collects color images, depth images and human skeleton joints synchronously.
2 papers · 1 benchmark
OADAT (OADAT: Experimental and Synthetic Clinical Optoacoustic Data for Standardized Image Processing)
An experimental and synthetic (simulated) OA raw signals and reconstructed image domain datasets rendered with different experimental parameters and tomographic acquisition geometries.
2 papers · 0 benchmarks
OCR-IDL (OCR Annotations for Industry Document Library Dataset)
The OCR-IDL dataset comprises the OCR annotations for a subset of 26M pages of the large-scale IDL document library.
2 papers · 0 benchmarks
ODMS (Object Depth via Motion and Segmentation)
ODMS is a dataset for learning Object Depth via Motion and Segmentation.
2 papers · 0 benchmarks
Occluded COCO is automatically generated subset of COCO val dataset, collecting partially occluded objects for a large variety of categories in real images in a scalable manner, where target object is partially occluded but the…
2 papers · 1 benchmark
OmniCity is a dataset for omnipotent city understanding from multi-level and multi-view images.
2 papers · 0 benchmarks
(L)ifel(O)ng (R)obotic V(IS)ion (OpenLORIS) - Object Recognition Dataset (OpenLORIS-Object) is designed for accelerating the lifelong/continual/incremental learning research and application,currently focusing on improving the continuous…
2 papers · 0 benchmarks
OpenViDial 2.0 is a larger-scale open-domain multi-modal dialogue dataset compared to the previous version OpenViDial 1.0.
2 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.