Home › Datasets › modality › Images

Images datasets

archive 2025-07-28

3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 32 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Images datasets 1489–1536 of 3,239

F-SIOL-310 (Few-Shot Incremental Object Learning)
F-SIOL-310 is a robotic dataset and benchmark for Few-Shot Incremental Object Learning, which is used to test incremental learning capabilities for robotic vision from a few examples.
4 papers · 0 benchmarks
FACTIFY (a dataset on multi-modal fact verification)
FACTIFY is a dataset on multi-modal fact verification.
4 papers · 0 benchmarks
FES (Fisheye Evaluation Suite)
FES is an indoor dataset that can be used for evaluation of deep learning approaches.
4 papers · 0 benchmarks
The FIGR-8 database is a dataset containing 17,375 classes of 1,548,256 images representing pictograms, ideograms, icons, emoticons or object or conception depictions.
4 papers · 0 benchmarks
FIRE (Fundus Image Registration Dataset)
Fundus Image Registration Dataset (FIRE) is a dataset consisting of 129 retinal images forming 134 image pairs.
4 papers · 1 benchmark
FewSOL (A Dataset for Few-Shot Object Learning in Robotic Environments)
The Few-Shot Object Learning (FewSOL) dataset can be used for object recognition with a few images per object.
4 papers · 0 benchmarks
The FieldSAFE dataset is a multi-modal dataset for obstacle detection in agriculture.
4 papers · 0 benchmarks
Pretrain: 200k Instruction: 100k
4 papers · 0 benchmarks
The Fraunhofer IPA Bin-Picking dataset is a large-scale dataset comprising both simulated and real-world scenes for various objects (potentially having symmetries) and is fully annotated with 6D poses.
4 papers · 0 benchmarks
Fruits 360 (A dataset of images containing fruits, vegetables, nuts and seeds)
Fruits-360 dataset: A dataset of images containing fruits, vegetables, nuts and seeds Version: 2025.03.24.0 Content The following fruits, vegetables and nuts and are included: Apples (different varieties: Crimson Snow, Golden, Golden-Red,…
4 papers · 0 benchmarks
GVLM (Global Very-High-Resolution Landslide Mapping)
For change detection tasks, current open-source datasets mainly focus on building extraction (e.g., WHU building dataset and LEVIR-CD dataset) (Chen and Shi, 2020; Ji et al., 2018) and urban development monitoring (e.g., SECOND dataset,…
4 papers · 1 benchmark
HKR (Handwritten Kazakh and Russian (HKR) Database for Text Recognition)
The database is written in Cyrillic and shares the same 33 characters.
4 papers · 1 benchmark
Images with paired ground-truth caption hierarchies
4 papers · 0 benchmarks
Propose a dataset which adopts multi-channel visual input.
4 papers · 1 benchmark
ImageNet-1k vs NINCO (No ImageNet Class Objects)
The NINCO (No ImageNet Class Objects) dataset is introduced in the ICML 2023 paper In or Out?
4 papers · 1 benchmark
ImgEdit is a large-scale, high-quality image-editing dataset comprising 1.2 million carefully curated edit pairs, which contain both novel and complex single-turn edits, as well as challenging multi-turn tasks.
4 papers · 0 benchmarks
Imp1k is a new dataset of designs annotated with importance information.
4 papers · 0 benchmarks
Incidents1M is a large-scale multi-label dataset for incident detection which contains 977,088 images, with 43 incident and 49 place categories.
4 papers · 0 benchmarks
K-hairstyle is a novel large-scale Korean hairstyle dataset with 256,679 high-resolution images.
4 papers · 0 benchmarks
K-Lane (KAIST-Lane)
KAIST-Lane (K-Lane) is the world’s first and the largest public urban road and highway lane dataset for Lidar.
4 papers · 1 benchmark
From my knowledge, the dataset used in the project is the largest crack segmentation dataset so far.
4 papers · 2 benchmarks
Kitchen Scenes is a multi-view RGB-D dataset of nine kitchen scenes, each containing several objects in realistic cluttered environments including a subset of objects from the BigBird dataset.
4 papers · 0 benchmarks
Kitsune Network Attack Dataset This is a collection of nine network attack datasets captured from a either an IP-based commercial surveillance system or a network full of IoT devices.
4 papers · 0 benchmarks
Kuzushiji-Kanji is an imbalanced dataset of total 3832 Kanji characters (64x64 grayscale, 140,426 images), ranging from 1,766 examples to only a single example per class.
4 papers · 0 benchmarks
LAM(line-level) (The LAM Dataset: A Novel Benchmark for Line-Level Handwritten Text Recognition)
Handwritten Text Recognition (HTR) is an open problem at the intersection of Computer Vision and Natural Language Processing.
4 papers · 1 benchmark
LIMUC (Labeled Images for Ulcerative Colitis)
The LIMUC dataset is the largest publicly available labeled ulcerative colitis dataset that compromises 11276 images from 564 patients and 1043 colonoscopy procedures.
4 papers · 1 benchmark
LLM-Seg40K dataset contains 14K images in total.
4 papers · 0 benchmarks
LayoutBench-COCO is a diagnostic benchmark that examines layout-guided image generation models on arbitrary, unseen layouts.
4 papers · 1 benchmark
LayoutBench-COCO is a diagnostic benchmark that examines layout-guided image generation models on arbitrary, unseen layouts.
4 papers · 1 benchmark
LayoutBench-COCO is a diagnostic benchmark that examines layout-guided image generation models on arbitrary, unseen layouts.
4 papers · 1 benchmark
LayoutBench-COCO is a diagnostic benchmark that examines layout-guided image generation models on arbitrary, unseen layouts.
4 papers · 1 benchmark
This is a 4D light-field dataset of materials.
4 papers · 0 benchmarks
M³-VOS (M³-VOS: Multi-Phase, Multi-Transition, and Multi-Scenery Video Object Segmentation)
💡 Description A new benchmark, Multi-Phase, Multi-Transition, and Multi-Scenery Video Object Segmentation (M³-VOS), to verify the ability of models to understand object phases, which consists of 479 high-resolution videos spanning over 10…
4 papers · 1 benchmark
The M5Product dataset is a large-scale multi-modal pre-training dataset with coarse and fine-grained annotations for E-products.
4 papers · 0 benchmarks
Description The consists of 92 groups of video clips with 113, 918 high resolution frames taken by two drones and 63 groups of video clips with 145, 875 high resolution frames taken by three drones.
4 papers · 0 benchmarks
MERL-RAV (MERL Reannotation of AFLW with Visibility)
The MERL-RAV (MERL Reannotation of AFLW with Visibility) Dataset contains over 19,000 face images in a full range of head poses.
4 papers · 2 benchmarks
The METU Trademark Dataset is a large dataset (the largest publicly available logo dataset as of 2014, and the largest one not requiring any preprocessing as of 2017), which is composed of more than 900K real logos belonging to real…
4 papers · 0 benchmarks
MIMIC-CXR-LT (long-tailed version of MIMIC-CXR)
MIMIC-CXR-LT.
4 papers · 1 benchmark
MJU-Waste is an RGBD waste object segmentation dataset that is made public to facilitate future research in this area.
4 papers · 1 benchmark
MLFP (Multispectral Latex Mask based Video Face Presentation Attack)
The MLFP dataset consists of face presentation attacks captured with seven 3D latex masks and three 2D print attacks.
4 papers · 1 benchmark
MMToM-QA (Multimodal Theory of Mind Question Answering)
MMToM-QA is the first multimodal benchmark to evaluate machine Theory of Mind (ToM), the ability to understand people's minds.
4 papers · 0 benchmarks
The MNIST Large Scale dataset is based on the classic MNIST dataset, but contains large scale variations up to a factor of 16.
4 papers · 1 benchmark
MPV (Multi-Pose Virtual try on)
Consists of 37,723/14,360 person/clothes images, with the resolution of 256x192.
4 papers · 1 benchmark
MUAD (Multiple Uncertainties for Autonomous Driving)
The MUAD dataset (Multiple Uncertainties for Autonomous Driving), consisting of 10,413 realistic synthetic images with diverse adverse weather conditions (night, fog, rain, snow), out-of-distribution objects, and annotations for semantic…
4 papers · 0 benchmarks
MUSES: MUlti-SEnsor Semantic perception dataset (The Multi-Sensor Semantic Perception Dataset for Driving under Uncertainty)
MUSES offers 2500 multi-modal scenes, evenly distributed across various combinations of weather conditions (clear, fog, rain, and snow) and types of illumination (daytime, nighttime).
4 papers · 5 benchmarks
The Middlebury 2001 is a stereo dataset of indoor scenes with multiple handcrafted layouts.
4 papers · 0 benchmarks
MovingFashion is a dataset for video-to-shop, the task of retrieving clothes which are worn in social media videos.
4 papers · 1 benchmark
MuMu is a new dataset of more than 31k albums classified into 250 genre classes.
4 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.