Home › Datasets › modality › Images
Images datasets
archive 2025-07-28
3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 36 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Images datasets 1681–1728 of 3,239
The ICVL dataset is a hand pose estimation dataset that consists of 330K training frames and 2 testing sequences with each 800 frames.
3 papers · 0 benchmarks
Please refer: https://github.com/google/imageinwords/blob/main/datasets/IIW-400/README.md
3 papers · 0 benchmarks
We have cleaned the noisy IMDB-WIKI dataset using a constrained clustering method, resulting this new benchmark for in-the-wild age estimation.
3 papers · 1 benchmark
IllusionVQA is a Visual Question Answering (VQA) dataset with two sub-tasks.
3 papers · 2 benchmarks
transform the ImageNet-1K classification datatset for Chinese models by translating labels and prompts into Chinese.
3 papers · 1 benchmark
Imgur5k is a large-scale handwritten in-the-wild dataset, containing challenging real world handwritten samples from nearly 5K writers.
3 papers · 0 benchmarks
The Indoor-6 dataset was created from multiple sessions captured in six indoor scenes over multiple days.
3 papers · 0 benchmarks
This Dataset consists of 2120 sequences of binary masks of pedestrians.
3 papers · 1 benchmark
KOHTD (Kazakh Offline Handwritten Text Dataset)
Kazakh offline Handwritten Text dataset (KOHTD) has 3000 handwritten exam papers and more than 140335 segmented images and there are approximately 922010 symbols.
3 papers · 1 benchmark
The Kenyan Food Type Dataset (KenyanFood13) is an image classification dataset for Kenyan food.
3 papers · 0 benchmarks
Konzil dataset was created by specialists of the University of Greifswald.
3 papers · 0 benchmarks
The Kvasir-VQA dataset is an extended dataset derived from the HyperKvasir and Kvasir-Instrument datasets, augmented with question-and-answer annotations.
3 papers · 0 benchmarks
The dataset contains a Video capsule endoscopy dataset for polyp segmentation.
3 papers · 1 benchmark
LAGENDA (Layer Age and Gender Dataset)
The LAGENDA dataset is a large-scale dataset with age and gender annotations for face and body bounding boxes.
3 papers · 4 benchmarks
LAION-COCO is the world’s largest dataset of 600M generated high-quality captions for publicly available web-images.
3 papers · 1 benchmark
A Large Dataset for Remote Sensing Image Change Captioning.
3 papers · 0 benchmarks
The dataset was proposed in LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.
3 papers · 0 benchmarks
LOOK is a large-scale dataset for eye contact detection in the wild, which focuses on diverse and unconstrained scenarios for real-world generalization.
3 papers · 0 benchmarks
LSA16 (Lengua de Señas Argentina - 16 Handshapes classes)
This database contains images of 16 handshapes of the Argentinian Sign Language (LSA), each performed 5 times by 10 different subjects, for a total of 800 images.
3 papers · 1 benchmark
Large Age-Gap (LAG) is a dataset for face verification, The dataset contains 3,828 images of 1,010 celebrities.
3 papers · 0 benchmarks
LayoutBench is a diagnostic benchmark that examines 4 spatial control skills (number, position, size, shape), where each skill consists of 2 OOD layout splits, i.e., in total 8 tasks = 4 skills x 2 splits.
3 papers · 1 benchmark
MCVQA (Multilingual and Code-mixed Visual Question Answering)
The MCVQA dataset consists of 248, 349 training questions and 121, 512 validation questions for real images in Hindi and Code-mixed.
3 papers · 0 benchmarks
MDID (Multimodal Document Intent Dataset)
The Multimodal Document Intent Dataset (MDID) is a dataset for computing author intent from multimodal data from Instagram.
3 papers · 0 benchmarks
Operating rooms (ORs) are complex, high-stakes environments requiring precise understanding of interactions among medical staff, tools, and equipment for enhancing surgical assistance, situational awareness, and patient safety.
3 papers · 3 benchmarks
MMCode is a multi-modal code generation dataset designed to evaluate the problem-solving skills of code language models in visually rich contexts (i.e.
3 papers · 0 benchmarks
The main goal of the data collection is to acquire highly natural conversations that cover a wide variety of styles and scenarios.
3 papers · 2 benchmarks
MULTI-Benchmark is a cutting-edge benchmark for evaluating Multimodal Large Language Models (MLLMs).
3 papers · 0 benchmarks
MUTE (Multimodal Bengali Hateful Memes Dataset)
MUTE This is the first open-source Bengali Hateful Meme dataset, consisting of around 4200 memes annotated with two labels: hate and not hate.
3 papers · 0 benchmarks
The “Medico automatic polyp segmentation challenge” aims to develop computer-aided diagnosis systems for automatic polyp segmentation to detect all types of polyps (for example, irregular polyp, smaller or flat polyps) with high efficiency…
3 papers · 1 benchmark
Mirrored-Human is a dataset for 3D pose estimation from a single view.
3 papers · 0 benchmarks
We sample 2025 frames of images from the original KITTI for Mono3DRefer, containing 41,140 expressions in total and a vocabulary of 5,271 words.
3 papers · 0 benchmarks
The dataset, comprising 1204 meticulously curated images, serves as a comprehensive resource for advancing real-time mosquito detection models.
3 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
3 papers · 0 benchmarks
Contains squared blocks of 48×48 pixels including 13 Sentinel-2 bands.
3 papers · 1 benchmark
NILUT (NILUT 3D LUT Dataset and enhanced images from MIT5K)
Read all the details about the dataset in our paper "NILUT: Conditional Neural Implicit 3D Lookup Tables for Image Enhancement" We host the dataset in Kaggle: https://www.kaggle.com/datasets/photolab/nilut-3d-lut-dataset More information…
3 papers · 0 benchmarks
This dataset provides the VCIP 2020 Grand Challenge on the NIR Image Colorization dataset.
3 papers · 1 benchmark
NKL (short for NanKai Lines) is a dataset for semantic line detection.
3 papers · 1 benchmark
NOD (Night Object Detection)
This is a high-quality large-scale Night Object Detection (NOD) dataset of outdoor images targeting low-light object detection.
3 papers · 0 benchmarks
Replay data from human players and AI agents navigating in a 3D game environment.
3 papers · 0 benchmarks
OAM-TCD is a dataset of around 5k aerial images from around the world to support robust tree detection algorithms.
3 papers · 0 benchmarks
The OCTAGON dataset is a set of Angiography by Octical Coherence Tomography images (OCT-A) used to the segmentation of the Foveal Avascular Zone (FAZ).
3 papers · 0 benchmarks
OFDIW (OnFocus Detection In the Wild)
OnFocus Detection In the Wild (OFDIW) is an onfocus detection dataset.
3 papers · 0 benchmarks
ORVS (Online Retinal image for Vessel Segmentation (ORVS))
The ORVS dataset has been newly established as a collaboration between the computer science and visual-science departments at the University of Calgary.
3 papers · 0 benchmarks
The OpeReid dataset is a person re-identification dataset that consists of 7,413 images of 200 persons.
3 papers · 0 benchmarks
OpenCHAIR is a benchmark for evaluating open-vocabulary hallucinations in image captioning models.
3 papers · 0 benchmarks
OpenLane-V2 is the world's first perception and reasoning benchmark for scene structure in autonomous driving.
3 papers · 1 benchmark
A benchmark designed to evaluate MLLMs’ proficiency in understanding inter-object relationships and textual content.
3 papers · 0 benchmarks
P3 (Psychophysical Patterns Dataset)
A set of patterns used in psychophysical research to evaluate the ability of saliency algorithms to find targets distinct from distractors in orientation, color and size.
3 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.