Home › Datasets › modality › Images
Images datasets
archive 2025-07-28
3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 62 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Images datasets 2929–2976 of 3,239
A visible-light and thermal-infrared images dataset for dual-spectrum depth estimation.
1 paper · 0 benchmarks
VISOR is a dataset of pixel annotations and a benchmark suite for segmenting hands and active objects in egocentric video.
1 paper · 0 benchmarks
VME & CDSI (Vehicles in the Middle East (VME) & Car Detection in Satellite Imagery (CDSI) datasets)
Vehicles in the Middle East (VME) dataset, designed explicitly for vehicle detection in high-resolution satellite images from Middle Eastern countries.
1 paper · 1 benchmark
VQA-MHUG is a 49-participant dataset of multimodal human gaze on both images and questions during visual question answering (VQA) collected using a high-speed eye tracker.
1 paper · 0 benchmarks
VR-Folding contains garment meshes of 4 categories from CLOTH3D dataset, namely Shirt, Pants, Top and Skirt.
1 paper · 0 benchmarks
VTQA (Visual Text Question Answering)
VTQA is a dataset containing open-ended questions about image-text pairs.
1 paper · 0 benchmarks
Vastextures (Vast Dataset for textures and PBR materials)
VasTexture is a free giant repository of textures and PBR materials extracted from real-world images.
1 paper · 0 benchmarks
The ViCoS Towel Dataset is a state-of-the-art benchmark for grasp point localization on cloth objects, specifically towels.
1 paper · 1 benchmark
We propose the first standardized benchmark in multimodal continual learning for video data, defining protocols for training and metrics for evaluation.
1 paper · 0 benchmarks
Introduction This dataset was gathered during the Vid2RealHRI study of humans’ perception of robots' intelligence in the context of an incidental Human-Robot encounter.
1 paper · 0 benchmarks
Dataset Introduction This dataset leverages VideoDB's Public Collection to offer a diverse range of videos featuring text-containing scenes.
1 paper · 1 benchmark
A pioneering dataset for vignette removal.
1 paper · 0 benchmarks
VinDr-PCXR is an open, large-scale pediatric chest X-ray dataset for interpretation of common thoracic diseases in children.
1 paper · 0 benchmarks
Vinegar Fly is a pose estimation dataset for fruit flies.
1 paper · 1 benchmark
Molecules represent tokens of the language of chemistry, which underlies not only chemistry itself, but also scientific fields that use chemical information such as pharmacy, material science, and molecular biology.
1 paper · 0 benchmarks
VisA-AC is a refined benchmark based on the VisA dataset, tailored for the task of anomaly classification—distinguishing between different types of anomalies rather than simply detecting whether an image is anomalous.
1 paper · 1 benchmark
VisAlign is a dataset for measuring AI-human visual alignment in terms of image classification, a fundamental task in machine perception.
1 paper · 0 benchmarks
VisArgs is a densely annotated benchmark for visual argument understanding.
1 paper · 0 benchmarks
VisCon-100K is a dataset specially designed to facilitate fine-tuning of vision-language models (VLMs) by leveraging interleaved image-text web documents.
1 paper · 0 benchmarks
VisCon-100K is a dataset specially designed to facilitate fine-tuning of vision-language models (VLMs) by leveraging interleaved image-text web documents.
1 paper · 0 benchmarks
Despite its importance for assessing the effectiveness of communicating information visually, fine-grained recallability of information visualisations has not been studied quantitatively so far.
1 paper · 0 benchmarks
Dataset for testing the ability of Vision Language Models (LVM) to recognize and match 3D objects of the exact same 3D shapes but with different orientation/materials/textures/ environments and light conditions.
1 paper · 0 benchmarks
Visual Commonsense Immorality benchmark is a benchmark designed to evaluate commonsense immorality.
1 paper · 0 benchmarks
Visual Haystacks (VHs) is a "visual-centric" Needle-In-A-Haystack (NIAH) benchmark specifically designed to evaluate the capabilities of Large Multimodal Models (LMMs) in visual retrieval and reasoning over sets of unrelated images.
1 paper · 0 benchmarks
This data contains about 2500 trajectories (with images and actions) of a Sawyer robot interacting with various objects.
1 paper · 0 benchmarks
VizWiz-FewShot is a a few-shot localization dataset originating from photographers who authentically were trying to learn about the visual content in the images they took.
1 paper · 0 benchmarks
The WIDER Attribute dataset is a human attribute recognition dataset with human attribute and image event annotations.
1 paper · 0 benchmarks
The WORC database consists in total of 930 patients composed of six datasets gathered at the Erasmus MC, consisting of patients with: 1) well-differentiated liposarcoma or lipoma (115 patients); 2) desmoid-type fibromatosis or extremity…
1 paper · 0 benchmarks
Wood plate bark removal processing is critical for ensuring the quality of wood processing and its products.
1 paper · 0 benchmarks
WebGen-Bench WebGen-Bench is created to benchmark LLM-based agent's ability to generate websites from scratch.
1 paper · 0 benchmarks
The Western Mediterranean Wetlands Bird Dataset is a collection of birds' vocalizations of different lengths that primarily consists of 5,795 labelled audio clips derived from 1,098 recordings, totalling 201.6 minutes or 12,096 seconds…
1 paper · 0 benchmarks
WiFiCam dataset for through-wall imaging based on WiFi channel state information.
1 paper · 0 benchmarks
WiRLD (Wikidata Reference Logo Dataset)
The Wikidata Reference Logo Dataset (WiRLD), a comprehensive collection of reference logos specifically designed to address the challenges of large-scale logo identification.
1 paper · 0 benchmarks
WiRLD_ (Wikidata Reference Logo Dataset)
The Wikidata Reference Logo Dataset (WiRLD), a comprehensive collection of reference logos specifically designed to address the challenges of large-scale logo identification.
1 paper · 0 benchmarks
This Wider-Test-200 dataset is introduced in the following paper: "Towards Unsupervised Blind Face Restoration using Diffusion Prior" Please visit our website and refer to our paper for more information on the dataset and our method:…
1 paper · 0 benchmarks
The Wiki-Flick Event dataset for cross-modal event retrieval is a well-labelled but weakly-aligned dataset collected for cross-modality event retrieval.
1 paper · 0 benchmarks
WikiChurches is a dataset for architectural style classification, consisting of 9,485 images of church buildings.
1 paper · 0 benchmarks
Wikipedia Webpage 2M (WikiWeb2M) is a multimodal open source dataset consisting of over 2 million English Wikipedia articles.
1 paper · 0 benchmarks
WildQA is a video understanding dataset of videos recorded in outside settings.
1 paper · 1 benchmark
WildestFaces is tailored to study cross-domain recognition under a variety of adverse conditions.
1 paper · 0 benchmarks
WinSyn (WinSyn: A High Resolution Testbed for Synthetic Data)
75k photos of windows + 21k synthetic renders of building windows.
1 paper · 0 benchmarks
We present the World Wide Dishes dataset which seeks to assess disparities in representations of food through a decentralised data collection effort to gather perspectives directly from people with a wide variety of backgrounds from around…
1 paper · 0 benchmarks
Vision Language Models (VLMs) often struggle with culture-specific knowledge, particularly in languages other than English and in underrepresented cultural contexts.
1 paper · 0 benchmarks
The X-MARS dataset proposes new splits for the MARS dataset, to allow for cross-evaluation with the Market-1501 dataset without training and test overlap between the two datasets.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Synthetic dataset intended for benchmarking disentanglement frameworks.
1 paper · 0 benchmarks
The scales of the data accessible through internet search engines can reach hundreds of millions, or even billions.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.