Home › Datasets › modality › Images
Images datasets
archive 2025-07-28
3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 34 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Images datasets 1585–1632 of 3,239
UTA-RLDD (University of Texas at Arlington Real-Life Drowsiness Dataset)
Consists of around 30 hours of video, with contents ranging from subtle signs of drowsiness to more obvious ones.
4 papers · 0 benchmarks
The ukiyo-e faces dataset comprises of 5209 images of faces from ukiyo-e prints.
4 papers · 0 benchmarks
VASR (Visual Analogies of Situation Recognition)
Visual Analogies of Situation Recognition (VASR) is a dataset for visual analogical mapping, adapting the classical word-analogy task into the visual domain.
4 papers · 1 benchmark
Covers 5 generic driving scenarios, with a total of 25 distinct action classes.
4 papers · 0 benchmarks
Visuelle 2.0 is a dataset containing real data for 5355 clothing products of the retail fast-fashion Italian company, Nuna Lie.
4 papers · 2 benchmarks
VIVA (Vision for Intelligent Vehicles and Applications)
The VIVA challenge’s dataset is a multimodal dynamic hand gesture dataset specifically designed with difficult settings of cluttered background, volatile illumination, and frequent occlusion for studying natural human activities in…
4 papers · 2 benchmarks
A Large Vision-Language Model Knowledge Editing Benchmark
4 papers · 0 benchmarks
VOT2019 is a Visual Object Tracking benchmark for short-term tracking in RGB.
4 papers · 1 benchmark
Verse is a new dataset that augments existing multimodal datasets (COCO and TUHOI) with sense labels.
4 papers · 0 benchmarks
The VizWiz-VQA-Grounding dataset is a dataset that visually grounds answers to visual questions asked by people with visual impairments.
4 papers · 0 benchmarks
WHU-Hi (Wuhan UAV-borne hyperspectral image)
WHU-Hi dataset (Wuhan UAV-borne hyperspectral image) is collected and shared by the RSIDEA research group of Wuhan University, and it could serve as a benchmark dataset for precise crop classification and hyperspectral image classification…
4 papers · 0 benchmarks
WHU-RS19 is a set of satellite images exported from Google Earth, which provides high-resolution satellite images up to 0.5 m.
4 papers · 0 benchmarks
WIKIPerson is a high-quality human-annotated visual person linking dataset based on Wikipedia.
4 papers · 0 benchmarks
XTD10 is a dataset for cross-lingual image retrieval and tagging consisting of the MSCOCO2014 caption test dataset annotated in 7 languages that were collected using a crowdsourcing platform.
4 papers · 0 benchmarks
The eBDtheque database is a selection of one hundred comic pages from America, Japan (manga) and Europe.
4 papers · 1 benchmark
iShape is an irregular shape dataset for instance segmentation.
4 papers · 1 benchmark
The dataset is designed specifically to solve a range of computer vision problems (2D-3D tracking, posture) faced by biologists while designing behavior studies with animals.
3 papers · 0 benchmarks
ACL-Fig is a large-scale automatically annotated corpus consisting of 112,052 scientific figures extracted from 56K research papers in the ACL Anthology.
3 papers · 0 benchmarks
ADE-Affordance is a new dataset that builds upon ADE20k, which contains annotations enabling such rich visual reasoning.
3 papers · 0 benchmarks
The AND Dataset contains 13700 handwritten samples and 15 corresponding expert examined features for each sample.
3 papers · 1 benchmark
AS-V2 (The All-Seeing Dataset v2)
We propose a novel task, termed Relation Conversation (ReC), which unifies the formulation of text generation, object localization, and relation comprehension.
3 papers · 0 benchmarks
The Action-Camera Parking Dataset contains 293 images captured at a roughly 10-meter height using a GoPro Hero 6 camera.
3 papers · 1 benchmark
AdobeVFR real (Adobe Visual Font Recognition real-world images dataset)
Subset of AdobeVFR.
3 papers · 1 benchmark
AdobeVFR syn (Adobe Visual Font Recognition synthetic dataset)
Subset of AdobeVFR.
3 papers · 1 benchmark
The Aircraft Context Dataset, a composition of two inter-compatible large-scale and versatile image datasets focusing on manned aircraft and UAVs, is intended for training and evaluating classification, detection and segmentation models in…
3 papers · 0 benchmarks
We design an all-day semantic segmentation benchmark all-day CityScapes.
3 papers · 1 benchmark
Amateur Drawings is a dataset collected via the public demo of Animated Drawings, containing over 178,000 amateur drawings and corresponding user-accepted character bounding boxes, segmentation masks, and joint location annotations.
3 papers · 0 benchmarks
This dataset includes reviews (ratings, text, helpfulness votes), product metadata (descriptions, category information, price, brand, and image features), and links (also viewed/also bought graphs).
3 papers · 1 benchmark
A dataset for 2D pose estimation of anime/manga images.
3 papers · 0 benchmarks
AnoVox is a large-scale benchmark for ANOmaly detection in autonomous driving.
3 papers · 0 benchmarks
For a detailed description, we refer to Section 3 in our research article.
3 papers · 0 benchmarks
Atlas is a dataset for e-commerce clothing product categorization.
3 papers · 0 benchmarks
BAFMD (Bias-Aware Face Mask Detection Dataset)
BAFMD contains images posted on Twitter during the pandemic from around the world with more images from underrepresented race and age groups to mitigate the problem for the face mask detection task.
3 papers · 0 benchmarks
This image set is part of a high-throughput chemical screen on U2OS cells, with examples of 200 bioactive compounds.
3 papers · 0 benchmarks
BRIGHT is the first open-access, globally distributed, event-diverse multimodal dataset specifically curated to support AI-based disaster response.
3 papers · 1 benchmark
This dataset contains Bangla handwritten numerals, basic characters and compound characters.
3 papers · 2 benchmarks
23,000 cropped images of tree bark, for 23 species of trees around Quebec City, Canada.
3 papers · 0 benchmarks
This dataset contains images of individual hand-written Bengali characters.
3 papers · 1 benchmark
BnB is a large-scale and diverse in-domain VLN (Vision and Language Navigation) dataset.
3 papers · 0 benchmarks
Bongard-OpenWorld is a new benchmark for evaluating real-world few-shot reasoning for machine vision.
3 papers · 1 benchmark
✔️Abstract A Brain tumor is considered as one of the aggressive diseases, among children and adults.
3 papers · 0 benchmarks
CASIA-Face-Africa is a face image database which contains 38,546 images of 1,183 African subjects.
3 papers · 0 benchmarks
CC-19 is a small new dataset related to the latest family of coronavirus i.e.
3 papers · 0 benchmarks
We introduce FUNSD-r and CORD-r in Token Path Prediction, the revised VrD-NER datasets to reflect the real-world scenarios of NER on scanned VrDs.
3 papers · 1 benchmark
CORSMAL is a dataset for estimating the position and orientation in 3D (or 6D pose) of an object from a single view.
3 papers · 0 benchmarks
CTC (COCO-Text Captioned)
A dataset that allows exploration of cross-modal retrieval where images contain scene-text instances.
3 papers · 0 benchmarks
CUFS (CUHK Face Sketch Database)
CUHK Face Sketch database (CUFS) is for research on face sketch synthesis and face sketch recognition.
3 papers · 1 benchmark
CUHK Image Cropping is a dataset for image cropping.
3 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.