Home › Datasets › modality › Images

Images datasets

archive 2025-07-28

3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 19 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Images datasets 865–912 of 3,239

Food2K is a large food recognition dataset with 2,000 categories and over 1 million images.
14 papers · 0 benchmarks
HiFiMask (CASIA-SURF HiFiMask)
HiFiMask is a large-scale High-Fidelity Mask dataset, namely CASIA-SURF HiFiMask (briefly HiFiMask).
14 papers · 0 benchmarks
The IIIT5K dataset contains 5,000 text instance images: 2,000 for training and 3,000 for testing.
14 papers · 1 benchmark
IMC PhotoTourism (Image Matching Challenge Phototourism)
Dataset provided by the Image Matching Workshop https://www.cs.ubc.ca/research/image-matching-challenge/current/
14 papers · 1 benchmark
The KITTI-Depth dataset includes depth maps from projected LiDAR point clouds that were matched against the depth estimation from the stereo cameras.
14 papers · 0 benchmarks
MHP (Multiple-Human Parsing)
The MHP dataset contains multiple persons captured in real-world scenes with pixel-level fine-grained semantic annotations in an instance-aware setting.
14 papers · 3 benchmarks
MIntRec is a novel dataset for multimodal intent recognition.
14 papers · 1 benchmark
MMPD (Multi-Domain Mobile Video Physiology Dataset)
The Multi-domain Mobile Video Physiology Dataset (MMPD), comprising 11 hours(1152K frames) of recordings from mobile phones of 33 subjects.
14 papers · 0 benchmarks
MVOR (Multi-View Operating Room)
Multi-View Operating Room (MVOR) is a dataset recorded during real clinical interventions.
14 papers · 0 benchmarks
OVAD benchmark (Open-Vocabulary Attribute Detection)
Vision-language modeling has enabled open-vocabulary tasks where predictions can be queried using any text prompt in a zero-shot manner.
14 papers · 3 benchmarks
Open-Platypus is a family of fine-tuned and merged Large Language Models (LLMs) that achieves the strongest performance and currently stands at first place in HuggingFace's Open LLM Leaderboard.
14 papers · 0 benchmarks
PASS (Pictures without humAns for Self-Supervision)
PASS is a large-scale image dataset, containing 1.4 million images, that does not include any humans and which can be used for high-quality pretraining while significantly reducing privacy concerns.
14 papers · 0 benchmarks
A new large scale plane geometry problem solving dataset called PGPS9K, labeled both fine-grained diagram annotation and interpretable solution program.
14 papers · 1 benchmark
PPR10K (Portrait Photo Retouching dataset)
PPR10K is a dataset for portrait photo retouching (PPR), which aims to enhance the visual quality of a collection of flat-looking portrait photos.
14 papers · 0 benchmarks
The dataset is based on the original MNIST dataset.
14 papers · 0 benchmarks
REFUGE Challenge (Retinal Fundus Glaucoma Challenge)
REFUGE Challenge provides a data set of 1200 fundus images with ground truth segmentations and clinical glaucoma labels, currently the largest existing one.
14 papers · 4 benchmarks
SIBR (SIBR Dataset for VIE in the Wild)
SIBR是面向自然场景视觉信息抽取的数据集。 1)SIBR总的有1000张图片,400张测试,600张训练,包括中文、英文两种语言。 2)包含images.zip、label.zip、train.txt、test.txt四个文件,images.zip、label.zip中包含所有图片和标签,通过train.txt和test.txt区分训练和测试。…
14 papers · 1 benchmark
Purpose Medical imaging has become increasingly important in diagnosing and treating oncological patients, particularly in radiotherapy.
14 papers · 0 benchmarks
TransNAS-Bench-101 is a Neural Architecture Search (NAS) benchmark dataset containing network performance across seven tasks, covering classification, regression, pixel-level prediction, and self-supervised tasks.
14 papers · 0 benchmarks
VegFru is a domain-specific dataset for fine-grained visual categorization.
14 papers · 0 benchmarks
Created for MVS tasks and is a large-scale multi-view aerial dataset generated from a highly accurate 3D digital surface model produced from thousands of real aerial images with precise camera parameters.
14 papers · 0 benchmarks
This dataset accompanies our paper on synthesizing the 3D Ken Burns effect from a single image.
13 papers · 0 benchmarks
Amazon Baby (Amazon Baby 5-core)
This dataset includes reviews (ratings, text, helpfulness votes), product metadata (descriptions, category information, price, brand, and image features), and links (also viewed/also bought graphs).
13 papers · 2 benchmarks
BRATS 2016 is a brain tumor segmentation dataset.
13 papers · 0 benchmarks
BrixIA (BrixIA Covid-19)
BrixIA Covid-19 is a large dataset of CXR images corresponding to the entire amount of images taken for both triage and patient monitoring in sub-intensive and intensive care units during one month (between March 4th and April 4th 2020) of…
13 papers · 0 benchmarks
Detecting vehicles and representing their position and orientation in the three dimensional space is a key technology for autonomous driving.
13 papers · 3 benchmarks
DNA-Rendering is a large-scale, high-fidelity repository of human performance data for neural actor rendering.
13 papers · 0 benchmarks
The Dayton dataset is a dataset for ground-to-aerial (or aerial-to-ground) image translation, or cross-view image synthesis.
13 papers · 4 benchmarks
The EgoDexter dataset provides both 2D and 3D pose annotations for 4 testing video sequences with 3190 frames.
13 papers · 0 benchmarks
Flare7K, the first nighttime flare removal dataset, which is generated based on the observation and statistic of real-world nighttime lens flares.
13 papers · 1 benchmark
Logo-2K+:A Large-Scale Logo Dataset for Scalable Logo Classification The Logo-2K+ dataset contains a diverse range of logo classes from real-world logo images.
13 papers · 0 benchmarks
PISC (People in Social Context)
The People in Social Context (PISC) dataset is a dataset that focuses on social relationships.
13 papers · 1 benchmark
SEN12MS-CR is a multi-modal and mono-temporal data set for cloud removal.
13 papers · 1 benchmark
Next generation task-oriented dialog systems need to understand conversational contexts with their perceived surroundings, to effectively help users in the real-world multimodal environment.
13 papers · 2 benchmarks
SynWoodScape (Synthetic Surround-view Fisheye Camera Dataset for Autonomous Driving)
SynWoodScape is a synthetic version of the surround-view dataset covering many of its weaknesses and extending it.
13 papers · 0 benchmarks
TTPLA (Transmission Towers and Power Lines (TTPLA))
TTPLA is a public dataset which is a collection of aerial images on Transmission Towers (TTs) and Power Lines (PLs).
13 papers · 0 benchmarks
The COLOSSEUM (The COLOSSEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation)
To realize effective large-scale, real-world robotic applications, we must evaluate how well our robot policies adapt to changes in environmental conditions.
13 papers · 1 benchmark
This dataset encompasses a diverse range of tactile features that are instrumental in bifurcating various material properties.
13 papers · 0 benchmarks
UFDD (Unconstrained Face Detection Dataset)
Unconstrained Face Detection Dataset (UFDD) aims to fuel further research in unconstrained face detection.
13 papers · 0 benchmarks
The VQA-CP dataset was constructed by reorganizing VQA v2 such that the correlation between the question type and correct answer differs in the training and test splits.
13 papers · 1 benchmark
ViQuAE is a dataset for KVQAE (Knowledge-based Visual Question Answering about named Entities), a task which consists in answering questions about named entities grounded in a visual context using a Knowledge Base.
13 papers · 0 benchmarks
Visual Madlibs is a dataset consisting of 360,001 focused natural language descriptions for 10,738 images.
13 papers · 0 benchmarks
The Watch-n-Patch dataset was created with the focus on modeling human activities, comprising multiple actions in a completely unsupervised setting.
13 papers · 0 benchmarks
A Multi-Task 4D Radar-Camera Fusion Dataset for Autonomous Driving on Water Surfaces description of the dataset WaterScenes, the first multi-task 4D radar-camera fusion dataset on water surfaces, which offers data from multiple sensors,…
13 papers · 2 benchmarks
Ecoset, an ecologically motivated image dataset, is a large-scale image dataset designed for human visual neuroscience, which consists of over 1.5 million images from 565 basic-level categories.
13 papers · 0 benchmarks
The ACNE04 dataset includes 3756 Chinese face images with Acne.
12 papers · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.