Home › Datasets › modality › Images

Images datasets

archive 2025-07-28

3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 28 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Images datasets 1297–1344 of 3,239

A dataset for building models that detect people Looking At Each Other (LAEO) in video sequences.
6 papers · 0 benchmarks
UIT-ViIC contains manually written captions for images from Microsoft COCO dataset relating to sports played with ball.
6 papers · 0 benchmarks
We present a further analysis of visual modality incompleteness, benchmarking latest MMEA models on our proposed dataset MMEA-UMVM.
6 papers · 3 benchmarks
Ultra-high definition benchmark (UHDBench) includes 2293 images at 2k resolution sourced from the ground-truth test sets of HRSOD, LIU4k, UAVid, UHDM, and UHRSD.
6 papers · 1 benchmark
The Urban Environments dataset is a dataset of 20 land use classes across 300 European cities paired with satellite imagery data.
6 papers · 0 benchmarks
VEDAI (Vehicle Detection in Aerial Imagery)
VEDAI is a dataset for Vehicle Detection in Aerial Imagery, provided as a tool to benchmark automatic target recognition algorithms in unconstrained environments.
6 papers · 1 benchmark
VIPER is a benchmark suite for visual perception.
6 papers · 0 benchmarks
VMRD (Visual Manipulation Relationship Dataset)
VMRD is a multi-object grasp dataset.
6 papers · 0 benchmarks
VideoCube is a high-quality and large-scale benchmark to create a challenging real-world experimental environment for Global Instance Tracking (GIT).
6 papers · 1 benchmark
VisPro dataset contains coreference annotation of 29,722 pronouns from 5,000 dialogues.
6 papers · 0 benchmarks
WHOI-Plankton is a collection of annotated plankton images.
6 papers · 0 benchmarks
WebLINX (Real-World Website Navigation with Multi-Turn)
WebLINX is a large-scale benchmark of 100K interactions across 2300 expert demonstrations of conversational web navigation.
6 papers · 1 benchmark
Wild-Time is a benchmark of 5 datasets that reflect temporal distribution shifts arising in a variety of real-world applications, including patient prognosis and news classification.
6 papers · 0 benchmarks
XImageNet-12 (XIMAGENET-12: An Explainable AI Benchmark Dataset for Model Robustness Evaluation)
Enlarge the dataset to understand how image background effect the Computer Vision ML model.
6 papers · 1 benchmark
A new dataset with significant occlusions related to object manipulation.
6 papers · 0 benchmarks
Dataset for large-scale yoga pose recognition with 82 classes.
6 papers · 0 benchmarks
The ZS-F-VQA dataset is a new split of the F-VQA dataset for zero-shot problem.
6 papers · 1 benchmark
25kTrees (Individual Tree Crown Annotations)
Manual crown delineation of individual trees in two countries: Denmark and Finland.
5 papers · 0 benchmarks
2DeteCT (2DeteCT - A large 2D expandable, trainable, experimental Computed Tomography dataset for machine learning)
Maximilian B.
5 papers · 0 benchmarks
AHP (Amodal Human Perception)
The AHP dataset consists of 56,599 images in total which are collected from several large-scale instance segmentation and detection datasets, including COCO, VOC (w/ SBD), LIP, Objects365 and OpenImages.
5 papers · 0 benchmarks
ARID (Autonomous Robot Indoor Dataset)
ARID is a large-scale, multi-view object dataset collected with an RGB-D camera mounted on a mobile robot.
5 papers · 0 benchmarks
Dataset to address the problem of detecting people Looking At Each Other (LAEO) in video sequences.
5 papers · 0 benchmarks
BIRD (Blocksworld Image Reasoning Dataset)
Blocksworld Image Reasoning Dataset (BIRD) contains images of wooden blocks in different configurations, and the sequence of moves to rearrange one configuration to the other.
5 papers · 1 benchmark
BTS3.1 (Expanding Accurate Person Recognition to New Altitudes and Ranges: The BRIAR Dataset)
Large, multimodal biometric dataset: It contains still images and videos of over 1,000 people captured at various ranges (up to 1,000 meters) and elevations (up to 400 meters) using a diverse set of cameras (commercial, military-grade,…
5 papers · 2 benchmarks
BUP20 (Sweet Pepper 2020 University of Bonn)
Video sequences from a glasshouse environment in Campus Kleinaltendorf(CKA), University of Bonn, captured by PATHoBot, a glasshouse monitoring robot.
5 papers · 0 benchmarks
This dataset consists of images and annotations in Bengali.
5 papers · 1 benchmark
The BanglaWriting dataset contains single-page handwritings of 260 individuals of different personalities and ages.
5 papers · 1 benchmark
BRATS 2014 is a brain tumor segmentation dataset.
5 papers · 1 benchmark
This brain anatomy segmentation dataset has 1300 2D US scans for training and 329 for testing.
5 papers · 1 benchmark
CBC (Complete Blood Count)
The complete blood count (CBC) dataset contains 360 blood smear images along with their annotation files splitting into Training, Testing, and Validation sets.
5 papers · 0 benchmarks
The dataset offers tag and mask annotations for image-text pairs from the CC3M validation set.
5 papers · 2 benchmarks
CLEVR-X is a dataset that extends the CLEVR dataset with natural language explanations in the context of VQA.
5 papers · 1 benchmark
CLIC (Challenge on Learned Image Compression)
CLIC is a dataset for learned image compression.
5 papers · 0 benchmarks
A database of over 1.4 billion 3x3 convolution filters extracted from hundreds of diverse CNN models with relevant meta information.
5 papers · 0 benchmarks
Request access: cadpath.ai@impdiagnostics.com The CRC dataset contains 1133 colorectal biopsy and polypectomy slides and is the result of our ongoing efforts to contribute to CRC diagnosis with a reference dataset.
5 papers · 0 benchmarks
CURE-OR (Challenging Unreal and Real Environments for Object Recognition)
CURE-OR is a large-scale, controlled, and multi-platform object recognition dataset denoted as Challenging Unreal and Real Environments for Object Recognition.
5 papers · 0 benchmarks
Cata7 is the first cataract surgical instrument dataset for semantic segmentation.
5 papers · 0 benchmarks
Chinese Text in the Wild is a dataset of Chinese text with about 1 million Chinese characters from 3850 unique ones annotated by experts in over 30000 street view images.
5 papers · 0 benchmarks
Cube++ is a novel dataset for the color constancy problem that continues on the Cube+ dataset.
5 papers · 1 benchmark
DIMO (Dataset of Industrial Metal Objects)
The Industrial Metal Objects dataset is a diverse dataset of industrial metal objects.
5 papers · 0 benchmarks
DREAM-dataset (Deep Robot-to-camera Extrinsics for Articulated Manipulators)
The DREAM dataset is introduce by the paper "Camera-to-Robot Pose Estimation from a Single Image" (ICRA 2020).
5 papers · 1 benchmark
This basketball dataset was acquired under the Walloon region project DeepSport, using the Keemotion system installed in multiple arenas.
5 papers · 0 benchmarks
DurLAR (A High-Fidelity 128-Channel LiDAR Dataset with Panoramic Ambient and Reflectivity Imagery)
DurLAR is a high-fidelity 128-channel 3D LiDAR dataset with panoramic ambient (near infrared) and reflectivity imagery for multi-modal autonomous driving applications.
5 papers · 0 benchmarks
ELAS is a dataset for lane detection.
5 papers · 0 benchmarks
FS2K is a high-quality Facial Sketch Synthesis (FSS).
5 papers · 0 benchmarks
The Flick Cropping Dataset consists of high quality cropping and pairwise ranking annotations used to evaluate the performance of automatic image cropping approaches.
5 papers · 0 benchmarks
Contains 8k flickr Images with captions.
5 papers · 2 benchmarks
We construct the ForgeryNet dataset, an extremely large face forgery dataset with unified annotations in image- and video-level data across four tasks: 1) Image Forgery Classification, including two-way (real / fake), three-way (real /…
5 papers · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.