Home › Datasets › modality › Images
Images datasets
archive 2025-07-28
3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 30 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Images datasets 1393–1440 of 3,239
PSI (IUPUI-CSRC Pedestrian Situated Intent)
The IUPUI-CSRC Pedestrian Situated Intent (PSI) benchmark dataset has two innovative labels besides comprehensive computer vision annotations.
5 papers · 0 benchmarks
PedX is a large-scale multi-modal collection of pedestrians at complex urban intersections.
5 papers · 0 benchmarks
Contains 10,000 fine-grained SKU-level products frequently bought by online customers in JD.com.
5 papers · 0 benchmarks
This dataset arises from the READ project (Horizon 2020).
5 papers · 1 benchmark
The evaluation of object detection models is usually performed by optimizing a single metric, e.g.
5 papers · 1 benchmark
RITE (Retinal Images vessel Tree Extraction)
The RITE (Retinal Images vessel Tree Extraction) is a database that enables comparative studies on segmentation or classification of arteries and veins on retinal fundus images, which is established based on the public available DRIVE…
5 papers · 2 benchmarks
ROF (Real World Occluded Faces)
ROF is a dataset for occluded face recognition that contains faces with both upper face occlusion, due to sunglasses, and lower face occlusion, due to masks.
5 papers · 0 benchmarks
RefRef (RefRef: A Synthetic Dataset and Benchmark for Reconstructing Refractive and Reflective Objects)
RefRef is a synthetic dataset and benchmark designed for the task of reconstructing scenes with complex refractive and reflective objects.
5 papers · 1 benchmark
The semantic line (SEL) dataset contains 1,750 outdoor images in total, which are split into 1,575 training and 175 testing images.
5 papers · 1 benchmark
SODA-D is a large-scale dataset tailored for small object detection in driving scenario, which is built on top of MVD dataset and owned data, where the former is a dataset dedicated to pixel-level understanding of street scenes, and the…
5 papers · 1 benchmark
SYNS-Patches dataset, which is a subset of SYNS.
5 papers · 0 benchmarks
Sewer-ML is a sewer defect dataset.
5 papers · 0 benchmarks
Simitate is a hybrid benchmarking suite targeting the evaluation of approaches for imitation learning.
5 papers · 0 benchmarks
SkyEyeGPT: Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language Model
5 papers · 0 benchmarks
SpaceNet 1: Building Detection v1 is a dataset for building footprint detection.
5 papers · 2 benchmarks
- Games dataset containing 100,000 Gameplay Images of 175 Video Games across 10 Sports Genres - AMERICAN FOOTBALL, BASKETBALL, BIKE RACING, CAR RACING, FIGHTING, HOCKEY, SOCCER, TABLE TENNIS, TENNIS.
5 papers · 2 benchmarks
StreetStyle is a large-scale dataset of photos of people annotated with clothing attributes, and use this dataset to train attribute classifiers via deep learning.
5 papers · 0 benchmarks
The SUGARCREPE++ dataset evaluates the sensitivity of vision language models (VLMs) and unimodal language models (ULMs) to semantic and lexical alterations.
5 papers · 0 benchmarks
SyRIP is a hybrid synthetic and real infant pose (SyRIP) dataset with small yet diverse real infant images as well as generated synthetic infant poses and (2) a multi-stage invariant representation learning strategy that could transfer the…
5 papers · 0 benchmarks
TAS500 is a semantic segmentation dataset for autonomous driving in unstructured environments.
5 papers · 0 benchmarks
This data collection consists of images acquired during chemoradiotherapy of 20 locally-advanced, non-small cell lung cancer patients.
5 papers · 0 benchmarks
TRANCE (Transformation Driven Visual Reasoning)
TRANCE extends CLEVR by asking a uniform question, i.e.
5 papers · 0 benchmarks
Tencent ML-Images is a large open-source multi-label image database, including 17,609,752 training and 88,739 validation image URLs, which are annotated with up to 11,166 categories.
5 papers · 0 benchmarks
A Dense-text Image Benchmark to evaluate large generation model's ability on text generation.
5 papers · 1 benchmark
Thyroid is a dataset for detection of thyroid diseases, in which patients diagnosed with hypothyroid or subnormal are anomalies against normal patients.
5 papers · 1 benchmark
Twitter-MEL is a multimodal entity linking (MEL) dataset built from Twitter.
5 papers · 0 benchmarks
Twitter100k is a large-scale dataset for weakly supervised cross-media retrieval.
5 papers · 0 benchmarks
UPLight is an underwater RGB-Polarization multimodal semantic segmentation dataset with 12 typical underwater semantic classes.
5 papers · 1 benchmark
UruDendro (UruDendro, a public dataset of cross-section images of pinus taeda)
UruDendro is a database of wood cross section images of commercially grown Pinus taeda trees from northern Uruguay.
5 papers · 1 benchmark
The dataset uses VGG-Sound which consists of 10s clips collected from YouTube for 309 sound classes.
5 papers · 0 benchmarks
We present the VIS30K dataset, a collection of 29,689 images that represents 30 years of figures and tables from each track of the IEEE Visualization conference series (Vis, SciVis, InfoVis, VAST).
5 papers · 0 benchmarks
VQA-VS (a new VQA benchmark considering Varying Shortcuts)
The current OOD benchmark VQA-CP v2 only considers one type of shortcut (from question type to answer) and thus still cannot guarantee that the modelrelies on the intended solution rather than a solution specific to this shortcut.
5 papers · 0 benchmarks
VinDr-RibCXR is a benchmark dataset for automatic segmentation and labeling of individual ribs from chest X-ray (CXR) scans.
5 papers · 0 benchmarks
WGISD (Embrapa Wine Grape Instance Segmentation Dataset)
Embrapa Wine Grape Instance Segmentation Dataset (WGISD) contains grape clusters properly annotated in 300 images and a novel annotation methodology for segmentation of complex objects in natural images.
5 papers · 0 benchmarks
WWW Crowd provides 10,000 videos with over 8 million frames from 8,257 diverse scenes, therefore offering a comprehensive dataset for the area of crowd understanding.
5 papers · 0 benchmarks
This dataset is collected via the WinoGAViL game to collect challenging vision-and-language associations.
5 papers · 2 benchmarks
ZeroWaste is a dataset for automatic waste detection and segmentation.
5 papers · 0 benchmarks
aiMotive dataset is a multimodal dataset for robust autonomous driving with long-range perception.
5 papers · 1 benchmark
mTVR is a large-scale multilingual video moment retrieval dataset, containing 218K English and Chinese queries from 21.8K TV show video clips.
5 papers · 0 benchmarks
7,672 human written natural language navigation instructions for routes in OpenStreetMap with a focus on visual landmarks.
5 papers · 2 benchmarks
Description: 1,995 People Face Images Data (Asian race).
4 papers · 0 benchmarks
Provides a large-scale synthetic dataset which contains accurate ground truth depth of various photo-realistic scenes.
4 papers · 0 benchmarks
The 3DNet dataset is a free resource for object class recognition and 6DOF pose estimation from point cloud data.
4 papers · 0 benchmarks
A Game Of Sorts is a collaborative image ranking task.
4 papers · 0 benchmarks
ADE-OoD is a public benchmark for dense out-of-distribution detection in general natural images.
4 papers · 1 benchmark
ALTO (Aerial-view Large-scale Terrain-Oriented)
ALTO is a vision-focused dataset for the development and benchmarking of Visual Place Recognition and Localization methods for Unmanned Aerial Vehicles.
4 papers · 0 benchmarks
AMA (Articulated Mesh Animation)
Articulated Mesh Animation (AMA) is a real-world dataset containing 10 mesh sequences depicting 3 different humans performing various actions
4 papers · 0 benchmarks
ARMBench is a large-scale, object-centric benchmark dataset for robotic manipulation in the context of a warehouse.
4 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.