Home › Datasets › task › Image Retrieval

Image Retrieval datasets

archive 2025-07-28

87 datasets carry the task tag "Image Retrieval" (the task itself: Image Retrieval), ordered by the archive's paper count. Page 2 of 2: 39 shown of 87. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Image Retrieval datasets 49–87 of 87

The dataset consists of over 350,000 public domain patent drawings collected from the United States Patent and Trademark Office (USPTO).
4 papers · 1 benchmark
Kitchen Scenes is a multi-view RGB-D dataset of nine kitchen scenes, each containing several objects in realistic cluttered environments including a subset of objects from the BigBird dataset.
4 papers · 0 benchmarks
RetVQA (Retrieval-Based Visual Question Answering)
The RetVQA dataset is a large-scale dataset designed for Retrieval-Based Visual Question Answering (RetVQA).
4 papers · 1 benchmark
STAIR Captions is a large-scale dataset containing 820,310 Japanese captions.
4 papers · 0 benchmarks
Verse is a new dataset that augments existing multimodal datasets (COCO and TUHOI) with sense labels.
4 papers · 0 benchmarks
CV-Cities comprises $223,736$ ground panoramic images and an equal number of satellite images all accompanied by high-precision GPS coordinates.
3 papers · 1 benchmark
Most publications that aim to optimize neural networks for CBIR, train and test their models on domain specific datasets.
3 papers · 0 benchmarks
We introduce here our Large Time Lags Location (LTLL) dataset containing pictures of 25 locations captured over a range of more than 150 years.
3 papers · 1 benchmark
MMID (Massively Multilingual Image Dataset)
A large-scale multilingual corpus of images, each labeled with the word it represents.
3 papers · 0 benchmarks
SCapRepo (Google Play Screenshot Caption)
A screenshot-caption dataset containing 135k pairs of screenshots and captions extracted from Google Play.
3 papers · 0 benchmarks
SpaGBOL (Spatial-Graph-Based Orientated Cross-View Geo-Localisation)
Cross-View Geo-Localisation within urban regions is challenging in part due to the lack of geo-spatial structuring within current datasets and techniques.
3 papers · 1 benchmark
COFAR (Commonsense and Factual Reasoning in Image Search)
The COFAR (COmmonsense and FActual Reasoning) dataset is a collection of images and text queries specifically designed to challenge and evaluate image search models that aim to go beyond simple visual matching.
2 papers · 1 benchmark
A fundamental characteristic common to both human vision and natural language is their compositional nature.
2 papers · 1 benchmark
The appearance of the world varies dramatically not only from place to place but also from hour to hour and month to month.
2 papers · 1 benchmark
DyML-Animal (Dynamic Metric Learning Animal)
DyML-Animal is based on animal images selected from ImageNet-5K [1].
2 papers · 1 benchmark
DyML-Product (Dynamic Metric Learning Product)
DyML-Product is derived from iMaterialist-2019, a hierarchical online product dataset.
2 papers · 1 benchmark
DyML-Vehicle (Dynamic Metric Learning Vehicle)
DyML-Vehicle merges two vehicle re-ID datasets PKU VehicleID [1], VERI-Wild [1].
2 papers · 1 benchmark
This dataset consists of 3,710 flood images, annotated by domain experts regarding their relevance with respect to three tasks (determining the flooded area, inundation depth, water pollution).
2 papers · 0 benchmarks
INSTRE is a benchmark for INSTance-level visual object REtrieval and REcognition (INSTRE).
2 papers · 1 benchmark
A salient object subitizing image dataset of about 14K everyday images which are annotated using an online crowdsourcing marketplace.
2 papers · 0 benchmarks
A Large-Scale Chinese Image-Text Benchmark for Real-World Short Video Search Scenario
1 paper · 1 benchmark
ConQA (Conceptual Query Answering)
ConQA is a dataset created using the intersection between VisualGenome and MS-COCO.
1 paper · 2 benchmarks
The standard evaluation protocol of Cross-View Time dataset allows for certain cameras to be shared between training and testing sets.
1 paper · 1 benchmark
Provides 450, 000 relevance annotations and 53 structured queries.
1 paper · 0 benchmarks
DialogCC is a large-scale multi-modal dialogue dataset, which covers diverse real-world topics and various images per dialogue.
1 paper · 0 benchmarks
FETA Car-Manuals (FETA Car-Manuals dataset, image-text retrieval for foundation models' expert data performance.)
FETA benchmark focuses on text-to-image and image-to-text retrieval in public car manuals and sales catalogue brochures.
1 paper · 2 benchmarks
FooDI-ML (Food Drinks and groceries Images Multi Lingual)
Food Drinks and groceries Images Multi Lingual (FooDI-ML) is a dataset that contains over 1.5M unique images and over 9.5M store names, product names descriptions, and collection sections gathered from the Glovo application.
1 paper · 2 benchmarks
IAPR TC-12 (IAPR TC-12 Benchmark)
The image collection of the IAPR TC-12 Benchmark consists of 20,000 still natural images taken from locations around the world and comprising an assorted cross-section of still natural images.
1 paper · 0 benchmarks
IAW Dataset (Ikea Assembly In The Wild Dataset)
The IAW dataset contains 420 Ikea furniture pieces from 14 common categories e.g.
1 paper · 0 benchmarks
ILIAS (ILIAS: Instance-Level Image retrieval At Scale)
ILIAS is a large-scale test dataset for evaluation on Instance-Level Image retrieval At Scale.
1 paper · 0 benchmarks
IMPACT Patent (A Large-scale Integrated Multimodal Patent Analysis and Creation Dataset for Design Patents)
It is a large-scale multimodal patent dataset with detailed captions for design patent figures.
1 paper · 1 benchmark
InstaCities1M is a dataset of social media images with associated text.
1 paper · 0 benchmarks
It is composed of around 770k of color 256x256 RGB images extracted from the European Union Intellectual Property Office (EUIPO) open registry.
1 paper · 1 benchmark
MELON (Melodic Design)
1.
1 paper · 0 benchmarks
The NAVER LABS localization datasets are 5 new indoor datasets for visual localization in challenging real-world environments.
1 paper · 0 benchmarks
The PKU Sketch Re-ID dataset is constructed by National Engineering Laboratory for Video Technology (NELVT), Peking University.
1 paper · 1 benchmark
Visual Haystacks (VHs) is a "visual-centric" Needle-In-A-Haystack (NIAH) benchmark specifically designed to evaluate the capabilities of Large Multimodal Models (LMMs) in visual retrieval and reasoning over sets of unrelated images.
1 paper · 0 benchmarks
WebLI (Web Language Image)
WebLI (Web Language Image) is a web-scale multilingual image-text dataset, designed to support Google’s vision-language research, such as the large-scale pre-training for image understanding, image captioning, visual question answering,…
1 paper · 0 benchmarks
fruit-SALAD is a synthetic image dataset with 10,000 generated images of fruit depictions.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.