Home › Datasets › task › Image Retrieval

Image Retrieval datasets

archive 2025-07-28

87 datasets carry the task tag "Image Retrieval" (the task itself: Image Retrieval), ordered by the archive's paper count. Page 1 of 2: 48 shown of 87. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Image Retrieval datasets 1–48 of 87

description withheld: archive row vandalised before snapshot
16,145 papers · 91 benchmarks
The COCO (Common Objects in Context) dataset is a large-scale object detection, segmentation, and captioning dataset.
11,922 papers · 77 benchmarks
CUB-200-2011 (Caltech-UCSD Birds-200-2011)
The Caltech-UCSD Birds-200-2011 (CUB-200-2011) dataset is the most widely-used dataset for fine-grained visual categorization task.
2,235 papers · 47 benchmarks
The Flickr30k dataset contains 31,000 images collected from Flickr, together with 5 reference sentences provided by human annotators.
880 papers · 9 benchmarks
The iNaturalist 2017 dataset (iNat) contains 675,170 training and validation images from 5,089 natural fine-grained categories.
603 papers · 12 benchmarks
DeepFashion is a dataset containing around 800K diverse fashion images with their rich annotations (46 categories, 1,000 descriptive attributes, bounding boxes and landmark information) ranging from well-posed product images to…
397 papers · 5 benchmarks
The NUS-WIDE dataset contains 269,648 images with a total of 5,018 tags collected from Flickr.
348 papers · 3 benchmarks
YFCC100M is a that dataset contains a total of 100 million media objects, of which approximately 99.2 million are photos and 0.8 million are videos, all of which carry a Creative Commons license.
243 papers · 0 benchmarks
Stanford Online Products (SOP) dataset has 22,634 classes with 120,053 product images.
231 papers · 5 benchmarks
In-Shop (In-shop Clothes Retrieval Benchmark)
In-shop Clothes Retrieval Benchmark evaluates the performance of in-shop Clothes Retrieval.
154 papers · 2 benchmarks
Oxford5k (Oxford Buildings)
Oxford5K is the Oxford Buildings Dataset, which contains 5062 images collected from Flickr.
137 papers · 1 benchmark
The image dataset TinyImages contains 80 million images of size 32×32 collected from the Internet, crawling the words in WordNet.
103 papers · 0 benchmarks
Fashion IQ support and advance research on interactive fashion image retrieval.
102 papers · 6 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
90 papers · 13 benchmarks
WIT (Wikipedia-based Image Text)
Wikipedia-based Image Text (WIT) Dataset is a large multimodal multilingual dataset.
79 papers · 1 benchmark
misc @inproceedings{RITAC18, author = {Radenovi\'{c}, F.
73 papers · 3 benchmarks
InLoc is a dataset with reference 6DoF poses for large-scale indoor localization.
67 papers · 1 benchmark
We propose Localized Narratives, a new form of multimodal image annotations connecting vision and language.
63 papers · 3 benchmarks
CIRR (Compose Image Retrieval on Real-life images)
Composed Image Retrieval (or, Image Retreival conditioned on Language Feedback) is a relatively new retrieval task, where an input query consists of an image and short textual description of how to modify the image.
61 papers · 3 benchmarks
The Quick Draw Dataset is a collection of 50 million drawings across 345 categories, contributed by players of the game Quick, Draw!.
50 papers · 0 benchmarks
Oxford105k is the combination of the Oxford5k dataset and 99782 negative images crawled from Flickr using 145 most popular tags.
44 papers · 0 benchmarks
CARS196 is composed of 16,185 car images of 196 classes.
43 papers · 4 benchmarks
Fashion-Gen consists of 293,008 high definition (1360 x 1360 pixels) fashion images paired with item descriptions provided by professional stylists.
36 papers · 0 benchmarks
CIRCO (Composed Image Retrieval on Common Objects in context)
CIRCO (Composed Image Retrieval on Common Objects in context) is an open-domain benchmarking dataset for Composed Image Retrieval (CIR) based on real-world images from COCO 2017 unlabeled set.
35 papers · 1 benchmark
This is the second version of the Google Landmarks dataset (GLDv2), which contains images annotated with labels representing human-made and natural landmarks.
35 papers · 4 benchmarks
Spot-the-diff is a dataset consisting of 13,192 image pairs along with corresponding human provided text annotations stating the differences between the two images.
35 papers · 0 benchmarks
DeepFashion2 is a versatile benchmark of four tasks including clothes detection, pose estimation, segmentation, and retrieval.
30 papers · 0 benchmarks
ShoeV2 is a dataset of 2,000 photos and 6648 sketches of shoes.
30 papers · 0 benchmarks
The Google Landmarks dataset contains 1,060,709 images from 12,894 landmarks, and 111,036 additional query images.
24 papers · 0 benchmarks
COCO-CN is a bilingual image description dataset enriching MS-COCO with manually written Chinese sentences and tags.
21 papers · 1 benchmark
PhotoChat, the first dataset that casts light on the photo sharing behavior in online messaging.
20 papers · 2 benchmarks
OVEN (Open-domain Visual Entity Recognition)
In this project, we formally present the task of Open-domain Visual Entity recognitioN (OVEN), where a model need to link an image onto a Wikipedia entity with respect to a text query.
19 papers · 1 benchmark
ICFG-PEDES (Identity-Centric and Fine-Grained Person Description Dataset)
One large-scale database for Text-to-Image Person Re-identification, i.e., Text-based Person Retrieval.
16 papers · 3 benchmarks
A dataset containing 404,683 shop photos collected from 25 different online retailers and 20,357 street photos, providing a total of 39,479 clothing item matches between street and shop photos.
15 papers · 1 benchmark
VegFru is a domain-specific dataset for fine-grained visual categorization.
14 papers · 0 benchmarks
Visual Madlibs is a dataset consisting of 360,001 focused natural language descriptions for 10,738 images.
13 papers · 0 benchmarks
ImageCoDe (Image Retrieval from Contextual Descriptions)
Given 10 minimally contrastive (highly similar) images and a complex description for one of them, the task is to retrieve the correct image.
11 papers · 1 benchmark
AmsterTime (AmsterTime: A Visual Place Recognition Benchmark Dataset for Severe Domain Shift)
AmsterTime dataset offers a collection of 2,500 well-curated images matching the same scene from a street view matched to historical archival image data from Amsterdam city.
9 papers · 3 benchmarks
Consists of 330,000 sketches and 204,000 photos spanning across 110 categories.
9 papers · 0 benchmarks
SketchyScene is a large-scale dataset of scene sketches to advance research on sketch understanding at both the object and scene level.
9 papers · 0 benchmarks
ETH SfM (ETH Structure-from-Motion)
The ETH SfM (structure-from-motion) dataset is a dataset for 3D Reconstruction.
8 papers · 0 benchmarks
Flickr30k-CNA (Flickr30k-Chinese All)
Former Flickr30k-CN translates the training and validation sets of Flickr30k using machine translation and manually translates the test set.
7 papers · 1 benchmark
The Hotels-50K dataset consists of over 1 million images from 50,000 different hotels around the world.
7 papers · 0 benchmarks
Large Scale Composed Image Retrieval (LaSCo) is a new dataset for Composed Image Retrieval (CoIR), x10 times larger than current ones.
7 papers · 1 benchmark
The Oxford-Affine dataset is a small dataset containing 8 scenes with sequence of 6 images per scene.
7 papers · 0 benchmarks
The Retrieval-SFM dataset is used for instance image retrieval.
7 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
6 papers · 1 benchmark
This dataset contains 114 individuals including 1824 images captured from two disjoint camera views.
5 papers · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.