Browse State-of-the-Art › Image Retrieval
Image Retrieval
835 papers with code · 56 benchmarks · 87 datasets archive 2025-07-28
Image Retrieval is a fundamental and long-standing computer vision task that involves finding images similar to a given query from a large database. It is often considered a form of fine-grained, instance-level classification. The task is integral to image recognition alongside classification and cross-modal retrieval. By leveraging visual similarity and other criteria, image retrieval enables users to efficiently discover relevant images, making it a crucial tool in applications such as search and recommendation.
Extending CLIP for Category-to-image Retrieval in E-commerce
( Image credit: DELF )
Description from the archive archive 2025-07-28; Papers-with-Code links inside it are rewritten to this site.
Benchmarks archive 2025-07-28
56 leaderboard tables shown for this task, 56 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 56 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
87 datasets whose archive record lists this task, ordered by the archive's paper count. 30 shown of 87 until expanded.
Subtasks archive 2025-07-28
10 subtasks in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 835 papers with code (2,239 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
29 Apr 2021 32 repositories listed Syntology ran 5 of 20 samples · 15 unverified · 2 pointer-only (licence)In this paper, we question if self-supervised learning provides new properties to Vision Transformer (ViT) that stand out compared to convolutional networks (convnets).
-
14 Apr 2023 26 repositories listed Syntology ran 21 of 46 samples · 25 unverified · 12 pointer-only (licence)The recent breakthroughs in natural language processing for model pretraining on large quantities of data have opened the way for similar foundation models in computer vision.
-
23 Oct 2017 24 repositories listed Syntology ran 5 of 8 samples · 3 unverified · 5 pointer-only (licence)The dataset was collected with three goals in mind: (i) to have both a large number of identities and also a large number of images for each identity; (ii) to cover a large range of pose, age and ethnicity; and (iii) to…
-
30 Jan 2023 17 repositories listed Syntology ran 4 of 8 samples · 4 unverified · 1 pointer-only (licence)The cost of vision-and-language pre-training has become increasingly prohibitive due to end-to-end training of large-scale models.
-
25 Feb 2020 16 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 1 pointer-only (licence)This paper provides a pair similarity optimization viewpoint on deep feature learning, aiming to maximize the within-class similarity sₚ and minimize the between-class similarity sₙ.
-
3 Nov 2017 14 repositories listed Syntology ran 3 of 24 samples · 21 unverified · 2 pointer-only (licence)We show that both hard-positive and hard-negative examples, selected by exploiting the geometry and the camera positions available from the 3D models, enhance the performance of particular-object retrieval.
-
23 Nov 2015 14 repositories listed Syntology ran 0 of 12 samples · 12 unverifiedWe tackle the problem of large scale visual place recognition, where the task is to quickly and accurately recognize the location of a given query photograph.
-
17 Apr 2023 13 repositories listed Syntology ran 16 of 51 samples · 35 unverifiedInstruction tuning large language models (LLMs) using machine-generated instruction-following data has improved zero-shot capabilities on new tasks, but the idea is less explored in the multimodal field.
-
19 Dec 2016 13 repositories listed Syntology ran 3 of 12 samples · 9 unverifiedWe propose an attentive local feature descriptor suitable for large-scale image retrieval, referred to as DELF (DEep Local Feature).
-
15 Mar 2023 11 repositories listed Syntology ran 2 of 5 samples · 3 unverified · 1 pointer-only (licence)We report the development of GPT-4, a large-scale, multimodal model which can accept image and text inputs and produce text outputs.
-
6 Aug 2019 11 repositories listed Syntology ran 10 of 34 samples · 24 unverified · 34 pointer-only (licence)We present ViLBERT (short for Vision-and-Language BERT), a model for learning task-agnostic joint representations of image content and natural language.
-
18 Jul 2017 10 repositories listed Syntology ran 2 of 5 samples · 3 unverified · 4 pointer-only (licence)We present a new technique for learning visual-semantic embeddings for cross-modal retrieval.
-
4 Mar 2017 9 repositories listedThis paper extends fully-convolutional neural networks (FCN) for the clothing parsing problem.
-
17 May 2016 9 repositories listedState-of-the-art methods for zero-shot visual recognition formulate learning as a joint embedding problem of images and side information.
-
26 Mar 2019 7 repositories listed Syntology ran 0 of 5 samples · 5 unverifiedRecent studies in image retrieval task have shown that ensembling different models and combining multiple global descriptors lead to performance improvement.
-
5 Feb 2021 6 repositories listed Syntology ran 1 of 4 samples · 3 unverified · 1 pointer-only (licence)Vision-and-Language Pre-training (VLP) has improved performance on various joint vision-and-language downstream tasks.
-
21 Mar 2018 6 repositories listed Syntology ran 7 of 16 samples · 9 unverified · 1 pointer-only (licence)Prior work either simply aggregates the similarity of all possible pairs of regions and words without attending differentially to more and less important words or regions, or uses a multi-step attentional process to…
-
23 Jun 2017 6 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedIn addition, we show that a simple margin based loss is sufficient to outperform all other loss functions.
-
18 Nov 2015 6 repositories listed Syntology ran 0 of 4 samples · 4 unverifiedRecently, image representation built upon Convolutional Neural Network (CNN) has been shown to provide effective descriptors for image search, outperforming pre-CNN features as short-vector representations.
-
6 Aug 2021 5 repositories listed Syntology ran 10 of 23 samples · 13 unverifiedComponents orthogonal to the global image representation are then extracted from the local information.
-
3 Apr 2020 5 repositories listedGLDv2 is the largest such dataset to date by a large margin, including over 5M images and 200k distinct instance labels.
-
15 Mar 2020 5 repositories listed Syntology ran 2 of 7 samples · 5 unverifiedWe employ the Earth Mover's Distance (EMD) as a metric to compute a structural distance between dense image representations to determine image relevance.
-
14 Jan 2020 5 repositories listedImage retrieval is the problem of searching an image database for items that are similar to a query image.
-
14 Dec 2019 5 repositories listedThis suggests that the features of instances computed at preceding iterations can be used to considerably approximate their features extracted by the current model.
-
5 Dec 2019 5 repositories listed Syntology ran 4 of 20 samples · 16 unverified · 20 pointer-only (licence)Much of vision-and-language research focuses on a small but diverse set of independent tasks and supporting datasets often studied in isolation; however, the visually-grounded language understanding skills required for…
-
5 Oct 2019 5 repositories listed Syntology ran 0 of 20 samples · 20 unverifiedThis work presents Kornia -- an open source computer vision library which consists of a set of differentiable routines and modules to solve generic computer vision problems.
-
11 Sep 2019 5 repositories listed Syntology ran 2 of 4 samples · 2 unverified · 2 pointer-only (licence)The set of triplet constraints has to be sampled within the mini-batch.
-
17 Nov 2018 5 repositories listed Syntology ran 3 of 12 samples · 9 unverifiedIn this paper, we propose the Batch DropBlock (BDB) Network which is a two branch network composed of a conventional ResNet-50 as the global branch and a feature dropping branch.
-
8 Apr 2016 5 repositories listed Syntology ran 0 of 2 samples · 2 unverified · 2 pointer-only (licence)Convolutional Neural Networks (CNNs) achieve state-of-the-art performance in many computer vision tasks.
-
20 Dec 2014 5 repositories listedThe zero-shot paradigm exploits vector-based word representations extracted from text corpora with unsupervised methods to learn general mapping functions from other feature spaces onto word space, where the words…
Syntology lines on 24 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections