Browse State-of-the-Art › Image Retrieval

Image Retrieval

835 papers with code · 56 benchmarks · 87 datasets archive 2025-07-28

Computer Vision

Image Retrieval is a fundamental and long-standing computer vision task that involves finding images similar to a given query from a large database. It is often considered a form of fine-grained, instance-level classification. The task is integral to image recognition alongside classification and cross-modal retrieval. By leveraging visual similarity and other criteria, image retrieval enables users to efficiently discover relevant images, making it a crucial tool in applications such as search and recommendation.

Extending CLIP for Category-to-image Retrieval in E-commerce

( Image credit: DELF )

Description from the archive archive 2025-07-28; Papers-with-Code links inside it are rewritten to this site.

Benchmarks archive 2025-07-28

56 leaderboard tables shown for this task, 56 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 56 until expanded.

DatasetBest model (first row in archive order)PaperCodeSyntologyCompare
ROxford (Hard) (23 rows) SuperGlobal Global Features are All You Need for Image Retrieval and Reranking code Syntology ran 13 of 18 samples · 5 unverified Compare
ROxford (Medium) (23 rows) AMES AMES: Asymmetric and Memory-Efficient Similarity Estimation for... code Syntology ran 13 of 13 samples · 0 unverified Compare
RParis (Hard) (23 rows) AMES AMES: Asymmetric and Memory-Efficient Similarity Estimation for... code Syntology ran 13 of 13 samples · 0 unverified Compare
RParis (Medium) (23 rows) AMES AMES: Asymmetric and Memory-Efficient Similarity Estimation for... code Syntology ran 13 of 13 samples · 0 unverified Compare
CREPE (Compositional REPresentation Evaluation) (22 rows) ViT-L-14 (LAION400M) CREPE: Can Vision-Language Foundation Models Reason Compositionally? code Syntology ran 3 of 3 samples · 0 unverified Compare
Fashion IQ (22 rows) DQU-CIR — — — Compare
Flickr30K 1K test (18 rows) X-VLM (base) Multi-Grained Vision Language Pre-Training: Aligning Texts with... code Syntology ran 1 of 1 samples · 0 unverified Compare
CIRR (17 rows) TMCIR TMCIR: Token Merge Benefits Composed Image Retrieval — — Compare
SOP (14 rows) Unicom+ViT-L@336px Unicom: Universal and Compact Representation Learning for Image Retrieval code Syntology ran 3 of 6 samples · 3 unverified Compare
Flickr30k-CN (11 rows) InternVL-G-FT InternVL: Scaling up Vision Foundation Models and Aligning for... code Syntology ran 2 of 2 samples · 0 unverified Compare
Oxf5k (11 rows) Offline Diffusion Efficient Image Retrieval via Decoupling Diffusion into Online and... code — Compare
iNaturalist (10 rows) Unicom+ViT-L@336px Unicom: Universal and Compact Representation Learning for Image Retrieval code Syntology ran 3 of 6 samples · 3 unverified Compare
COCO-CN (9 rows) CN-CLIP (ViT-H/14) Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese code Syntology ran 4 of 7 samples · 3 unverified Compare
Flickr30k (9 rows) BLIP-2 ViT-G (zero-shot, 1K test set) BLIP-2: Bootstrapping Language-Image Pre-training with Frozen... code Syntology ran 4 of 8 samples · 4 unverified Compare
MUGE Retrieval (9 rows) CN-CLIP (ViT-H/14) Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese code Syntology ran 4 of 7 samples · 3 unverified Compare
Oxf105k (9 rows) Offline Diffusion Efficient Image Retrieval via Decoupling Diffusion into Online and... code — Compare
CARS196 (8 rows) CGD (MG/SG) Combination of Multiple Global Descriptors for Image Retrieval code Syntology ran 0 of 5 samples · 5 unverified Compare
CUB-200-2011 (8 rows) CGD (MG/SG) Combination of Multiple Global Descriptors for Image Retrieval code Syntology ran 0 of 5 samples · 5 unverified Compare
In-Shop (7 rows) CGD (SG/GS) Combination of Multiple Global Descriptors for Image Retrieval code Syntology ran 0 of 5 samples · 5 unverified Compare
Par106k (7 rows) Offline Diffusion Efficient Image Retrieval via Decoupling Diffusion into Online and... code — Compare
Par6k (7 rows) Offline Diffusion Efficient Image Retrieval via Decoupling Diffusion into Online and... code — Compare
COCO (Common Objects in Context) (6 rows) BLIP-2 ViT-G (fine-tuned) BLIP-2: Bootstrapping Language-Image Pre-training with Frozen... code Syntology ran 4 of 8 samples · 4 unverified Compare
AmsterTime (5 rows) DINOv2 distilled (ViT-L/14 frozen) DINOv2: Learning Robust Visual Features without Supervision code Syntology ran 21 of 46 samples · 25 unverified Compare
ConQA Conceptual (5 rows) CLIP Does the Performance of Text-to-Image Retrieval Models Generalize... code — Compare
ConQA Descriptive (5 rows) CLIP Does the Performance of Text-to-Image Retrieval Models Generalize... code — Compare
PhotoChat (5 rows) PaCE PaCE: Unified Multi-modal Dialogue Pre-training with Progressive... code Syntology ran 0 of 1 samples · 1 unverified Compare
DeepFashion - Consumer-to-shop (4 rows) CTL Model (ResNet50-IBN-A, 320x320) On the Unreasonable Effectiveness of Centroids in Image Retrieval code — Compare
DeepPatent (4 rows) SwinV2 Patent image retrieval using transformer-based deep metric learning code — Compare
Google Landmarks Dataset v2 (retrieval, testing) (4 rows) AMES AMES: Asymmetric and Memory-Efficient Similarity Estimation for... code Syntology ran 13 of 13 samples · 0 unverified Compare
24/7 Tokyo (3 rows) HED-N-GAN Dark Side Augmentation: Generating Diverse Night Examples for... code Syntology ran 2 of 3 samples · 1 unverified Compare
Exact Street2Shop (3 rows) CTL Model (ResNet50-IBN-A, 320x320) On the Unreasonable Effectiveness of Centroids in Image Retrieval code — Compare
Google Landmarks Dataset v2 (retrieval, validation) (3 rows) UNICOM-ViT-L-14-512px Unicom: Universal and Compact Representation Learning for Image Retrieval code Syntology ran 3 of 6 samples · 3 unverified Compare
LaSCo (3 rows) CASE Data Roaming and Quality Assessment for Composed Image Retrieval code — Compare
MSCOCO (3 rows) HADA HADA: A Graph-based Amalgamation Framework in Image-text Retrieval code — Compare
AIC-ICC (2 rows) ERNIE-ViL2.0 ERNIE-ViL 2.0: Multi-view Contrastive Learning for Image-Text Pre-training code — Compare
CBVS (2 rows) UniCLP CBVS: A Large-Scale Chinese Image-Text Benchmark for Real-World... code — Compare
INRIA Holidays (2 rows) MultiGrain R50 @ 800 MultiGrain: a unified image embedding for classes and instances code Syntology ran 1 of 4 samples · 3 unverified Compare
Oxford5k (2 rows) GNN-Reranking Understanding Image Retrieval Re-Ranking: A Graph Neural Network... code — Compare
Paris6k (2 rows) IME layer Iterative Manifold Embedding Layer Learned by Incomplete Data for... code — Compare
street2shop - topwear (2 rows) Ranknet Retrieving Similar E-Commerce Images Using Deep Learning code — Compare
WIT (2 rows) WIT-ALL WIT: Wikipedia-based Image Text Dataset for Multimodal... code — Compare
CIFAR-10 (1 row) Custom: 3 conv + 2 fcn Deep Supervised Hashing for Fast Image Retrieval code — Compare
COFAR (1 row) KRAMT COFAR: Commonsense and Factual Reasoning in Image Search code — Compare
DeepFashion (1 row) RCCapsNet Fashion Image Retrieval with Capsule Networks code — Compare
FETA Car-Manuals (1 row) FETA's CLIP-MIL (Many-Shot Image-to-text) FETA: Towards Specializing Foundation Models for Expert Task Applications code — Compare
FooDI-ML (Global) (1 row) ADAPT-I2T FooDI-ML: a large multi-language dataset of food, drinks and... code — Compare
FooDI-ML (Spain) (1 row) ADAPT-I2T FooDI-ML: a large multi-language dataset of food, drinks and... code — Compare
ICFG-PEDES (1 row) SSAN Semantically Self-Aligned Network for Text-to-Image Part-aware... code — Compare
ImageCoDe (1 row) ContextualCLIP Image Retrieval from Contextual Descriptions code Syntology ran 2 of 2 samples · 0 unverified Compare
INSTRE (1 row) IME layer Iterative Manifold Embedding Layer Learned by Incomplete Data for... code — Compare
Localized Narratives (1 row) OPT OPT: Omni-Perception Pre-Trainer for Cross-Modal Understanding and... code — Compare
NUS-WIDE (1 row) DTQ Deep Triplet Quantization code Syntology ran 0 of 7 samples · 7 unverified Compare
PKU-Reid (1 row) IHDA Instance-level Heterogeneous Domain Adaptation for Limited-labeled... code — Compare
PKU SketchRe-ID Dataset (1 row) IHDA Instance-level Heterogeneous Domain Adaptation for Limited-labeled... code — Compare
ROxford Medium without fine-tuning (1 row) HesAff–rSIFT–VLAD Revisiting Oxford and Paris: Large-Scale Image Retrieval Benchmarking code — Compare
RUC-CAS-WenLan (1 row) CMCL WenLan: Bridging Vision and Language by Large-Scale Multi-Modal... code — Compare

Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.

Libraries

Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.

Datasets archive 2025-07-28

87 datasets whose archive record lists this task, ordered by the archive's paper count. 30 shown of 87 until expanded.

Subtasks archive 2025-07-28

10 subtasks in the archive's task tree.

Most implemented papers archive 2025-07-28

30 shown of 835 papers with code (2,239 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.

Syntology lines on 24 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections