Browse State-of-the-Art › Image-to-Text Retrieval
Image-to-Text Retrieval
37 papers with code · 8 benchmarks · 8 datasets archive 2025-07-28
Image-text retrieval is the process of retrieving relevant images based on textual descriptions or finding corresponding textual descriptions for a given image. This task is interdisciplinary, combining techniques from computer vision, and natural language processing. The primary challenge lies in bridging the semantic gap — the difference between how visual data is represented in images and how humans describe that information using language. To address this, many methods focus on learning a shared embedding space where both images and text can be represented in a comparable way, allowing their similarities to be measured and facilitating more accurate retrieval.
Source: Extending CLIP for Category-to-Image Retrieval in E-commerce
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
8 leaderboard tables shown for this task, 8 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
8 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 37 papers with code (59 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
26 Feb 2021 82 repositories listed Syntology ran 16 of 20 samples · 4 unverified · 16 pointer-only (licence)State-of-the-art computer vision systems are trained to predict a fixed set of predetermined object categories.
-
30 Jan 2023 17 repositories listed Syntology ran 4 of 8 samples · 4 unverified · 1 pointer-only (licence)The cost of vision-and-language pre-training has become increasingly prohibitive due to end-to-end training of large-scale models.
-
27 Mar 2023 11 repositories listed Syntology ran 12 of 29 samples · 17 unverified · 25 pointer-only (licence)We propose a simple pairwise Sigmoid loss for Language-Image Pre-training (SigLIP).
-
16 Jul 2021 6 repositories listed Syntology ran 3 of 5 samples · 2 unverified · 3 pointer-only (licence)Most existing methods employ a transformer-based multimodal encoder to jointly model visual tokens (region-based image features) and word tokens.
-
8 Dec 2021 4 repositories listedState-of-the-art vision and vision-and-language models rely on large-scale visio-linguistic pretraining for obtaining good performance on a variety of downstream tasks.
-
13 Apr 2020 4 repositories listed Syntology ran 11 of 23 samples · 12 unverifiedLarge-scale pre-training methods of learning cross-modal representations on image-text pairs are becoming popular for vision-language tasks.
-
7 Dec 2014 4 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedOur approach leverages datasets of images and their sentence descriptions to learn about the inter-modal correspondences between language and visual data.
-
27 Jan 2022 3 repositories listed Syntology ran 8 of 21 samples · 13 unverifiedOur benchmark enables the evaluation of multilingual multimodal models for transfer learning, not only in a zero-shot setting, but also in newly defined few-shot learning setups.
-
21 Dec 2023 2 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)However, the progress in vision and vision-language foundation models, which are also critical elements of multi-modal AGI, has not kept pace with LLMs.
-
10 Dec 2023 2 repositories listedSince clean samples are easier distinguished by GMM with increasing noise, the memory bank can still maintain high quality at a high noise ratio.
-
15 Aug 2023 2 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)In this work, we design the first vision-language dataset distillation method, building on the idea of trajectory matching.
-
18 May 2023 2 repositories listed Syntology ran 2 of 7 samples · 5 unverifiedIn this work, we explore a scalable way for building a general representation model toward unlimited modalities.
-
31 Jan 2023 2 repositories listed Syntology ran 0 of 2 samples · 2 unverifiedReal-world data contains a vast amount of multimodal information, among which vision and language are the two most representative modalities.
-
11 Jan 2023 2 repositories listedIt is not easy to propose a new model with a novel architecture and intensively train it on a massive dataset with many GPUs to surpass many SOTA models, which are already available to use on the Internet.
-
6 Dec 2022 2 repositories listedTo verify the effectiveness of our approach, extensive experiments are conducted on MS-COCO, CUB Captions, and Flickr30K, which are commonly used in cross-modal retrieval.
-
12 Nov 2022 2 repositories listed Syntology ran 2 of 11 samples · 9 unverifiedIn this work, we present a conceptually simple and effective method to train a strong bilingual/multilingual multimodal representation model.
-
1 Jul 2021 2 repositories listedIn this paper, we propose an Omni-perception Pre-Trainer (OPT) for cross-modal understanding and generation, by jointly modeling visual, text and audio resources.
-
11 Mar 2021 2 repositories listedWe further construct a large Chinese multi-source image-text dataset called RUC-CAS-WenLan for pre-training our BriVL model.
-
21 Dec 2017 2 repositories listedFinally, a comprehensive review is presented on the proposed data set to fully advance the task of remote sensing caption.
-
10 Jun 2025 1 repository listedALTA achieves superior performance in vision-language matching tasks like retrieval and zero-shot classification by adapting the pretrained vision model from masked record modeling.
-
30 Jul 2024 1 repository listedWe refer to this bias in associating an activity with the gender of its actual performer in an image or text as the Gender-Activity Binding (GAB) bias and analyze how this bias is internalized in VLMs.
-
10 Jul 2024 1 repository listedHerein, we hypothesize that the pre-trained vision-language models can be utilized for quantitative histopathology image analysis through a simple image-to-text retrieval.
-
14 Jun 2024 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedThe novelty of BiVLC is to add a synthetic hard negative image generated from the synthetic text, resulting in two image-to-text retrieval examples (one for each image) and, more importantly, two text-to-image retrieval…
-
28 Apr 2024 1 repository listed Syntology ran 1 of 6 samples · 5 unverifiedWe hope this method can facilitate the efficient application of large models to a wide range of downstream tasks while significantly reducing the resource consumption.
-
1 Jan 2024 1 repository listedWe propose a novel Linguistic-Aware Patch Slimming (LAPS) framework for fine-grained alignment which explicitly identifies redundant visual patches with language supervision and rectifies their semantic and spatial…
-
29 Sep 2023 1 repository listed Syntology ran 12 of 19 samples · 7 unverifiedIn this paper, we propose a novel Prototype-based Aleatoric Uncertainty Quantification (PAU) framework to provide trustworthy predictions by quantifying the uncertainty arisen from the inherent data ambiguity.
-
24 Jul 2023 1 repository listed Syntology ran 2 of 3 samples · 1 unverifiedIn this paper, we present a prototype representation learning framework incorporating both global and local alignment between medical images and reports.
-
20 Jun 2023 1 repository listed Syntology ran 0 of 7 samples · 7 unverifiedMoreover, we present an image-text paired dataset in the field of remote sensing (RS), RS5M, which has 5 million RS images with English descriptions.
-
27 May 2023 1 repository listed Syntology ran 2 of 4 samples · 2 unverifiedTo pursue more efficient vision-language Transformers, this paper introduces Cross-Guided Ensemble of Tokens (CrossGET), a general acceleration framework for vision-language Transformers.
-
21 Apr 2023 1 repository listedThe reason is that a large amount of images and texts in the benchmarks are coarse-grained.
Syntology lines on 18 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections