Browse State-of-the-Art › Image-text matching
Image-text matching
102 papers with code · 1 benchmark · 1 dataset archive 2025-07-28
Image-Text Matching is a subtask within Cross-Modal Retrieval (CMR) that involves establishing associations between images and corresponding textual descriptions. The goal is to retrieve an image given a textual query or, conversely, retrieve a textual description given an image query. This task is challenging due to the heterogeneity gap between image and text data representations. Image-text matching is used in applications such as content-based image search, visual question answering, and multimodal summarization.
Assessing Brittleness of Image-Text Retrieval Benchmarks from Vision-Language Models Perspective
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
1 leaderboard table shown for this task, 1 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| CommercialAdsDataset (8 rows) | AlignCMSS | Align before Search: Aligning Ads Image to Text for Accurate... | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
1 dataset whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 102 papers with code (188 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
28 Nov 2017 20 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedIn this paper, we propose an Attentional Generative Adversarial Network (AttnGAN) that allows attention-driven, multi-stage refinement for fine-grained text-to-image generation.
-
28 Jan 2022 9 repositories listedFurthermore, performance improvement has been largely achieved by scaling up the dataset with noisy image-text pairs collected from the web, which is a suboptimal source of supervision.
-
2 Jan 2021 7 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 1 pointer-only (licence)In our experiments we feed the visual features generated by the new object detection model into a Transformer-based VL fusion model \oscar \cite{li2020oscar}, and utilize an improved approach \short\ to pre-train the VL…
-
25 Sep 2019 7 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 2 pointer-only (licence)Different from previous work that applies joint random masking to both modalities, we use conditional masking on pre-training tasks (i.
-
16 Jul 2021 6 repositories listed Syntology ran 3 of 5 samples · 2 unverified · 3 pointer-only (licence)Most existing methods employ a transformer-based multimodal encoder to jointly model visual tokens (region-based image features) and word tokens.
-
21 Mar 2018 6 repositories listed Syntology ran 7 of 16 samples · 9 unverified · 1 pointer-only (licence)Prior work either simply aggregates the similarity of all possible pairs of regions and words without attending differentially to more and less important words or regions, or uses a multi-step attentional process to…
-
13 Apr 2020 4 repositories listed Syntology ran 11 of 23 samples · 12 unverifiedLarge-scale pre-training methods of learning cross-modal representations on image-text pairs are becoming popular for vision-language tasks.
-
6 May 2023 3 repositories listed Syntology ran 0 of 9 samples · 9 unverifiedIn this paper, we present an end-to-end framework Structure-CLIP, which integrates Scene Graph Knowledge (SGK) to enhance multi-modal structured representations.
-
22 Aug 2019 3 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedWe introduce a new pre-trainable generic representation for visual-linguistic tasks, called Visual-Linguistic BERT (VL-BERT for short).
-
10 Dec 2023 2 repositories listedSince clean samples are easier distinguished by GMM with increasing noise, the memory bank can still maintain high quality at a high noise ratio.
-
24 Jul 2023 2 repositories listedThis paper aims to provide a comprehensive survey of cutting-edge research in prompt engineering on three types of vision-language models: multimodal-to-text generation models (e.
-
6 Dec 2022 2 repositories listedTo verify the effectiveness of our approach, extensive experiments are conducted on MS-COCO, CUB Captions, and Flickr30K, which are commonly used in cross-modal retrieval.
-
24 Nov 2022 2 repositories listed Syntology ran 6 of 10 samples · 4 unverifiedMedical image visual question answering (VQA) is a task to answer clinical questions, given a radiographic image, which is a challenging problem that requires a model to integrate both vision and language information.
-
21 Oct 2022 2 repositories listedIn the event that the gradients are not integrable to a valid loss function, we implement our proposed objectives such that they would directly operate in the gradient space instead of on the losses in the embedding…
-
7 Apr 2022 2 repositories listed Syntology ran 1 of 3 samples · 2 unverified · 3 pointer-only (licence)Image-Text matching (ITM) is a common task for evaluating the quality of Vision and Language (VL) models.
-
7 Oct 2020 2 repositories listed Syntology ran 3 of 11 samples · 8 unverifiedFurthermore, we introduce a new polynomial loss under the universal weighting framework, which defines a weight function for the positive and negative informative pairs respectively.
-
6 Sep 2019 2 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)It outperforms the current best method by 6.
-
12 Aug 2019 2 repositories listedWe propose a novel framework that achieves remarkable matching performance with acceptable model complexity.
-
2 Nov 2016 2 repositories listedWe propose Dual Attention Networks (DANs) which jointly leverage visual and textual attention mechanisms to capture fine-grained interplay between vision and language.
-
10 Jun 2025 1 repository listedALTA achieves superior performance in vision-language matching tasks like retrieval and zero-shot classification by adapting the pretrained vision model from masked record modeling.
-
19 Mar 2025 1 repository listedEnabling Visual Semantic Models to effectively handle multi-view description matching has been a longstanding challenge.
-
5 Mar 2025 1 repository listed Syntology ran 2 of 4 samples · 2 unverified · 4 pointer-only (licence)Our paradigm is simple and training-free, providing the first method to defend CLIP from adversarial attacks at test time, which is orthogonal to existing methods aiming to boost zero-shot adversarial robustness of CLIP.
-
2 Mar 2025 1 repository listedTo address these challenges, we propose IteRPrimE (Iterative Grad-CAM Refinement and Primary word Emphasis), which leverages a saliency heatmap through Grad-CAM from a Vision-Language Pre-trained (VLP) model for…
-
27 Feb 2025 1 repository listed Syntology ran 5 of 5 samples · 0 unverified · 5 pointer-only (licence)Contrastive Language-Image Pre-training (CLIP) models excel in zero-shot classification, yet face challenges in complex multi-object scenarios.
-
27 Feb 2025 1 repository listed Syntology ran 3 of 5 samples · 2 unverifiedTo address this problem, we propose a general Relation Consistency learning framework, namely ReCon, to accurately discriminate the true correspondences among the multimodal data and thus effectively mitigate the…
-
17 Jan 2025 1 repository listedHowever, their handcrafted generic descriptions fail to capture the diverse range of anomalies that may emerge in different objects, and simple patch-level image-text matching often struggles to localize anomalous…
-
25 Aug 2024 1 repository listedCurrently, large vision-language models have gained promising progress on many downstream tasks.
-
29 Jul 2024 1 repository listedWe show that both the Hungarian Matching and the proposed BERT-based model outperform a fuzzy string matching baseline, and we highlight inherent limitations of the matching algorithms as the target increases in size,…
-
11 Jul 2024 1 repository listedCross-modal matching has recently gained significant popularity to facilitate retrieval across multi-modal data, and existing works are highly relied on an implicit assumption that the training data pairs are perfectly…
-
18 Jun 2024 1 repository listed Syntology ran 4 of 6 samples · 2 unverified · 6 pointer-only (licence)For efficient adaptation, we treat the CLIP model as a black box and leverage the extracted features to obtain visual and textual prototypes for prediction.
Syntology lines on 16 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections