Papers › Integrating Language Guidance Into Image-Text Matching for Correcting False Negatives

Integrating Language Guidance Into Image-Text Matching for Correcting False Negatives

24 Mar 2023IEEE Transactions on Multimedia 2023 3archive 2025-07-28

Zheng Li, Caili Guo, Zerun Feng, Jenq-Neng Hwang, Zhongtian Du

Image-Text Matching (ITM) aims to establish the correspondence between images and sentences. ITM is fundamental to various vision and language understanding tasks. However, there are limitations in the way existing ITM benchmarks are constructed. The ITM benchmark collects pairs of images and sentences during construction. Therefore, only samples that are paired at collection are annotated as positive. All other samples are annotated as negative. Many correlations are missed in these samples that are annotated as negative. For example, a sentence matches only one image at the time of collection. Only this image is annotated as positive for the sentence. All other images are annotated as negative. However, these negative images may contain images that correspond to the sentences. These mislabeled samples are called false negatives . Existing ITM models are optimized based on annotations containing mislabels, which can introduce noise during training. In this paper, we propose an ITM framework integrating Language Guidance ( LG ) for correcting false negatives. A language pre-training model is introduced into the ITM framework to identify false negatives. To correct false negatives, we propose language guidance loss, which adaptively corrects the locations of false negatives in the visual-semantic embedding space. Extensive experiments on two ITM benchmarks show that our method can improve the performance of existing ITM models. To verify the performance of correcting false negatives, we conduct further experiments on ECCV Caption. ECCV Caption is a verified dataset where false negatives in annotations have been corrected. The experimental results show that our method can recall more relevant false negatives.

PaperPDFCode

Code

AAA-Zheng/LG_ITM officialpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Cross-modal retrieval with noisy correspondenceImage-text matchingSentenceText Matching

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Cross-modal retrieval with noisy correspondence COCO-Noisy LG-ITM-SGARF Image-to-text R@1 79.6 #6 of 17 Archive leaderboard report
Cross-modal retrieval with noisy correspondence COCO-Noisy LG-ITM-SGARF Image-to-text R@10 98.5 #6 of 17 Archive leaderboard report
Cross-modal retrieval with noisy correspondence COCO-Noisy LG-ITM-SGARF Image-to-text R@5 96.5 #6 of 17 Archive leaderboard report
Cross-modal retrieval with noisy correspondence COCO-Noisy LG-ITM-SGARF R-Sum 524.9 #6 of 17 Archive leaderboard report
Cross-modal retrieval with noisy correspondence COCO-Noisy LG-ITM-SGARF Text-to-image R@1 64.4 #6 of 17 Archive leaderboard report
Cross-modal retrieval with noisy correspondence COCO-Noisy LG-ITM-SGARF Text-to-image R@10 95.9 #6 of 17 Archive leaderboard report
Cross-modal retrieval with noisy correspondence COCO-Noisy LG-ITM-SGARF Text-to-image R@5 90.0 #6 of 17 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections