Papers › Cross-Modal Retrieval with Partially Mismatched Pairs

Cross-Modal Retrieval with Partially Mismatched Pairs

22 Feb 2023IEEE Transactions on Pattern Analysis and Machine Intelligence 2023 2archive 2025-07-28

Peng Hu, Zhenyu Huang, Dezhong Peng, Xu Wang, Xi Peng

In this paper, we study a challenging but less-touched problem in cross-modal retrieval, i.e., partially mismatched pairs (PMPs). Specifically, in real-world scenarios, a huge number of multimedia data (e.g., the Conceptual Captions dataset) are collected from the Internet, and thus it is inevitable to wrongly treat some irrelevant cross-modal pairs as matched. Undoubtedly, such a PMP problem will remarkably degrade the cross-modal retrieval performance. To tackle this problem, we derive a unified theoretical Robust Cross-modal Learning framework (RCL) with an unbiased estimator of the cross-modal retrieval risk, which aims to endow the cross-modal retrieval methods with robustness against PMPs. In detail, our RCL adopts a novel complementary contrastive learning paradigm to address the following two challenges, i.e., the overfitting and underfitting issues. On the one hand, our method only utilizes the negative information which is much less likely false compared with the positive information, thus avoiding the overfitting issue to PMPs. However, these robust strategies could induce underfitting issues, thus making training models more difficult. On the other hand, to address the underfitting issue brought by weak supervision, we present to leverage of all available negative pairs to enhance the supervision contained in the negative information. Moreover, to further improve the performance, we propose to minimize the upper bounds of the risk to pay more attention to hard samples. To verify the effectiveness and robustness of the proposed method, we carry out comprehensive experiments on five widely-used benchmark datasets compared with nine state-of-the-art approaches w.r.t. the image-text and video-text retrieval tasks. The code is available at https://github.com/penghu-cs/RCL .

PaperPDFCode

Code

penghu-cs/RCL mentioned in paperpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Contrastive LearningCross-Modal RetrievalCross-modal retrieval with noisy correspondenceRetrievalText RetrievalVideo-Text Retrieval

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Cross-modal retrieval with noisy correspondence CC152K RCL-SGRAF Image-to-text R@1 41.7 #12 of 15 Archive leaderboard report
Cross-modal retrieval with noisy correspondence CC152K RCL-SGRAF Image-to-text R@10 73.6 #12 of 15 Archive leaderboard report
Cross-modal retrieval with noisy correspondence CC152K RCL-SGRAF Image-to-text R@5 66.0 #12 of 15 Archive leaderboard report
Cross-modal retrieval with noisy correspondence CC152K RCL-SGRAF R-Sum 364.4 #12 of 15 Archive leaderboard report
Cross-modal retrieval with noisy correspondence CC152K RCL-SGRAF Text-to-image R@1 41.6 #12 of 15 Archive leaderboard report
Cross-modal retrieval with noisy correspondence CC152K RCL-SGRAF Text-to-image R@10 75.1 #12 of 15 Archive leaderboard report
Cross-modal retrieval with noisy correspondence CC152K RCL-SGRAF Text-to-image R@5 66.4 #12 of 15 Archive leaderboard report
Cross-modal retrieval with noisy correspondence COCO-Noisy RCL-SGR Image-to-text R@1 77.0 #17 of 17 Archive leaderboard report
Cross-modal retrieval with noisy correspondence COCO-Noisy RCL-SGR Image-to-text R@10 98.1 #17 of 17 Archive leaderboard report
Cross-modal retrieval with noisy correspondence COCO-Noisy RCL-SGR Image-to-text R@5 95.5 #17 of 17 Archive leaderboard report
Cross-modal retrieval with noisy correspondence COCO-Noisy RCL-SGR R-Sum 515.5 #17 of 17 Archive leaderboard report
Cross-modal retrieval with noisy correspondence COCO-Noisy RCL-SGR Text-to-image R@1 61.3 #17 of 17 Archive leaderboard report
Cross-modal retrieval with noisy correspondence COCO-Noisy RCL-SGR Text-to-image R@10 94.8 #17 of 17 Archive leaderboard report
Cross-modal retrieval with noisy correspondence COCO-Noisy RCL-SGR Text-to-image R@5 88.8 #17 of 17 Archive leaderboard report
Cross-modal retrieval with noisy correspondence Flickr30K-Noisy RCL-SGR Image-to-text R@1 74.2 #16 of 16 Archive leaderboard report
Cross-modal retrieval with noisy correspondence Flickr30K-Noisy RCL-SGR Image-to-text R@10 96.9 #16 of 16 Archive leaderboard report
Cross-modal retrieval with noisy correspondence Flickr30K-Noisy RCL-SGR Image-to-text R@5 91.8 #16 of 16 Archive leaderboard report
Cross-modal retrieval with noisy correspondence Flickr30K-Noisy RCL-SGR R-Sum 487.2 #16 of 16 Archive leaderboard report
Cross-modal retrieval with noisy correspondence Flickr30K-Noisy RCL-SGR Text-to-image R@1 55.6 #16 of 16 Archive leaderboard report
Cross-modal retrieval with noisy correspondence Flickr30K-Noisy RCL-SGR Text-to-image R@10 87.5 #16 of 16 Archive leaderboard report
Cross-modal retrieval with noisy correspondence Flickr30K-Noisy RCL-SGR Text-to-image R@5 81.2 #16 of 16 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections