Papers › Noisy Correspondence Learning with Self-Reinforcing Errors Mitigation

Noisy Correspondence Learning with Self-Reinforcing Errors Mitigation

27 Dec 2023arXiv:2312.16478archive 2025-07-28

Zhuohang Dang, Minnan Luo, Chengyou Jia, Guang Dai, Xiaojun Chang, Jingdong Wang

Cross-modal retrieval relies on well-matched large-scale datasets that are laborious in practice. Recently, to alleviate expensive data collection, co-occurring pairs from the Internet are automatically harvested for training. However, it inevitably includes mismatched pairs, \ie, noisy correspondences, undermining supervision reliability and degrading performance. Current methods leverage deep neural networks' memorization effect to address noisy correspondences, which overconfidently focus on \emph{similarity-guided training with hard negatives} and suffer from self-reinforcing errors. In light of above, we introduce a novel noisy correspondence learning framework, namely \textbf{S}elf-\textbf{R}einforcing \textbf{E}rrors \textbf{M}itigation (SREM). Specifically, by viewing sample matching as classification tasks within the batch, we generate classification logits for the given sample. Instead of a single similarity score, we refine sample filtration through energy uncertainty and estimate model's sensitivity of selected clean samples using swapped classification entropy, in view of the overall prediction distribution. Additionally, we propose cross-modal biased complementary learning to leverage negative matches overlooked in hard-negative training, further improving model optimization stability and curbing self-reinforcing errors. Extensive experiments on challenging benchmarks affirm the efficacy and efficiency of SREM.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Cross-Modal RetrievalCross-modal retrieval with noisy correspondenceMemorizationModel OptimizationRetrieval

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Cross-modal retrieval with noisy correspondence CC152K SREM Image-to-text R@1 40.9 #8 of 15 Archive leaderboard report
Cross-modal retrieval with noisy correspondence CC152K SREM Image-to-text R@10 77.1 #8 of 15 Archive leaderboard report
Cross-modal retrieval with noisy correspondence CC152K SREM Image-to-text R@5 67.5 #8 of 15 Archive leaderboard report
Cross-modal retrieval with noisy correspondence CC152K SREM R-Sum 372.2 #8 of 15 Archive leaderboard report
Cross-modal retrieval with noisy correspondence CC152K SREM Text-to-image R@1 41.5 #8 of 15 Archive leaderboard report
Cross-modal retrieval with noisy correspondence CC152K SREM Text-to-image R@10 77.0 #8 of 15 Archive leaderboard report
Cross-modal retrieval with noisy correspondence CC152K SREM Text-to-image R@5 68.2 #8 of 15 Archive leaderboard report
Cross-modal retrieval with noisy correspondence COCO-Noisy SREM Image-to-text R@1 78.5 #10 of 17 Archive leaderboard report
Cross-modal retrieval with noisy correspondence COCO-Noisy SREM Image-to-text R@10 98.8 #10 of 17 Archive leaderboard report
Cross-modal retrieval with noisy correspondence COCO-Noisy SREM Image-to-text R@5 96.8 #10 of 17 Archive leaderboard report
Cross-modal retrieval with noisy correspondence COCO-Noisy SREM R-Sum 524.1 #10 of 17 Archive leaderboard report
Cross-modal retrieval with noisy correspondence COCO-Noisy SREM Text-to-image R@1 63.8 #10 of 17 Archive leaderboard report
Cross-modal retrieval with noisy correspondence COCO-Noisy SREM Text-to-image R@10 95.8 #10 of 17 Archive leaderboard report
Cross-modal retrieval with noisy correspondence COCO-Noisy SREM Text-to-image R@5 90.4 #10 of 17 Archive leaderboard report
Cross-modal retrieval with noisy correspondence Flickr30K-Noisy SREM Image-to-text R@1 79.5 #3 of 16 Archive leaderboard report
Cross-modal retrieval with noisy correspondence Flickr30K-Noisy SREM Image-to-text R@10 97.9 #3 of 16 Archive leaderboard report
Cross-modal retrieval with noisy correspondence Flickr30K-Noisy SREM Image-to-text R@5 94.2 #3 of 16 Archive leaderboard report
Cross-modal retrieval with noisy correspondence Flickr30K-Noisy SREM R-Sum 507.8 #3 of 16 Archive leaderboard report
Cross-modal retrieval with noisy correspondence Flickr30K-Noisy SREM Text-to-image R@1 61.2 #3 of 16 Archive leaderboard report
Cross-modal retrieval with noisy correspondence Flickr30K-Noisy SREM Text-to-image R@10 90.2 #3 of 16 Archive leaderboard report
Cross-modal retrieval with noisy correspondence Flickr30K-Noisy SREM Text-to-image R@5 84.8 #3 of 16 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Focus

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections