{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/deep-evidential-learning-with-noisy","title":"Deep Evidential Learning with Noisy Correspondence for Cross-Modal Retrieval","arxiv_id":null,"date":"2022-10-10","proceeding":"ACM International Conference on Multimedia 2022 10","authors":["Yang Qin","Dezhong Peng","Xi Peng","Xu Wang","Peng Hu"],"abstract":"Cross-modal retrieval has been a compelling topic in the multimodal community. Recently, to mitigate the high cost of data collection, the co-occurred pairs (e.g., image and text) could be collected from the Internet as a large-scaled cross-modal dataset, e.g., Conceptual Captions. However, it will unavoidably introduce noise (i.e., mismatched pairs) into training data, dubbed noisy correspondence. Unquestionably, such noise will make supervision information unreliable/uncertain and remarkably degrade the performance. Besides, most existing methods focus training on hard negatives, which will amplify the unreliability of noise. To address the issues, we propose a generalized Deep Evidential Cross-modal Learning framework (DECL), which integrates a novel Cross-modal Evidential Learning paradigm (CEL) and a Robust Dynamic Hinge loss (RDH) with positive and negative learning. CEL could capture and learn the uncertainty brought by noise to improve the robustness and reliability of cross-modal retrieval. Specifically, the bidirectional evidence based on cross-modal similarity is first modeled and parameterized into the Dirichlet distribution, which not only provides accurate uncertainty estimation but also imparts resilience to perturbations against noisy correspondence. To address the amplification problem, RDH smoothly increases the hardness of negatives focused on, thus embracing higher robustness against high noise. Extensive experiments are conducted on three image-text benchmark datasets, i.e., Flickr30K, MS-COCO, and Conceptual Captions, to verify the effectiveness and efficiency of the proposed method. The code is available at \\urlhttps://github.com/QinYang79/DECL.","url_abs":"https://dl.acm.org/doi/abs/10.1145/3503161.3547922","url_pdf":"https://dl.acm.org/doi/pdf/10.1145/3503161.3547922","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"deep-evidential-learning-with-noisy","repo_url":"https://github.com/qinyang79/decl","is_official":0,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"cross-modal-retrieval","task_name":"Cross-Modal Retrieval"},{"task_slug":"cross-modal-retrieval-with-noisy","task_name":"Cross-modal retrieval with noisy correspondence"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"text-based-person-retrieval-with-noisy","task_name":"Text-based Person Retrieval with Noisy Correspondence"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/cross-modal-retrieval-with-noisy-1","task":"Cross-modal retrieval with noisy correspondence","dataset":"CC152K","model":"DECL-SGRAF","rank_in_archive_order":13,"of":15,"metrics":{"Image-to-text R@1":"39.0","Image-to-text R@10":"75.5","Image-to-text R@5":"66.1","R-Sum":"364.3","Text-to-image R@1":"40.7","Text-to-image R@10":"76.7","Text-to-image R@5":"66.3"},"uses_additional_data":false},{"leaderboard":"/sota/cross-modal-retrieval-with-noisy-3","task":"Cross-modal retrieval with noisy correspondence","dataset":"COCO-Noisy","model":"DECL-SGARF","rank_in_archive_order":16,"of":17,"metrics":{"Image-to-text R@1":"77.5","Image-to-text R@10":"98.4","Image-to-text R@5":"95.9","R-Sum":"518.2","Text-to-image R@1":"61.7","Text-to-image R@10":"95.4","Text-to-image R@5":"89.3"},"uses_additional_data":false},{"leaderboard":"/sota/cross-modal-retrieval-with-noisy-2","task":"Cross-modal retrieval with noisy correspondence","dataset":"Flickr30K-Noisy","model":"DECL-SGRAF","rank_in_archive_order":15,"of":16,"metrics":{"Image-to-text R@1":"77.5","Image-to-text R@10":"97.0","Image-to-text R@5":"93.8","R-Sum":"494.7","Text-to-image R@1":"56.1","Text-to-image R@10":"88.5","Text-to-image R@5":"81.8"},"uses_additional_data":false},{"leaderboard":"/sota/text-based-person-retrieval-with-noisy","task":"Text-based Person Retrieval with Noisy Correspondence","dataset":"CUHK-PEDES","model":"DECL","rank_in_archive_order":2,"of":6,"metrics":{"Rank 10":"91.93","Rank-1":"70.29","Rank-5":"87.04","mAP":"62.84","mINP":"46.54"},"uses_additional_data":true},{"leaderboard":"/sota/text-based-person-retrieval-with-noisy-1","task":"Text-based Person Retrieval with Noisy Correspondence","dataset":"ICFG-PEDES","model":"DECL","rank_in_archive_order":2,"of":6,"metrics":{"Rank 1":"61.95","Rank-10":"83.88","Rank-5":"78.36","mAP":"36.08","mINP":"6.25"},"uses_additional_data":true},{"leaderboard":"/sota/text-based-person-retrieval-with-noisy-2","task":"Text-based Person Retrieval with Noisy Correspondence","dataset":"RSTPReid","model":"DECL","rank_in_archive_order":2,"of":6,"metrics":{"Rank 1":"61.75","Rank 10":"86.90","Rank 5":"80.70","mAP":"47.70","mINP":"26.07"},"uses_additional_data":true}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}