{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/deep-cross-modal-projection-learning-for","title":"Deep Cross-Modal Projection Learning for Image-Text Matching","arxiv_id":null,"date":"2018-09-01","proceeding":"ECCV 2018 9","authors":["Ying Zhang","Huchuan Lu "],"abstract":"The key point of image-text matching is how to accurately measure the similarity between visual and textual inputs. Despite the great progress of associating the deep cross-modal embeddings with the bi-directional ranking loss, developing the strategies for mining useful triplets and selecting appropriate margins remains a challenge in real applications. In this paper, we propose a cross-modal projection matching (CMPM) loss and a cross-modal projection classification (CMPC) loss for learning discriminative image-text embeddings. The CMPM loss minimizes the KL divergence between the projection compatibility distributions and the normalized matching distributions defined with all the positive and negative samples in a mini-batch. The CMPC loss attempts to categorize the vector projection of representations from one modality onto another with the improved norm-softmax loss, for further enhancing the feature compactness of each class. Extensive analysis and experiments on multiple datasets demonstrate the superiority of the proposed approach.","url_abs":"http://openaccess.thecvf.com/content_ECCV_2018/html/Ying_Zhang_Deep_Cross-Modal_Projection_ECCV_2018_paper.html","url_pdf":"http://openaccess.thecvf.com/content_ECCV_2018/papers/Ying_Zhang_Deep_Cross-Modal_Projection_ECCV_2018_paper.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"deep-cross-modal-projection-learning-for","repo_url":"https://github.com/YingZhangDUT/Cross-Modal-Projection-Learning","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"cross-modal-retrieval","task_name":"Cross-Modal Retrieval"},{"task_slug":"image-text-matching","task_name":"Image-text matching"},{"task_slug":"text-matching","task_name":"Text Matching"},{"task_slug":"nlp-based-person-retrival","task_name":"Text based Person Retrieval"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/cross-modal-retrieval-on-flickr30k","task":"Cross-Modal Retrieval","dataset":"Flickr30k","model":"CMPL\n  (ResNet)","rank_in_archive_order":25,"of":27,"metrics":{"Image-to-text R@1":"49.6","Image-to-text R@10":"86.1","Image-to-text R@5":"76.8","Text-to-image R@1":"37.3","Text-to-image R@10":"75.5","Text-to-image R@5":"65.7"},"uses_additional_data":false},{"leaderboard":"/sota/nlp-based-person-retrival-on-cuhk-pedes","task":"Text based Person Retrieval","dataset":"CUHK-PEDES","model":"CMPM+CMPC","rank_in_archive_order":18,"of":21,"metrics":{"R@1":"49.37","R@10":"79.27","R@5":"-"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}