{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-two-branch-neural-networks-for-image","title":"Learning Two-Branch Neural Networks for Image-Text Matching Tasks","arxiv_id":"1704.03470","date":"2017-04-11","proceeding":null,"authors":["Liwei Wang","Yin Li","Jing Huang","Svetlana Lazebnik"],"abstract":"Image-language matching tasks have recently attracted a lot of attention in\nthe computer vision field. These tasks include image-sentence matching, i.e.,\ngiven an image query, retrieving relevant sentences and vice versa, and\nregion-phrase matching or visual grounding, i.e., matching a phrase to relevant\nregions. This paper investigates two-branch neural networks for learning the\nsimilarity between these two data modalities. We propose two network structures\nthat produce different output representations. The first one, referred to as an\nembedding network, learns an explicit shared latent embedding space with a\nmaximum-margin ranking loss and novel neighborhood constraints. Compared to\nstandard triplet sampling, we perform improved neighborhood sampling that takes\nneighborhood information into consideration while constructing mini-batches.\nThe second network structure, referred to as a similarity network, fuses the\ntwo branches via element-wise product and is trained with regression loss to\ndirectly predict a similarity score. Extensive experiments show that our\nnetworks achieve high accuracies for phrase localization on the Flickr30K\nEntities dataset and for bi-directional image-sentence retrieval on Flickr30K\nand MSCOCO datasets.","url_abs":"http://arxiv.org/abs/1704.03470v4","url_pdf":"http://arxiv.org/pdf/1704.03470v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-two-branch-neural-networks-for-image","repo_url":"https://github.com/BryanPlummer/cite","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"image-text-matching","task_name":"Image-text matching"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"sentence","task_name":"Sentence"},{"task_slug":"sentence-retrieval","task_name":"Sentence Retrieval"},{"task_slug":"text-matching","task_name":"Text Matching"},{"task_slug":null,"task_name":"Triplet"},{"task_slug":"visual-grounding","task_name":"Visual Grounding"},{"task_slug":"two","task_name":"Vocal Bursts Valence Prediction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1704.03470","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}