{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/instance-aware-image-and-sentence-matching","title":"Instance-aware Image and Sentence Matching with Selective Multimodal LSTM","arxiv_id":"1611.05588","date":"2016-11-17","proceeding":"CVPR 2017 7","authors":["Yan Huang","Wei Wang","Liang Wang"],"abstract":"Effective image and sentence matching depends on how to well measure their\nglobal visual-semantic similarity. Based on the observation that such a global\nsimilarity arises from a complex aggregation of multiple local similarities\nbetween pairwise instances of image (objects) and sentence (words), we propose\na selective multimodal Long Short-Term Memory network (sm-LSTM) for\ninstance-aware image and sentence matching. The sm-LSTM includes a multimodal\ncontext-modulated attention scheme at each timestep that can selectively attend\nto a pair of instances of image and sentence, by predicting pairwise\ninstance-aware saliency maps for image and sentence. For selected pairwise\ninstances, their representations are obtained based on the predicted saliency\nmaps, and then compared to measure their local similarity. By similarly\nmeasuring multiple local similarities within a few timesteps, the sm-LSTM\nsequentially aggregates them with hidden states to obtain a final matching\nscore as the desired global similarity. Extensive experiments show that our\nmodel can well match image and sentence with complex content, and achieve the\nstate-of-the-art results on two public benchmark datasets.","url_abs":"http://arxiv.org/abs/1611.05588v1","url_pdf":"http://arxiv.org/pdf/1611.05588v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"semantic-similarity","task_name":"Semantic Similarity"},{"task_slug":"semantic-textual-similarity","task_name":"Semantic Textual Similarity"},{"task_slug":"sentence","task_name":"Sentence"}],"methods":[{"method_slug":"memory-network","method_name":"Memory Network"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/image-retrieval-on-flickr30k-1k-test","task":"Image Retrieval","dataset":"Flickr30K 1K test","model":"SM-LSTM (VGG)","rank_in_archive_order":14,"of":18,"metrics":{"R@1":"30.2","R@10":"72.3"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1611.05588","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}