{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-semantic-concepts-and-order-for","title":"Learning Semantic Concepts and Order for Image and Sentence Matching","arxiv_id":"1712.02036","date":"2017-12-06","proceeding":"CVPR 2018 6","authors":["Yan Huang","Qi Wu","Liang Wang"],"abstract":"Image and sentence matching has made great progress recently, but it remains\nchallenging due to the large visual-semantic discrepancy. This mainly arises\nfrom that the representation of pixel-level image usually lacks of high-level\nsemantic information as in its matched sentence. In this work, we propose a\nsemantic-enhanced image and sentence matching model, which can improve the\nimage representation by learning semantic concepts and then organizing them in\na correct semantic order. Given an image, we first use a multi-regional\nmulti-label CNN to predict its semantic concepts, including objects,\nproperties, actions, etc. Then, considering that different orders of semantic\nconcepts lead to diverse semantic meanings, we use a context-gated sentence\ngeneration scheme for semantic order learning. It simultaneously uses the image\nglobal context containing concept relations as reference and the groundtruth\nsemantic order in the matched sentence as supervision. After obtaining the\nimproved image representation, we learn the sentence representation with a\nconventional LSTM, and then jointly perform image and sentence matching and\nsentence generation for model learning. Extensive experiments demonstrate the\neffectiveness of our learned semantic concepts and order, by achieving the\nstate-of-the-art results on two public benchmark datasets.","url_abs":"http://arxiv.org/abs/1712.02036v1","url_pdf":"http://arxiv.org/pdf/1712.02036v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"cross-modal-retrieval","task_name":"Cross-Modal Retrieval"},{"task_slug":"sentence","task_name":"Sentence"}],"methods":[{"method_slug":"1x1-convolution","method_name":"1x1 Convolution"},{"method_slug":"average-pooling","method_name":"Average Pooling"},{"method_slug":"batch-normalization","method_name":"Batch Normalization"},{"method_slug":"bottleneck-residual-block","method_name":"Bottleneck Residual Block"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"global-average-pooling","method_name":"Global Average Pooling"},{"method_slug":"kaiming-initialization","method_name":"Kaiming Initialization"},{"method_slug":"lstm","method_name":"LSTM"},{"method_slug":"max-pooling","method_name":"Max Pooling"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"residual-block","method_name":"Residual Block"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/cross-modal-retrieval-on-coco-2014","task":"Cross-Modal Retrieval","dataset":"COCO 2014","model":"SCO (ResNet)","rank_in_archive_order":33,"of":36,"metrics":{"Image-to-text R@1":"42.8","Image-to-text R@10":"83.0","Image-to-text R@5":"72.3","Text-to-image R@1":"33.1","Text-to-image R@10":"75.5","Text-to-image R@5":"62.9"},"uses_additional_data":false},{"leaderboard":"/sota/cross-modal-retrieval-on-flickr30k","task":"Cross-Modal Retrieval","dataset":"Flickr30k","model":"SCO\n  (ResNet)","rank_in_archive_order":23,"of":27,"metrics":{"Image-to-text R@1":"55.5","Image-to-text R@10":"89.3","Image-to-text R@5":"82.0","Text-to-image R@1":"41.1","Text-to-image R@10":"80.1","Text-to-image R@5":"70.5"},"uses_additional_data":false},{"leaderboard":"/sota/image-retrieval-on-flickr30k-1k-test","task":"Image Retrieval","dataset":"Flickr30K 1K test","model":"SCO","rank_in_archive_order":11,"of":18,"metrics":{"R@1":"41.1","R@10":"80.1","R@5":"70.5"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1712.02036","atlas_url":"https://app.syntology.ai/?focus=1712.02036","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}