{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/simnet-stepwise-image-topic-merging-network","title":"simNet: Stepwise Image-Topic Merging Network for Generating Detailed and Comprehensive Image Captions","arxiv_id":"1808.08732","date":"2018-08-27","proceeding":"EMNLP 2018 10","authors":["Fenglin Liu","Xuancheng Ren","Yuanxin Liu","Houfeng Wang","Xu sun"],"abstract":"The encode-decoder framework has shown recent success in image captioning.\nVisual attention, which is good at detailedness, and semantic attention, which\nis good at comprehensiveness, have been separately proposed to ground the\ncaption on the image. In this paper, we propose the Stepwise Image-Topic\nMerging Network (simNet) that makes use of the two kinds of attention at the\nsame time. At each time step when generating the caption, the decoder\nadaptively merges the attentive information in the extracted topics and the\nimage according to the generated context, so that the visual information and\nthe semantic information can be effectively combined. The proposed approach is\nevaluated on two benchmark datasets and reaches the state-of-the-art\nperformances.(The code is available at https://github.com/lancopku/simNet)","url_abs":"http://arxiv.org/abs/1808.08732v1","url_pdf":"http://arxiv.org/pdf/1808.08732v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"simnet-stepwise-image-topic-merging-network","repo_url":"https://github.com/lancopku/simNet","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"image-captioning","task_name":"Image Captioning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1808.08732","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}