{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/gated-hierarchical-attention-for-image","title":"Gated Hierarchical Attention for Image Captioning","arxiv_id":"1810.12535","date":"2018-10-30","proceeding":null,"authors":["Qingzhong Wang","Antoni B. Chan"],"abstract":"Attention modules connecting encoder and decoders have been widely applied in\nthe field of object recognition, image captioning, visual question answering\nand neural machine translation, and significantly improves the performance. In\nthis paper, we propose a bottom-up gated hierarchical attention (GHA) mechanism\nfor image captioning. Our proposed model employs a CNN as the decoder which is\nable to learn different concepts at different layers, and apparently, different\nconcepts correspond to different areas of an image. Therefore, we develop the\nGHA in which low-level concepts are merged into high-level concepts and\nsimultaneously low-level attended features pass to the top to make predictions.\nOur GHA significantly improves the performance of the model that only applies\none level attention, for example, the CIDEr score increases from 0.923 to\n0.999, which is comparable to the state-of-the-art models that employ\nattributes boosting and reinforcement learning (RL). We also conduct extensive\nexperiments to analyze the CNN decoder and our proposed GHA, and we find that\ndeeper decoders cannot obtain better performance, and when the convolutional\ndecoder becomes deeper the model is likely to collapse during training.","url_abs":"http://arxiv.org/abs/1810.12535v2","url_pdf":"http://arxiv.org/pdf/1810.12535v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"gated-hierarchical-attention-for-image","repo_url":"https://github.com/qingzwang/GHA-ImageCaptioning","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"image-captioning","task_name":"Image Captioning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"translation","task_name":"Translation"},{"task_slug":"visual-question-answering-1","task_name":"Visual Question Answering"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}