{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/improved-fusion-of-visual-and-language","title":"Improved Fusion of Visual and Language Representations by Dense Symmetric Co-Attention for Visual Question Answering","arxiv_id":"1804.00775","date":"2018-04-03","proceeding":"CVPR 2018 6","authors":["Duy-Kien Nguyen","Takayuki Okatani"],"abstract":"A key solution to visual question answering (VQA) exists in how to fuse\nvisual and language features extracted from an input image and question. We\nshow that an attention mechanism that enables dense, bi-directional\ninteractions between the two modalities contributes to boost accuracy of\nprediction of answers. Specifically, we present a simple architecture that is\nfully symmetric between visual and language representations, in which each\nquestion word attends on image regions and each image region attends on\nquestion words. It can be stacked to form a hierarchy for multi-step\ninteractions between an image-question pair. We show through experiments that\nthe proposed architecture achieves a new state-of-the-art on VQA and VQA 2.0\ndespite its small size. We also present qualitative evaluation, demonstrating\nhow the proposed attention mechanism can generate reasonable attention maps on\nimages and questions, which leads to the correct answer prediction.","url_abs":"http://arxiv.org/abs/1804.00775v2","url_pdf":"http://arxiv.org/pdf/1804.00775v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"improved-fusion-of-visual-and-language","repo_url":"https://github.com/cvlab-tohoku/Dense-CoAttention-Network","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"visual-question-answering-1","task_name":"Visual Question Answering"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1804.00775","atlas_url":"https://app.syntology.ai/?focus=1804.00775","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}