{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/multimodal-residual-learning-for-visual-qa","title":"Multimodal Residual Learning for Visual QA","arxiv_id":"1606.01455","date":"2016-06-05","proceeding":"NeurIPS 2016 12","authors":["Jin-Hwa Kim","Sang-Woo Lee","Dong-Hyun Kwak","Min-Oh Heo","Jeonghee Kim","Jung-Woo Ha","Byoung-Tak Zhang"],"abstract":"Deep neural networks continue to advance the state-of-the-art of image\nrecognition tasks with various methods. However, applications of these methods\nto multimodality remain limited. We present Multimodal Residual Networks (MRN)\nfor the multimodal residual learning of visual question-answering, which\nextends the idea of the deep residual learning. Unlike the deep residual\nlearning, MRN effectively learns the joint representation from vision and\nlanguage information. The main idea is to use element-wise multiplication for\nthe joint residual mappings exploiting the residual learning of the attentional\nmodels in recent studies. Various alternative models introduced by\nmultimodality are explored based on our study. We achieve the state-of-the-art\nresults on the Visual QA dataset for both Open-Ended and Multiple-Choice tasks.\nMoreover, we introduce a novel method to visualize the attention effect of the\njoint representations for each learning block using back-propagation algorithm,\neven though the visual features are collapsed without spatial information.","url_abs":"http://arxiv.org/abs/1606.01455v2","url_pdf":"http://arxiv.org/pdf/1606.01455v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"multimodal-residual-learning-for-visual-qa","repo_url":"https://github.com/jnhwkim/nips-mrn-vqa","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"torch","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"multiple-choice","task_name":"Multiple-choice"},{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"visual-question-answering-1","task_name":"Visual Question Answering"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/visual-question-answering-on-coco-visual-1","task":"Visual Question Answering (VQA)","dataset":"COCO Visual Question Answering (VQA) real images 1.0 multiple choice","model":"MRN","rank_in_archive_order":6,"of":10,"metrics":{"Percentage correct":"66.3"},"uses_additional_data":false},{"leaderboard":"/sota/visual-question-answering-on-coco-visual-4","task":"Visual Question Answering (VQA)","dataset":"COCO Visual Question Answering (VQA) real images 1.0 open ended","model":"MRN + global features","rank_in_archive_order":7,"of":14,"metrics":{"Percentage correct":"61.8"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1606.01455","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}