{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/multi-modal-factorized-bilinear-pooling-with","title":"Multi-modal Factorized Bilinear Pooling with Co-Attention Learning for Visual Question Answering","arxiv_id":"1708.01471","date":"2017-08-04","proceeding":"ICCV 2017 10","authors":["Zhou Yu","Jun Yu","Jianping Fan","DaCheng Tao"],"abstract":"Visual question answering (VQA) is challenging because it requires a\nsimultaneous understanding of both the visual content of images and the textual\ncontent of questions. The approaches used to represent the images and questions\nin a fine-grained manner and questions and to fuse these multi-modal features\nplay key roles in performance. Bilinear pooling based models have been shown to\noutperform traditional linear models for VQA, but their high-dimensional\nrepresentations and high computational complexity may seriously limit their\napplicability in practice. For multi-modal feature fusion, here we develop a\nMulti-modal Factorized Bilinear (MFB) pooling approach to efficiently and\neffectively combine multi-modal features, which results in superior performance\nfor VQA compared with other bilinear pooling approaches. For fine-grained image\nand question representation, we develop a co-attention mechanism using an\nend-to-end deep network architecture to jointly learn both the image and\nquestion attentions. Combining the proposed MFB approach with co-attention\nlearning in a new network architecture provides a unified model for VQA. Our\nexperimental results demonstrate that the single MFB with co-attention model\nachieves new state-of-the-art performance on the real-world VQA dataset. Code\navailable at https://github.com/yuzcccc/mfb.","url_abs":"http://arxiv.org/abs/1708.01471v1","url_pdf":"http://arxiv.org/pdf/1708.01471v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"multi-modal-factorized-bilinear-pooling-with","repo_url":"https://github.com/yuzcccc/mfb","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"caffe2","reach":{"status":"unanswered"}},{"paper_slug":"multi-modal-factorized-bilinear-pooling-with","repo_url":"https://github.com/apugoneappu/ask_me_anything","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"multi-modal-factorized-bilinear-pooling-with","repo_url":"https://github.com/apugoneappu/vqa_visualise","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"multi-modal-factorized-bilinear-pooling-with","repo_url":"https://github.com/straightAYiJun/vqa-attention-visualize-system","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"multi-modal-factorized-bilinear-pooling-with","repo_url":"https://github.com/vikrantmane7781/detectroon2","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"multi-modal-factorized-bilinear-pooling-with","repo_url":"https://github.com/yuzcccc/vqa-mfb","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"caffe2","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"visual-question-answering-1","task_name":"Visual Question Answering"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1708.01471","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}