{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/progressive-attention-memory-network-for","title":"Progressive Attention Memory Network for Movie Story Question Answering","arxiv_id":"1904.08607","date":"2019-04-18","proceeding":"CVPR 2019 6","authors":["Junyeong Kim","Minuk Ma","Kyung-Su Kim","Sungjin Kim","Chang D. Yoo"],"abstract":"This paper proposes the progressive attention memory network (PAMN) for movie\nstory question answering (QA). Movie story QA is challenging compared to VQA in\ntwo aspects: (1) pinpointing the temporal parts relevant to answer the question\nis difficult as the movies are typically longer than an hour, (2) it has both\nvideo and subtitle where different questions require different modality to\ninfer the answer. To overcome these challenges, PAMN involves three main\nfeatures: (1) progressive attention mechanism that utilizes cues from both\nquestion and answer to progressively prune out irrelevant temporal parts in\nmemory, (2) dynamic modality fusion that adaptively determines the contribution\nof each modality for answering the current question, and (3) belief correction\nanswering scheme that successively corrects the prediction score on each\ncandidate answer. Experiments on publicly available benchmark datasets, MovieQA\nand TVQA, demonstrate that each feature contributes to our movie story QA\narchitecture, PAMN, and improves performance to achieve the state-of-the-art\nresult. Qualitative analysis by visualizing the inference mechanism of PAMN is\nalso provided.","url_abs":"http://arxiv.org/abs/1904.08607v1","url_pdf":"http://arxiv.org/pdf/1904.08607v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"video-story-qa","task_name":"Video Story QA"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"}],"methods":[{"method_slug":"memory-network","method_name":"Memory Network"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/video-story-qa-on-movieqa","task":"Video Story QA","dataset":"MovieQA","model":"PAMN","rank_in_archive_order":2,"of":3,"metrics":{"Accuracy":"42.53"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1904.08607","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}