{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/mmtf-multi-modal-temporal-fusion-for","title":"MMTF: Multi-Modal Temporal Fusion for Commonsense Video Question Answering","arxiv_id":null,"date":"2023-10-06","proceeding":"Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) Workshops, 2023 2023 10","authors":["Mobeen Ahmad","Geonwoo Park","Dongchan Park","Sanguk Park"],"abstract":"Video question answering is a challenging task that requires understanding the video and question in the same context. This becomes even harder when the questions involve reasoning, such as predicting future events or explaining counterfactual events, because they need knowledge not explicitly shown. Existing methods use coarse-grained fusion of video and language features, ignoring temporal information. To address this, we propose a novel vision-text fusion module that learns the temporal context of the video and question. Our module expands question tokens along the video's temporal axis and fuses them with video features to generate new representations with local and global context. We evaluated our method on four VideoQA datasets, including MSVD-QA, NExT-QA, Causal-VidQA, and AGQA-2.0.","url_abs":"https://openaccess.thecvf.com/content/ICCV2023W/VLAR/html/Ahmad_MMTF_Multi-Modal_Temporal_Fusion_for_Commonsense_Video_Question_Answering_ICCVW_2023_paper.html","url_pdf":"https://openaccess.thecvf.com/content/ICCV2023W/VLAR/papers/Ahmad_MMTF_Multi-Modal_Temporal_Fusion_for_Commonsense_Video_Question_Answering_ICCVW_2023_paper.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"video-question-answering","task_name":"Video Question Answering"},{"task_slug":null,"task_name":"counterfactual"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/video-question-answering-on-agqa-2-0-balanced","task":"Video Question Answering","dataset":"AGQA 2.0 balanced","model":"MMTF","rank_in_archive_order":8,"of":8,"metrics":{"Average Accuracy":"44.36"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}