{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/moviechat-from-dense-token-to-sparse-memory","title":"MovieChat: From Dense Token to Sparse Memory for Long Video Understanding","arxiv_id":"2307.16449","date":"2023-07-31","proceeding":"CVPR 2024 1","authors":["Enxin Song","Wenhao Chai","Guanhong Wang","Yucheng Zhang","Haoyang Zhou","Feiyang Wu","Haozhe Chi","Xun Guo","Tian Ye","Yanting Zhang","Yan Lu","Jenq-Neng Hwang","Gaoang Wang"],"abstract":"Recently, integrating video foundation models and large language models to build a video understanding system can overcome the limitations of specific pre-defined vision tasks. Yet, existing systems can only handle videos with very few frames. For long videos, the computation complexity, memory cost, and long-term temporal connection impose additional challenges. Taking advantage of the Atkinson-Shiffrin memory model, with tokens in Transformers being employed as the carriers of memory in combination with our specially designed memory mechanism, we propose the MovieChat to overcome these challenges. MovieChat achieves state-of-the-art performance in long video understanding, along with the released MovieChat-1K benchmark with 1K long video and 14K manual annotations for validation of the effectiveness of our method.","url_abs":"https://arxiv.org/abs/2307.16449v4","url_pdf":"https://arxiv.org/pdf/2307.16449v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"moviechat-from-dense-token-to-sparse-memory","repo_url":"https://github.com/rese1f/MovieChat","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"BSD-3-Clause"}}],"tasks":[{"task_slug":"multiple-choice","task_name":"Multiple-choice"},{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"video-question-answering","task_name":"Video Question Answering"},{"task_slug":"video-understanding","task_name":"Video Understanding"},{"task_slug":"video-based-generative-performance","task_name":"Video-based Generative Performance Benchmarking"},{"task_slug":"video-based-generative-performance-5","task_name":"Video-based Generative Performance Benchmarking (Consistency)"},{"task_slug":"video-based-generative-performance-3","task_name":"Video-based Generative Performance Benchmarking (Contextual Understanding)"},{"task_slug":"video-based-generative-performance-1","task_name":"Video-based Generative Performance Benchmarking (Correctness of Information)"},{"task_slug":"video-based-generative-performance-2","task_name":"Video-based Generative Performance Benchmarking (Detail Orientation))"},{"task_slug":"video-based-generative-performance-4","task_name":"Video-based Generative Performance Benchmarking (Temporal Understanding)"},{"task_slug":"zeroshot-video-question-answer","task_name":"Zero-Shot Video Question Answer"},{"task_slug":null,"task_name":"zero-shot long video breakpoint-mode question answering"},{"task_slug":null,"task_name":"zero-shot long video breakpoint-model question answering"},{"task_slug":null,"task_name":"zero-shot long video global-mode question answering"},{"task_slug":null,"task_name":"zero-shot long video global-model question answering"},{"task_slug":"zero-shot-long-video-question-answering","task_name":"zero-shot long video question answering"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/question-answering-on-next-qa-open-ended","task":"Question Answering","dataset":"NExT-QA (Open-ended VideoQA)","model":"MovieChat","rank_in_archive_order":6,"of":6,"metrics":{"Accuracy":"49.9","Confidence Score":"2.7"},"uses_additional_data":false},{"leaderboard":"/sota/video-question-answering-on-activitynet-qa","task":"Video Question Answering","dataset":"ActivityNet-QA","model":"MovieChat","rank_in_archive_order":15,"of":36,"metrics":{"Accuracy":"45.7","Confidence score":"3.1"},"uses_additional_data":false},{"leaderboard":"/sota/video-question-answering-on-ovbench","task":"Video Question Answering","dataset":"OVBench","model":"MovieChat (7B)","rank_in_archive_order":13,"of":16,"metrics":{"AVG":"30.9"},"uses_additional_data":false},{"leaderboard":"/sota/video-based-generative-performance-2","task":"Video-based Generative Performance Benchmarking (Consistency)","dataset":"VideoInstruct","model":"MovieChat","rank_in_archive_order":13,"of":18,"metrics":{"gpt-score":"2.42"},"uses_additional_data":false},{"leaderboard":"/sota/video-based-generative-performance-3","task":"Video-based Generative Performance Benchmarking (Contextual Understanding)","dataset":"VideoInstruct","model":"MovieChat","rank_in_archive_order":13,"of":18,"metrics":{"gpt-score":"3.01"},"uses_additional_data":false},{"leaderboard":"/sota/video-based-generative-performance-1","task":"Video-based Generative Performance Benchmarking (Correctness of Information)","dataset":"VideoInstruct","model":"MovieChat","rank_in_archive_order":12,"of":18,"metrics":{"gpt-score":"2.76"},"uses_additional_data":false},{"leaderboard":"/sota/video-based-generative-performance-4","task":"Video-based Generative Performance Benchmarking (Detail Orientation))","dataset":"VideoInstruct","model":"MovieChat","rank_in_archive_order":9,"of":18,"metrics":{"gpt-score":"2.93"},"uses_additional_data":false},{"leaderboard":"/sota/video-based-generative-performance-5","task":"Video-based Generative Performance Benchmarking (Temporal Understanding)","dataset":"VideoInstruct","model":"MovieChat","rank_in_archive_order":13,"of":18,"metrics":{"gpt-score":"2.24"},"uses_additional_data":false},{"leaderboard":"/sota/zeroshot-video-question-answer-on-activitynet","task":"Zero-Shot Video Question Answer","dataset":"ActivityNet-QA","model":"MovieChat","rank_in_archive_order":21,"of":28,"metrics":{"Accuracy":"45.7","Confidence Score":"3.1"},"uses_additional_data":false},{"leaderboard":"/sota/zeroshot-video-question-answer-on-msrvtt-qa","task":"Zero-Shot Video Question Answer","dataset":"MSRVTT-QA","model":"MovieChat","rank_in_archive_order":24,"of":30,"metrics":{"Accuracy":"52.7","Confidence Score":"2.6"},"uses_additional_data":false},{"leaderboard":"/sota/zeroshot-video-question-answer-on-msvd-qa","task":"Zero-Shot Video Question Answer","dataset":"MSVD-QA","model":"MovieChat","rank_in_archive_order":11,"of":28,"metrics":{"Accuracy":"75.2","Confidence Score":"2.9"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2307.16449","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}