Papers › MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

31 Jul 2023CVPR 2024 1arXiv:2307.16449archive 2025-07-28

Enxin Song, Wenhao Chai, Guanhong Wang, Yucheng Zhang, Haoyang Zhou, Feiyang Wu, Haozhe Chi, Xun Guo, Tian Ye, Yanting Zhang, Yan Lu, Jenq-Neng Hwang, Gaoang Wang

Recently, integrating video foundation models and large language models to build a video understanding system can overcome the limitations of specific pre-defined vision tasks. Yet, existing systems can only handle videos with very few frames. For long videos, the computation complexity, memory cost, and long-term temporal connection impose additional challenges. Taking advantage of the Atkinson-Shiffrin memory model, with tokens in Transformers being employed as the carriers of memory in combination with our specially designed memory mechanism, we propose the MovieChat to overcome these challenges. MovieChat achieves state-of-the-art performance in long video understanding, along with the released MovieChat-1K benchmark with 1K long video and 14K manual annotations for validation of the effectiveness of our method.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

rese1f/MovieChat officialmentioned on GitHubpytorchBSD-3-Clause report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Multiple-choiceQuestion AnsweringVideo Question AnsweringVideo UnderstandingVideo-based Generative Performance BenchmarkingVideo-based Generative Performance Benchmarking (Consistency)Video-based Generative Performance Benchmarking (Contextual Understanding)Video-based Generative Performance Benchmarking (Correctness of Information)Video-based Generative Performance Benchmarking (Detail Orientation))Video-based Generative Performance Benchmarking (Temporal Understanding)Zero-Shot Video Question Answerzero-shot long video question answering

4 archive task tags without a task page not shown.

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Question Answering NExT-QA (Open-ended VideoQA) MovieChat Accuracy 49.9 #6 of 6 Archive leaderboard report
Question Answering NExT-QA (Open-ended VideoQA) MovieChat Confidence Score 2.7 #6 of 6 Archive leaderboard report
Video Question Answering ActivityNet-QA MovieChat Accuracy 45.7 #15 of 36 Archive leaderboard report
Video Question Answering ActivityNet-QA MovieChat Confidence score 3.1 #15 of 36 Archive leaderboard report
Video Question Answering OVBench MovieChat (7B) AVG 30.9 #13 of 16 Archive leaderboard report
Video-based Generative Performance Benchmarking (Consistency) VideoInstruct MovieChat gpt-score 2.42 #13 of 18 Archive leaderboard report
Video-based Generative Performance Benchmarking (Contextual Understanding) VideoInstruct MovieChat gpt-score 3.01 #13 of 18 Archive leaderboard report
Video-based Generative Performance Benchmarking (Correctness of Information) VideoInstruct MovieChat gpt-score 2.76 #12 of 18 Archive leaderboard report
Video-based Generative Performance Benchmarking (Detail Orientation)) VideoInstruct MovieChat gpt-score 2.93 #9 of 18 Archive leaderboard report
Video-based Generative Performance Benchmarking (Temporal Understanding) VideoInstruct MovieChat gpt-score 2.24 #13 of 18 Archive leaderboard report
Zero-Shot Video Question Answer ActivityNet-QA MovieChat Accuracy 45.7 #21 of 28 Archive leaderboard report
Zero-Shot Video Question Answer ActivityNet-QA MovieChat Confidence Score 3.1 #21 of 28 Archive leaderboard report
Zero-Shot Video Question Answer MSRVTT-QA MovieChat Accuracy 52.7 #24 of 30 Archive leaderboard report
Zero-Shot Video Question Answer MSRVTT-QA MovieChat Confidence Score 2.6 #24 of 30 Archive leaderboard report
Zero-Shot Video Question Answer MSVD-QA MovieChat Accuracy 75.2 #11 of 28 Archive leaderboard report
Zero-Shot Video Question Answer MSVD-QA MovieChat Confidence Score 2.9 #11 of 28 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections