{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/ma-lmm-memory-augmented-large-multimodal","title":"MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding","arxiv_id":"2404.05726","date":"2024-04-08","proceeding":"CVPR 2024 1","authors":["Bo He","Hengduo Li","Young Kyun Jang","Menglin Jia","Xuefei Cao","Ashish Shah","Abhinav Shrivastava","Ser-Nam Lim"],"abstract":"With the success of large language models (LLMs), integrating the vision model into LLMs to build vision-language foundation models has gained much more interest recently. However, existing LLM-based large multimodal models (e.g., Video-LLaMA, VideoChat) can only take in a limited number of frames for short video understanding. In this study, we mainly focus on designing an efficient and effective model for long-term video understanding. Instead of trying to process more frames simultaneously like most existing work, we propose to process videos in an online manner and store past video information in a memory bank. This allows our model to reference historical video content for long-term analysis without exceeding LLMs' context length constraints or GPU memory limits. Our memory bank can be seamlessly integrated into current multimodal LLMs in an off-the-shelf manner. We conduct extensive experiments on various video understanding tasks, such as long-video understanding, video question answering, and video captioning, and our model can achieve state-of-the-art performances across multiple datasets. Code available at https://boheumd.github.io/MA-LMM/.","url_abs":"https://arxiv.org/abs/2404.05726v2","url_pdf":"https://arxiv.org/pdf/2404.05726v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"ma-lmm-memory-augmented-large-multimodal","repo_url":"https://github.com/boheumd/MA-LMM","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":null,"task_name":"GPU"},{"task_slug":"multiple-choice","task_name":"Multiple-choice"},{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"temporal-relation-extraction","task_name":"Temporal Relation Extraction"},{"task_slug":"video-captioning","task_name":"Video Captioning"},{"task_slug":"video-classification","task_name":"Video Classification"},{"task_slug":"video-question-answering","task_name":"Video Question Answering"},{"task_slug":"video-understanding","task_name":"Video Understanding"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"}],"methods":[{"method_slug":"focus","method_name":"Focus"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/temporal-relation-extraction-on-vinoground","task":"Temporal Relation Extraction","dataset":"Vinoground","model":"MA-LMM-Vicuna-7B","rank_in_archive_order":17,"of":24,"metrics":{"Group Score":"6.8","Text Score":"23.8","Video Score":"25.6"},"uses_additional_data":false},{"leaderboard":"/sota/video-captioning-on-youcook2","task":"Video Captioning","dataset":"YouCook2","model":"MA-LMM","rank_in_archive_order":14,"of":14,"metrics":{"CIDEr":"1.31","METEOR":"17.6"},"uses_additional_data":false},{"leaderboard":"/sota/video-classification-on-breakfast","task":"Video Classification","dataset":"Breakfast","model":"MA-LMM","rank_in_archive_order":2,"of":9,"metrics":{"Accuracy (%)":"93.0"},"uses_additional_data":false},{"leaderboard":"/sota/video-classification-on-coin-1","task":"Video Classification","dataset":"COIN","model":"MA-LMM","rank_in_archive_order":2,"of":7,"metrics":{"Accuracy (%)":"93.2"},"uses_additional_data":false},{"leaderboard":"/sota/video-question-answering-on-activitynet-qa","task":"Video Question Answering","dataset":"ActivityNet-QA","model":"MA-LMM","rank_in_archive_order":7,"of":36,"metrics":{"Accuracy":"49.8"},"uses_additional_data":false},{"leaderboard":"/sota/video-question-answering-on-msrvtt-qa","task":"Video Question Answering","dataset":"MSRVTT-QA","model":"MA-LMM","rank_in_archive_order":5,"of":14,"metrics":{"Accuracy":"48.5"},"uses_additional_data":false},{"leaderboard":"/sota/visual-question-answering-on-msvd-qa-1","task":"Visual Question Answering (VQA)","dataset":"MSVD-QA","model":"MA-LMM","rank_in_archive_order":2,"of":36,"metrics":{"Accuracy":"0.606"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2404.05726","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2404.05726"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/boheumd/MA-LMM","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran":7,"unverified":2},"by_repo_kind":{"official":{"samples":9,"ran":7,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"56f02812d66a0d11","entry":"generate_caption","repo":"boheumd/MA-LMM","repo_kind":"official","path":"app/caption.py","file_url":"https://github.com/boheumd/MA-LMM/blob/HEAD/app/caption.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"56f02812d66a0d11"}},{"code_sha256_prefix":"92f2e16ee0d3a24d","entry":"get_concat_v","repo":"boheumd/MA-LMM","repo_kind":"official","path":"app/dataset_browser.py","file_url":"https://github.com/boheumd/MA-LMM/blob/HEAD/app/dataset_browser.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"92f2e16ee0d3a24d"}},{"code_sha256_prefix":"13a43a7d815937e6","entry":"read_img","repo":"boheumd/MA-LMM","repo_kind":"official","path":"app/calculate_coco_features.py","file_url":"https://github.com/boheumd/MA-LMM/blob/HEAD/app/calculate_coco_features.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"13a43a7d815937e6"}},{"code_sha256_prefix":"926bb38f979a61ef","entry":"resize_img","repo":"boheumd/MA-LMM","repo_kind":"official","path":"app/utils.py","file_url":"https://github.com/boheumd/MA-LMM/blob/HEAD/app/utils.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"926bb38f979a61ef"}},{"code_sha256_prefix":"f809f2ecf34c1c99","entry":"resize_img_w","repo":"boheumd/MA-LMM","repo_kind":"official","path":"app/dataset_browser.py","file_url":"https://github.com/boheumd/MA-LMM/blob/HEAD/app/dataset_browser.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"f809f2ecf34c1c99"}},{"code_sha256_prefix":"3d4bbb5ca53fd24a","entry":"sample_dataset","repo":"boheumd/MA-LMM","repo_kind":"official","path":"app/dataset_browser.py","file_url":"https://github.com/boheumd/MA-LMM/blob/HEAD/app/dataset_browser.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"3d4bbb5ca53fd24a"}},{"code_sha256_prefix":"cb33571427334815","entry":"tile","repo":"boheumd/MA-LMM","repo_kind":"official","path":"lavis/models/base_model.py","file_url":"https://github.com/boheumd/MA-LMM/blob/HEAD/lavis/models/base_model.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"cb33571427334815"}},{"code_sha256_prefix":"0ec9fc2025c16f65","entry":"all_gather_with_grad","repo":"boheumd/MA-LMM","repo_kind":"official","path":"lavis/models/base_model.py","file_url":"https://github.com/boheumd/MA-LMM/blob/HEAD/lavis/models/base_model.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"0ec9fc2025c16f65"}},{"code_sha256_prefix":"3e12ea8c09674aa8","entry":"compute_gradcam_batch","repo":"boheumd/MA-LMM","repo_kind":"official","path":"app/multimodal_search.py","file_url":"https://github.com/boheumd/MA-LMM/blob/HEAD/app/multimodal_search.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"3e12ea8c09674aa8"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}