{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/when-large-multimodal-models-confront","title":"When Large Multimodal Models Confront Evolving Knowledge:Challenges and Pathways","arxiv_id":"2505.24449","date":"2025-05-30","proceeding":null,"authors":["Kailin Jiang","Yuntao Du","Yukai Ding","Yuchen Ren","Ning Jiang","Zhi Gao","Zilong Zheng","Lei Liu","Bin Li","Qing Li"],"abstract":"Large language/multimodal models (LLMs/LMMs) store extensive pre-trained knowledge but struggle to maintain consistency with real-world updates, making it difficult to avoid catastrophic forgetting while acquiring evolving knowledge. Previous work focused on constructing textual knowledge datasets and exploring knowledge injection in LLMs, lacking exploration of multimodal evolving knowledge injection in LMMs. To address this, we propose the EVOKE benchmark to evaluate LMMs' ability to inject multimodal evolving knowledge in real-world scenarios. Meanwhile, a comprehensive evaluation of multimodal evolving knowledge injection revealed two challenges: (1) Existing knowledge injection methods perform terribly on evolving knowledge. (2) Supervised fine-tuning causes catastrophic forgetting, particularly instruction following ability is severely compromised. Additionally, we provide pathways and find that: (1) Text knowledge augmentation during the training phase improves performance, while image augmentation cannot achieve it. (2) Continual learning methods, especially Replay and MoELoRA, effectively mitigate forgetting. Our findings indicate that current knowledge injection methods have many limitations on evolving knowledge, which motivates further research on more efficient and stable knowledge injection methods.","url_abs":"https://arxiv.org/abs/2505.24449v1","url_pdf":"https://arxiv.org/pdf/2505.24449v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"when-large-multimodal-models-confront","repo_url":"https://github.com/EVOKE-LMM/EVOKE","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"continual-learning","task_name":"Continual Learning"},{"task_slug":"image-augmentation","task_name":"Image Augmentation"},{"task_slug":"instruction-following","task_name":"Instruction Following"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2505.24449","atlas_url":"https://app.syntology.ai/?focus=2505.24449","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2505.24449"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/EVOKE-LMM/EVOKE","reach":{"status":"ok"}}],"summary":{"ran_fixture":1,"ran_draft_wrong":3,"unverified":2},"by_repo_kind":{"community":{"samples":4,"ran":4,"repositories":2}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":2,"samples":[{"code_sha256_prefix":"19f01570b0a18e2b","entry":"softmax","repo":"pjlab-sys4nlp/llama-moe","repo_kind":"community","path":"smoe/entrypoint/eval/eval_mmlu_moe_0.py","file_url":"https://github.com/pjlab-sys4nlp/llama-moe/blob/HEAD/smoe/entrypoint/eval/eval_mmlu_moe_0.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":2,"contract_check":"RAISES","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"19f01570b0a18e2b"}},{"code_sha256_prefix":"765ee58c7d78fb1b","entry":"build_instruction_prompt","repo":"deepseek-ai/DeepSeek-MoE","repo_kind":"community","path":"finetune/finetune.py","file_url":"https://github.com/deepseek-ai/DeepSeek-MoE/blob/HEAD/finetune/finetune.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"765ee58c7d78fb1b"}},{"code_sha256_prefix":"cd763eaf1ac287e7","entry":"format_example","repo":"pjlab-sys4nlp/llama-moe","repo_kind":"community","path":"smoe/entrypoint/eval/eval_mmlu_moe_0.py","file_url":"https://github.com/pjlab-sys4nlp/llama-moe/blob/HEAD/smoe/entrypoint/eval/eval_mmlu_moe_0.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"cd763eaf1ac287e7"}},{"code_sha256_prefix":"6ab745408cb8648b","entry":"format_subject","repo":"pjlab-sys4nlp/llama-moe","repo_kind":"community","path":"smoe/entrypoint/eval/eval_mmlu_moe_0.py","file_url":"https://github.com/pjlab-sys4nlp/llama-moe/blob/HEAD/smoe/entrypoint/eval/eval_mmlu_moe_0.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"6ab745408cb8648b"}},{"code_sha256_prefix":"e2e5c3a95a6aebb4","entry":"save_image_to_local","repo":null,"repo_kind":null,"path":null,"file_url":null,"link_basis":"identical_code_first_harvested_elsewhere","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":null,"inline_ok":false,"mcp_get_code":{"code_sha256":"e2e5c3a95a6aebb4"}},{"code_sha256_prefix":"22930d58d32e2d08","entry":"save_video_to_local","repo":null,"repo_kind":null,"path":null,"file_url":null,"link_basis":"identical_code_first_harvested_elsewhere","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":null,"inline_ok":false,"mcp_get_code":{"code_sha256":"22930d58d32e2d08"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}