{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/video-salmonn-2-captioning-enhanced-audio","title":"video-SALMONN 2: Captioning-Enhanced Audio-Visual Large Language Models","arxiv_id":"2506.15220","date":"2025-06-18","proceeding":null,"authors":["Changli Tang","Yixuan Li","Yudong Yang","Jimin Zhuang","Guangzhi Sun","Wei Li","Zejun Ma","Chao Zhang"],"abstract":"Videos contain a wealth of information, and generating detailed and accurate descriptions in natural language is a key aspect of video understanding. In this paper, we present video-SALMONN 2, an advanced audio-visual large language model (LLM) with low-rank adaptation (LoRA) designed for enhanced video (with paired audio) captioning through directed preference optimisation (DPO). We propose new metrics to evaluate the completeness and accuracy of video descriptions, which are optimised using DPO. To further improve training, we propose a novel multi-round DPO (MrDPO) approach, which involves periodically updating the DPO reference model, merging and re-initialising the LoRA module as a proxy for parameter updates after each training round (1,000 steps), and incorporating guidance from ground-truth video captions to stabilise the process. Experimental results show that MrDPO significantly enhances video-SALMONN 2's captioning accuracy, reducing the captioning error rates by 28\\%. The final video-SALMONN 2 model, with just 7 billion parameters, surpasses leading models such as GPT-4o and Gemini-1.5-Pro in video captioning tasks, while maintaining highly competitive performance to the state-of-the-art on widely used video question-answering benchmarks among models of similar size. Codes are available at \\href{https://github.com/bytedance/video-SALMONN-2}{https://github.com/bytedance/video-SALMONN-2}.","url_abs":"https://arxiv.org/abs/2506.15220v1","url_pdf":"https://arxiv.org/pdf/2506.15220v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"video-salmonn-2-captioning-enhanced-audio","repo_url":"https://github.com/bytedance/video-salmonn-2","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"audio-captioning","task_name":"Audio captioning"},{"task_slug":"large-language-model","task_name":"Large Language Model"},{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"video-captioning","task_name":"Video Captioning"},{"task_slug":"video-question-answering","task_name":"Video Question Answering"},{"task_slug":"video-understanding","task_name":"Video Understanding"}],"methods":[{"method_slug":"dpo","method_name":"DPO"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2506.15220","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2506.15220"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/bytedance/video-salmonn-2","reach":{"status":"ok","spdx":"Apache-2.0"}}],"summary":{"ran_fixture":1,"ran_draft_wrong":2,"unverified":5},"by_repo_kind":{"official":{"samples":8,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"3c76e52815c5401d","entry":"repeat_kv","repo":"bytedance/video-salmonn-2","repo_kind":"official","path":"video_SALMONN2_pro/qwenvl/model/modeling_qwen3_vl.py","file_url":"https://github.com/bytedance/video-salmonn-2/blob/HEAD/video_SALMONN2_pro/qwenvl/model/modeling_qwen3_vl.py","link_basis":"harvester_set","language":"python","status":"ran_fixture","verification_level":2,"contract_check":"RAISES","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"3c76e52815c5401d"}},{"code_sha256_prefix":"e03d53ba9d4f9ae5","entry":"rotate_half","repo":"bytedance/video-salmonn-2","repo_kind":"official","path":"video_SALMONN2_plus/qwenvl/model/modeling_qwen2_5_vl.py","file_url":"https://github.com/bytedance/video-salmonn-2/blob/HEAD/video_SALMONN2_plus/qwenvl/model/modeling_qwen2_5_vl.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"e03d53ba9d4f9ae5"}},{"code_sha256_prefix":"9dfcf8b13847c65f","entry":"shift_tokens_right","repo":"bytedance/video-salmonn-2","repo_kind":"official","path":"video_SALMONN2_pro/qwenvl/model/modeling_whisper.py","file_url":"https://github.com/bytedance/video-salmonn-2/blob/HEAD/video_SALMONN2_pro/qwenvl/model/modeling_whisper.py","link_basis":"plan_row","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"9dfcf8b13847c65f"}},{"code_sha256_prefix":"24acaeb8ad82253e","entry":"apply_rotary_pos_emb_flashatt","repo":"bytedance/video-salmonn-2","repo_kind":"official","path":"video_SALMONN2_plus/qwenvl/model/modeling_qwen2_5_vl.py","file_url":"https://github.com/bytedance/video-salmonn-2/blob/HEAD/video_SALMONN2_plus/qwenvl/model/modeling_qwen2_5_vl.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"24acaeb8ad82253e"}},{"code_sha256_prefix":"049747a79920f0f7","entry":"apply_rotary_pos_emb_vision","repo":"bytedance/video-salmonn-2","repo_kind":"official","path":"video_SALMONN2_pro/qwenvl/model/modeling_qwen3_vl.py","file_url":"https://github.com/bytedance/video-salmonn-2/blob/HEAD/video_SALMONN2_pro/qwenvl/model/modeling_qwen3_vl.py","link_basis":"plan_row","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"049747a79920f0f7"}},{"code_sha256_prefix":"5126650c89c15a91","entry":"eager_attention_forward","repo":"bytedance/video-salmonn-2","repo_kind":"official","path":"video_SALMONN2_pro/qwenvl/model/modeling_whisper.py","file_url":"https://github.com/bytedance/video-salmonn-2/blob/HEAD/video_SALMONN2_pro/qwenvl/model/modeling_whisper.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"5126650c89c15a91"}},{"code_sha256_prefix":"50982cb2b9cc7e0a","entry":"pad_and_stack","repo":"bytedance/video-salmonn-2","repo_kind":"official","path":"video_SALMONN2_plus/qwenvl/model/modeling_qwen2_5_vl.py","file_url":"https://github.com/bytedance/video-salmonn-2/blob/HEAD/video_SALMONN2_plus/qwenvl/model/modeling_qwen2_5_vl.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"50982cb2b9cc7e0a"}},{"code_sha256_prefix":"ec2159dcc8c21e1d","entry":"sinusoids","repo":"bytedance/video-salmonn-2","repo_kind":"official","path":"video_SALMONN2_plus/qwenvl/model/modeling_whisper.py","file_url":"https://github.com/bytedance/video-salmonn-2/blob/HEAD/video_SALMONN2_plus/qwenvl/model/modeling_whisper.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"ec2159dcc8c21e1d"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}