{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/fine-grained-audio-visual-joint","title":"Fine-grained Audio-Visual Joint Representations for Multimodal Large Language Models","arxiv_id":"2310.05863","date":"2023-10-09","proceeding":null,"authors":["Guangzhi Sun","Wenyi Yu","Changli Tang","Xianzhao Chen","Tian Tan","Wei Li","Lu Lu","Zejun Ma","Chao Zhang"],"abstract":"Audio-visual large language models (LLM) have drawn significant attention, yet the fine-grained combination of both input streams is rather under-explored, which is challenging but necessary for LLMs to understand general video inputs. To this end, a fine-grained audio-visual joint representation (FAVOR) learning framework for multimodal LLMs is proposed in this paper, which extends a text-based LLM to simultaneously perceive speech and audio events in the audio input stream and images or videos in the visual input stream, at the frame level. To fuse the audio and visual feature streams into joint representations and to align the joint space with the LLM input embedding space, we propose a causal Q-Former structure with a causal attention module to enhance the capture of causal relations of the audio-visual frames across time. An audio-visual evaluation benchmark (AVEB) is also proposed which comprises six representative single-modal tasks with five cross-modal tasks reflecting audio-visual co-reasoning abilities. While achieving competitive single-modal performance on audio, speech and image tasks in AVEB, FAVOR achieved over 20% accuracy improvements on the video question-answering task when fine-grained information or temporal causal reasoning is required. FAVOR, in addition, demonstrated remarkable video comprehension and reasoning abilities on tasks that are unprecedented by other multimodal LLMs. An interactive demo of FAVOR is available at https://github.com/BriansIDP/AudioVisualLLM.git, and the training code and model checkpoints will be released soon.","url_abs":"https://arxiv.org/abs/2310.05863v2","url_pdf":"https://arxiv.org/pdf/2310.05863v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"fine-grained-audio-visual-joint","repo_url":"https://github.com/briansidp/audiovisualllm","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null},{"paper_slug":"fine-grained-audio-visual-joint","repo_url":"https://github.com/the-anonymous-bs/favor","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"video-question-answering","task_name":"Video Question Answering"}],"methods":[{"method_slug":"align","method_name":"ALIGN"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2310.05863","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2310.05863"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/briansidp/audiovisualllm","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/the-anonymous-bs/favor","reach":{"status":"ok"}}],"summary":{"ran_fixture":2,"ran_violates":1,"ran":3,"ran_draft_wrong":2,"unverified":4},"by_repo_kind":{"official":{"samples":12,"ran":8,"repositories":2}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":10,"samples":[{"code_sha256_prefix":"9b4dff79d5e6102c","entry":"apply_rotary_pos_emb","repo":"the-anonymous-bs/favor","repo_kind":"official","path":"video_llama/models/modeling_llama.py","file_url":"https://github.com/the-anonymous-bs/favor/blob/HEAD/video_llama/models/modeling_llama.py","link_basis":"harvester_set","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"9b4dff79d5e6102c"}},{"code_sha256_prefix":"4cb732f513d69dfd","entry":"disabled_train","repo":"the-anonymous-bs/favor","repo_kind":"official","path":"video_llama/models/blip2.py","file_url":"https://github.com/the-anonymous-bs/favor/blob/HEAD/video_llama/models/blip2.py","link_basis":"harvester_set","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"4cb732f513d69dfd"}},{"code_sha256_prefix":"7d214655954e2bc6","entry":"getAttMap","repo":"the-anonymous-bs/favor","repo_kind":"official","path":"video_llama/common/gradcam.py","file_url":"https://github.com/the-anonymous-bs/favor/blob/HEAD/video_llama/common/gradcam.py","link_basis":"harvester_set","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"7d214655954e2bc6"}},{"code_sha256_prefix":"98589643273920ba","entry":"main_process","repo":"the-anonymous-bs/favor","repo_kind":"official","path":"video_llama/common/dist_utils.py","file_url":"https://github.com/the-anonymous-bs/favor/blob/HEAD/video_llama/common/dist_utils.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"98589643273920ba"}},{"code_sha256_prefix":"791c72070c0b1cfb","entry":"node_to_dict","repo":"the-anonymous-bs/favor","repo_kind":"official","path":"video_llama/common/config.py","file_url":"https://github.com/the-anonymous-bs/favor/blob/HEAD/video_llama/common/config.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"791c72070c0b1cfb"}},{"code_sha256_prefix":"826ceb601fb31250","entry":"parse_text","repo":"briansidp/audiovisualllm","repo_kind":"official","path":"inference.py","file_url":"https://github.com/briansidp/audiovisualllm/blob/HEAD/inference.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"826ceb601fb31250"}},{"code_sha256_prefix":"b99eea6376d1e212","entry":"rotate_half","repo":"the-anonymous-bs/favor","repo_kind":"official","path":"video_llama/models/modeling_llama.py","file_url":"https://github.com/the-anonymous-bs/favor/blob/HEAD/video_llama/models/modeling_llama.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"b99eea6376d1e212"}},{"code_sha256_prefix":"cb33571427334815","entry":"tile","repo":"the-anonymous-bs/favor","repo_kind":"official","path":"video_llama/models/base_model.py","file_url":"https://github.com/the-anonymous-bs/favor/blob/HEAD/video_llama/models/base_model.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"cb33571427334815"}},{"code_sha256_prefix":"0ec9fc2025c16f65","entry":"all_gather_with_grad","repo":"the-anonymous-bs/favor","repo_kind":"official","path":"video_llama/models/base_model.py","file_url":"https://github.com/the-anonymous-bs/favor/blob/HEAD/video_llama/models/base_model.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"0ec9fc2025c16f65"}},{"code_sha256_prefix":"0f1417dbe45c5b50","entry":"create_eva_vit_g","repo":"the-anonymous-bs/favor","repo_kind":"official","path":"video_llama/models/eva_vit.py","file_url":"https://github.com/the-anonymous-bs/favor/blob/HEAD/video_llama/models/eva_vit.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"0f1417dbe45c5b50"}},{"code_sha256_prefix":"643539116e049080","entry":"download_cached_file","repo":"the-anonymous-bs/favor","repo_kind":"official","path":"video_llama/common/dist_utils.py","file_url":"https://github.com/the-anonymous-bs/favor/blob/HEAD/video_llama/common/dist_utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"643539116e049080"}},{"code_sha256_prefix":"3310bb33a7fee07b","entry":"predict","repo":"briansidp/audiovisualllm","repo_kind":"official","path":"inference.py","file_url":"https://github.com/briansidp/audiovisualllm/blob/HEAD/inference.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"3310bb33a7fee07b"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}