{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/qwen2-audio-technical-report","title":"Qwen2-Audio Technical Report","arxiv_id":"2407.10759","date":"2024-07-15","proceeding":null,"authors":["Yunfei Chu","Jin Xu","Qian Yang","Haojie Wei","Xipin Wei","Zhifang Guo","Yichong Leng","YuanJun Lv","Jinzheng He","Junyang Lin","Chang Zhou","Jingren Zhou"],"abstract":"We introduce the latest progress of Qwen-Audio, a large-scale audio-language model called Qwen2-Audio, which is capable of accepting various audio signal inputs and performing audio analysis or direct textual responses with regard to speech instructions. In contrast to complex hierarchical tags, we have simplified the pre-training process by utilizing natural language prompts for different data and tasks, and have further expanded the data volume. We have boosted the instruction-following capability of Qwen2-Audio and implemented two distinct audio interaction modes for voice chat and audio analysis. In the voice chat mode, users can freely engage in voice interactions with Qwen2-Audio without text input. In the audio analysis mode, users could provide audio and text instructions for analysis during the interaction. Note that we do not use any system prompts to switch between voice chat and audio analysis modes. Qwen2-Audio is capable of intelligently comprehending the content within audio and following voice commands to respond appropriately. For instance, in an audio segment that simultaneously contains sounds, multi-speaker conversations, and a voice command, Qwen2-Audio can directly understand the command and provide an interpretation and response to the audio. Additionally, DPO has optimized the model's performance in terms of factuality and adherence to desired behavior. According to the evaluation results from AIR-Bench, Qwen2-Audio outperformed previous SOTAs, such as Gemini-1.5-pro, in tests focused on audio-centric instruction-following capabilities. Qwen2-Audio is open-sourced with the aim of fostering the advancement of the multi-modal language community.","url_abs":"https://arxiv.org/abs/2407.10759v1","url_pdf":"https://arxiv.org/pdf/2407.10759v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"qwen2-audio-technical-report","repo_url":"https://github.com/qwenlm/qwen2-audio","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"qwen2-audio-technical-report","repo_url":"https://github.com/MindCode-4/code-4/tree/main/qwen2","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null}],"tasks":[{"task_slug":"instruction-following","task_name":"Instruction Following"},{"task_slug":"language-modelling","task_name":"Language Modelling"}],"methods":[{"method_slug":"dpo","method_name":"DPO"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2407.10759","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2407.10759"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/qwenlm/qwen2-audio","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/MindCode-4/code-4/tree/main/qwen2","reach":null}],"summary":{"ran_draft_wrong":3,"ran_honours":1},"by_repo_kind":{"official":{"samples":4,"ran":4,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":4,"samples":[{"code_sha256_prefix":"08d897ce52e97690","entry":"add_file","repo":"qwenlm/qwen2-audio","repo_kind":"official","path":"demo/web_demo_audio.py","file_url":"https://github.com/qwenlm/qwen2-audio/blob/HEAD/demo/web_demo_audio.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"08d897ce52e97690"}},{"code_sha256_prefix":"467458f52cca4777","entry":"add_text","repo":"qwenlm/qwen2-audio","repo_kind":"official","path":"demo/web_demo_audio.py","file_url":"https://github.com/qwenlm/qwen2-audio/blob/HEAD/demo/web_demo_audio.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"467458f52cca4777"}},{"code_sha256_prefix":"1348aea7102035ce","entry":"remove_sp","repo":"qwenlm/qwen2-audio","repo_kind":"official","path":"eval_audio/evaluate_asr.py","file_url":"https://github.com/qwenlm/qwen2-audio/blob/HEAD/eval_audio/evaluate_asr.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"1348aea7102035ce"}},{"code_sha256_prefix":"257328f3459a2748","entry":"reset_state","repo":"qwenlm/qwen2-audio","repo_kind":"official","path":"demo/web_demo_audio.py","file_url":"https://github.com/qwenlm/qwen2-audio/blob/HEAD/demo/web_demo_audio.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"257328f3459a2748"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}