{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/cosyvoice-3-towards-in-the-wild-speech","title":"CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training","arxiv_id":"2505.17589","date":"2025-05-23","proceeding":null,"authors":["Zhihao Du","Changfeng Gao","Yuxuan Wang","Fan Yu","Tianyu Zhao","Hao Wang","Xiang Lv","Hui Wang","Chongjia Ni","Xian Shi","Keyu An","Guanrou Yang","Yabin Li","Yanni Chen","Zhifu Gao","Qian Chen","Yue Gu","Mengzhe Chen","Yafeng Chen","Shiliang Zhang","Wen Wang","Jieping Ye"],"abstract":"In our prior works, we introduced a scalable streaming speech synthesis model, CosyVoice 2, which integrates a large language model (LLM) and a chunk-aware flow matching (FM) model, and achieves low-latency bi-streaming speech synthesis and human-parity quality. Despite these advancements, CosyVoice 2 exhibits limitations in language coverage, domain diversity, data volume, text formats, and post-training techniques. In this paper, we present CosyVoice 3, an improved model designed for zero-shot multilingual speech synthesis in the wild, surpassing its predecessor in content consistency, speaker similarity, and prosody naturalness. Key features of CosyVoice 3 include: 1) A novel speech tokenizer to improve prosody naturalness, developed via supervised multi-task training, including automatic speech recognition, speech emotion recognition, language identification, audio event detection, and speaker analysis. 2) A new differentiable reward model for post-training applicable not only to CosyVoice 3 but also to other LLM-based speech synthesis models. 3) Dataset Size Scaling: Training data is expanded from ten thousand hours to one million hours, encompassing 9 languages and 18 Chinese dialects across various domains and text formats. 4) Model Size Scaling: Model parameters are increased from 0.5 billion to 1.5 billion, resulting in enhanced performance on our multilingual benchmark due to the larger model capacity. These advancements contribute significantly to the progress of speech synthesis in the wild. We encourage readers to listen to the demo at https://funaudiollm.github.io/cosyvoice3.","url_abs":"https://arxiv.org/abs/2505.17589v2","url_pdf":"https://arxiv.org/pdf/2505.17589v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"cosyvoice-3-towards-in-the-wild-speech","repo_url":"https://github.com/funaudiollm/cosyvoice","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}},{"paper_slug":"cosyvoice-3-towards-in-the-wild-speech","repo_url":"https://github.com/funaudiollm/cv3-eval","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"automatic-speech-recognition-2","task_name":"Automatic Speech Recognition"},{"task_slug":"emotion-recognition","task_name":"Emotion Recognition"},{"task_slug":"event-detection","task_name":"Event Detection"},{"task_slug":"language-identification","task_name":"Language Identification"},{"task_slug":"large-language-model","task_name":"Large Language Model"},{"task_slug":"speech-emotion-recognition","task_name":"Speech Emotion Recognition"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-synthesis","task_name":"Speech Synthesis"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2505.17589","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2505.17589"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/funaudiollm/cosyvoice","reach":{"status":"ok","spdx":"Apache-2.0"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/funaudiollm/cv3-eval","reach":{"status":"ok","spdx":"Apache-2.0"}}],"summary":{"ran_fixture":1,"unverified":3},"by_repo_kind":{"listed":{"samples":4,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"3e30d88eaef01190","entry":"off_diagonal","repo":"funaudiollm/cv3-eval","repo_kind":"listed","path":"utils/3D-Speaker/speakerlab/loss/dino_loss.py","file_url":"https://github.com/funaudiollm/cv3-eval/blob/HEAD/utils/3D-Speaker/speakerlab/loss/dino_loss.py","link_basis":"harvester_set","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"3e30d88eaef01190"}},{"code_sha256_prefix":"0cecf5b70e777fc6","entry":"compute_norm_counts","repo":"funaudiollm/cv3-eval","repo_kind":"listed","path":"utils/3D-Speaker/speakerlab/utils/score_metrics.py","file_url":"https://github.com/funaudiollm/cv3-eval/blob/HEAD/utils/3D-Speaker/speakerlab/utils/score_metrics.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"0cecf5b70e777fc6"}},{"code_sha256_prefix":"7948ea3900577d73","entry":"compute_pmiss_pfa","repo":"funaudiollm/cv3-eval","repo_kind":"listed","path":"utils/3D-Speaker/speakerlab/utils/score_metrics.py","file_url":"https://github.com/funaudiollm/cv3-eval/blob/HEAD/utils/3D-Speaker/speakerlab/utils/score_metrics.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"7948ea3900577d73"}},{"code_sha256_prefix":"9ea50e54c5b46c7d","entry":"compute_pmiss_pfa_rbst","repo":"funaudiollm/cv3-eval","repo_kind":"listed","path":"utils/3D-Speaker/speakerlab/utils/score_metrics.py","file_url":"https://github.com/funaudiollm/cv3-eval/blob/HEAD/utils/3D-Speaker/speakerlab/utils/score_metrics.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"9ea50e54c5b46c7d"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}