{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/multi-target-voice-conversion-without","title":"Multi-target Voice Conversion without Parallel Data by Adversarially Learning Disentangled Audio Representations","arxiv_id":"1804.02812","date":"2018-04-09","proceeding":null,"authors":["Ju-chieh Chou","Cheng-chieh Yeh","Hung-Yi Lee","Lin-shan Lee"],"abstract":"Recently, cycle-consistent adversarial network (Cycle-GAN) has been\nsuccessfully applied to voice conversion to a different speaker without\nparallel data, although in those approaches an individual model is needed for\neach target speaker. In this paper, we propose an adversarial learning\nframework for voice conversion, with which a single model can be trained to\nconvert the voice to many different speakers, all without parallel data, by\nseparating the speaker characteristics from the linguistic content in speech\nsignals. An autoencoder is first trained to extract speaker-independent latent\nrepresentations and speaker embedding separately using another auxiliary\nspeaker classifier to regularize the latent representation. The decoder then\ntakes the speaker-independent latent representation and the target speaker\nembedding as the input to generate the voice of the target speaker with the\nlinguistic content of the source utterance. The quality of decoder output is\nfurther improved by patching with the residual signal produced by another pair\nof generator and discriminator. A target speaker set size of 20 was tested in\nthe preliminary experiments, and very good voice quality was obtained.\nConventional voice conversion metrics are reported. We also show that the\nspeaker information has been properly reduced from the latent representations.","url_abs":"http://arxiv.org/abs/1804.02812v2","url_pdf":"http://arxiv.org/pdf/1804.02812v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"multi-target-voice-conversion-without","repo_url":"https://github.com/jjery2243542/voice_conversion","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"multi-target-voice-conversion-without","repo_url":"https://github.com/arshd91/multitarget-voice-conversion-vctk","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"multi-target-voice-conversion-without","repo_url":"https://github.com/BogiHsu/Voice-Conversion-PyTorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"voice-conversion","task_name":"Voice Conversion"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1804.02812","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1804.02812"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/arshd91/multitarget-voice-conversion-vctk","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/BogiHsu/Voice-Conversion-PyTorch","reach":{"status":"unanswered"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/jjery2243542/voice_conversion","reach":{"status":"unanswered"}}],"summary":{"unverified":2},"by_repo_kind":{"listed":{"samples":2,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":2,"samples":[{"code_sha256_prefix":"61e4acba753ed3a4","entry":"get_world_param","repo":"arshd91/multitarget-voice-conversion-vctk","repo_kind":"listed","path":"convert.py","file_url":"https://github.com/arshd91/multitarget-voice-conversion-vctk/blob/HEAD/convert.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"61e4acba753ed3a4"}},{"code_sha256_prefix":"639d82758c9e841c","entry":"synthesis","repo":"arshd91/multitarget-voice-conversion-vctk","repo_kind":"listed","path":"convert.py","file_url":"https://github.com/arshd91/multitarget-voice-conversion-vctk/blob/HEAD/convert.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"639d82758c9e841c"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}