{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/voice-conversion-based-on-cross-domain","title":"Voice Conversion Based on Cross-Domain Features Using Variational Auto Encoders","arxiv_id":"1808.09634","date":"2018-08-29","proceeding":null,"authors":["Wen-Chin Huang","Hsin-Te Hwang","Yu-Huai Peng","Yu Tsao","Hsin-Min Wang"],"abstract":"An effective approach to non-parallel voice conversion (VC) is to utilize\ndeep neural networks (DNNs), specifically variational auto encoders (VAEs), to\nmodel the latent structure of speech in an unsupervised manner. A previous\nstudy has confirmed the ef- fectiveness of VAE using the STRAIGHT spectra for\nVC. How- ever, VAE using other types of spectral features such as mel- cepstral\ncoefficients (MCCs), which are related to human per- ception and have been\nwidely used in VC, have not been prop- erly investigated. Instead of using one\nspecific type of spectral feature, it is expected that VAE may benefit from\nusing multi- ple types of spectral features simultaneously, thereby improving\nthe capability of VAE for VC. To this end, we propose a novel VAE framework\n(called cross-domain VAE, CDVAE) for VC. Specifically, the proposed framework\nutilizes both STRAIGHT spectra and MCCs by explicitly regularizing multiple\nobjectives in order to constrain the behavior of the learned encoder and de-\ncoder. Experimental results demonstrate that the proposed CD- VAE framework\noutperforms the conventional VAE framework in terms of subjective tests.","url_abs":"http://arxiv.org/abs/1808.09634v1","url_pdf":"http://arxiv.org/pdf/1808.09634v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"voice-conversion-based-on-cross-domain","repo_url":"https://github.com/unilight/cdvae-vc","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"voice-conversion","task_name":"Voice Conversion"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}