{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/statistical-parametric-speech-synthesis-1","title":"Statistical Parametric Speech Synthesis Incorporating Generative Adversarial Networks","arxiv_id":"1709.08041","date":"2017-09-23","proceeding":null,"authors":["Yuki Saito","Shinnosuke Takamichi","Hiroshi Saruwatari"],"abstract":"A method for statistical parametric speech synthesis incorporating generative\nadversarial networks (GANs) is proposed. Although powerful deep neural networks\n(DNNs) techniques can be applied to artificially synthesize speech waveform,\nthe synthetic speech quality is low compared with that of natural speech. One\nof the issues causing the quality degradation is an over-smoothing effect often\nobserved in the generated speech parameters. A GAN introduced in this paper\nconsists of two neural networks: a discriminator to distinguish natural and\ngenerated samples, and a generator to deceive the discriminator. In the\nproposed framework incorporating the GANs, the discriminator is trained to\ndistinguish natural and generated speech parameters, while the acoustic models\nare trained to minimize the weighted sum of the conventional minimum generation\nloss and an adversarial loss for deceiving the discriminator. Since the\nobjective of the GANs is to minimize the divergence (i.e., distribution\ndifference) between the natural and generated speech parameters, the proposed\nmethod effectively alleviates the over-smoothing effect on the generated speech\nparameters. We evaluated the effectiveness for text-to-speech and voice\nconversion, and found that the proposed method can generate more natural\nspectral parameters and $F_0$ than conventional minimum generation error\ntraining algorithm regardless its hyper-parameter settings. Furthermore, we\ninvestigated the effect of the divergence of various GANs, and found that a\nWasserstein GAN minimizing the Earth-Mover's distance works the best in terms\nof improving synthetic speech quality.","url_abs":"http://arxiv.org/abs/1709.08041v1","url_pdf":"http://arxiv.org/pdf/1709.08041v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"statistical-parametric-speech-synthesis-1","repo_url":"https://github.com/r9y9/gantts","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"statistical-parametric-speech-synthesis-1","repo_url":"https://github.com/rickyHong/GAN-TTS-repl2","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"statistical-parametric-speech-synthesis-1","repo_url":"https://github.com/rickyHong/GANTTS-update-repl","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"statistical-parametric-speech-synthesis-1","repo_url":"https://github.com/rickyHong/gantts-update-repl2","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"statistical-parametric-speech-synthesis-1","repo_url":"https://github.com/nafiuny/ICRCycleGAN-VC","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"speech-synthesis","task_name":"Speech Synthesis"},{"task_slug":"text-to-speech","task_name":"Text to Speech"},{"task_slug":"voice-conversion","task_name":"Voice Conversion"},{"task_slug":"text-to-speech-1","task_name":"text-to-speech"}],"methods":[{"method_slug":"convolution","method_name":"Convolution"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}