{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/wgansing","title":"WGANSing","arxiv_id":null,"date":"2019-03-01","proceeding":"Interspeech 2019 3","authors":["Pritish Chandna","Merlijn Blaauw"],"abstract":"We present a deep neural network based singing\r\nvoice synthesizer, inspired by the Deep Convolutions Generative\r\nAdversarial Networks (DCGAN) architecture and optimized using the Wasserstein-GAN algorithm. We use vocoder parameters\r\nfor acoustic modelling, to separate the influence of pitch and\r\ntimbre. This facilitates the modelling of the large variability\r\nof pitch in the singing voice. Our network takes a block of\r\nconsecutive frame-wise linguistic and fundamental frequency\r\nfeatures, along with global singer identity as input and outputs\r\nvocoder features, corresponding to the block of features. This\r\nblock-wise approach, along with the training methodology allows\r\nus to model temporal dependencies within the features of the\r\ninput block. For inference, sequential blocks are concatenated\r\nusing an overlap-add procedure. We show that the performance\r\nof our model is competitive with regards to the state-of-the-art\r\nand the original sample using objective metrics and a subjective\r\nlistening test. We also present examples of the synthesis on a\r\nsupplementary website and the source code via GitHub.","url_abs":"https://arxiv.org/pdf/1903.10729.pdf","url_pdf":"https://arxiv.org/pdf/1903.10729.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"wgansing","repo_url":"https://github.com/MTG/WGANSing","is_official":0,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"acoustic-modelling","task_name":"Acoustic Modelling"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}