{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/whispered-to-voiced-alaryngeal-speech","title":"Whispered-to-voiced Alaryngeal Speech Conversion with Generative Adversarial Networks","arxiv_id":"1808.10687","date":"2018-08-31","proceeding":null,"authors":["Santiago Pascual","Antonio Bonafonte","Joan Serrà","Jose A. Gonzalez"],"abstract":"Most methods of voice restoration for patients suffering from aphonia either\nproduce whispered or monotone speech. Apart from intelligibility, this type of\nspeech lacks expressiveness and naturalness due to the absence of pitch\n(whispered speech) or artificial generation of it (monotone speech). Existing\ntechniques to restore prosodic information typically combine a vocoder, which\nparameterises the speech signal, with machine learning techniques that predict\nprosodic information. In contrast, this paper describes an end-to-end neural\napproach for estimating a fully-voiced speech waveform from whispered\nalaryngeal speech. By adapting our previous work in speech enhancement with\ngenerative adversarial networks, we develop a speaker-dependent model to\nperform whispered-to-voiced speech conversion. Preliminary qualitative results\nshow effectiveness in re-generating voiced speech, with the creation of\nrealistic pitch contours.","url_abs":"http://arxiv.org/abs/1808.10687v2","url_pdf":"http://arxiv.org/pdf/1808.10687v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"whispered-to-voiced-alaryngeal-speech","repo_url":"https://github.com/develooper1994/MasterThesis","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"whispered-to-voiced-alaryngeal-speech","repo_url":"https://github.com/rickyHong/segan-pytorch-repl","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"whispered-to-voiced-alaryngeal-speech","repo_url":"https://github.com/santi-pdp/segan_pytorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"whispered-to-voiced-alaryngeal-speech","repo_url":"https://github.com/bacnguyenne/ASR_LibriSpeech","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"speech-enhancement","task_name":"Speech Enhancement"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}