{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/improved-speech-enhancement-with-the-wave-u","title":"Improved Speech Enhancement with the Wave-U-Net","arxiv_id":"1811.11307","date":"2018-11-27","proceeding":null,"authors":["Craig Macartney","Tillman Weyde"],"abstract":"We study the use of the Wave-U-Net architecture for speech enhancement, a\nmodel introduced by Stoller et al for the separation of music vocals and\naccompaniment. This end-to-end learning method for audio source separation\noperates directly in the time domain, permitting the integrated modelling of\nphase information and being able to take large temporal contexts into account.\nOur experiments show that the proposed method improves several metrics, namely\nPESQ, CSIG, CBAK, COVL and SSNR, over the state-of-the-art with respect to the\nspeech enhancement task on the Voice Bank corpus (VCTK) dataset. We find that a\nreduced number of hidden layers is sufficient for speech enhancement in\ncomparison to the original system designed for singing voice separation in\nmusic. We see this initial result as an encouraging signal to further explore\nspeech enhancement in the time-domain, both as an end in itself and as a\npre-processing step to speech recognition systems.","url_abs":"http://arxiv.org/abs/1811.11307v1","url_pdf":"http://arxiv.org/pdf/1811.11307v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"improved-speech-enhancement-with-the-wave-u","repo_url":"https://github.com/MattSegal/speech-enhancement","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"improved-speech-enhancement-with-the-wave-u","repo_url":"https://github.com/craigmacartney/Wave-U-Net-For-Speech-Enhancement","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"improved-speech-enhancement-with-the-wave-u","repo_url":"https://github.com/pheepa/DCUnet","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"audio-source-separation","task_name":"Audio Source Separation"},{"task_slug":"speech-enhancement","task_name":"Speech Enhancement"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/speech-enhancement-on-demand-1","task":"Speech Enhancement","dataset":"DEMAND","model":"Wave-U-Net","rank_in_archive_order":1,"of":1,"metrics":{"CBAK":"3.24","COVL":"2.96","CSIG":"3.52","PESQ":"2.4"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1811.11307","atlas_url":"https://app.syntology.ai/?focus=1811.11307","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}