{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/convolutional-neural-networks-to-enhance","title":"Convolutional Neural Networks to Enhance Coded Speech","arxiv_id":"1806.09411","date":"2018-06-25","proceeding":null,"authors":["Ziyue Zhao","Huijun Liu","Tim Fingscheidt"],"abstract":"Enhancing coded speech suffering from far-end acoustic background noise,\nquantization noise, and potentially transmission errors, is a challenging task.\nIn this work we propose two postprocessing approaches applying convolutional\nneural networks (CNNs) either in the time domain or the cepstral domain to\nenhance the coded speech without any modification of the codecs. The time\ndomain approach follows an end-to-end fashion, while the cepstral domain\napproach uses analysis-synthesis with cepstral domain features. The proposed\npostprocessors in both domains are evaluated for various narrowband and\nwideband speech codecs in a wide range of conditions. The proposed\npostprocessor improves speech quality (PESQ) by up to 0.25 MOS-LQO points for\nG.711, 0.30 points for G.726, 0.82 points for G.722, and 0.26 points for\nadaptive multirate wideband codec (AMR-WB). In a subjective CCR listening test,\nthe proposed postprocessor on G.711-coded speech exceeds the speech quality of\nan ITU-T-standardized postfilter by 0.36 CMOS points, and obtains a clear\npreference of 1.77 CMOS points compared to G.711, even en par with uncoded\nspeech. The source code for the cepstral domain approach to enhance G.711-coded\nspeech is made available.","url_abs":"http://arxiv.org/abs/1806.09411v1","url_pdf":"http://arxiv.org/pdf/1806.09411v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"convolutional-neural-networks-to-enhance","repo_url":"https://github.com/ifnspaml/Enhancement-Coded-Speech","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"quantization","task_name":"Quantization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}