{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/neural-vocoder-is-all-you-need-for-speech","title":"Neural Vocoder is All You Need for Speech Super-resolution","arxiv_id":"2203.14941","date":"2022-03-28","proceeding":null,"authors":["Haohe Liu","Woosung Choi","Xubo Liu","Qiuqiang Kong","Qiao Tian","DeLiang Wang"],"abstract":"Speech super-resolution (SR) is a task to increase speech sampling rate by generating high-frequency components. Existing speech SR methods are trained in constrained experimental settings, such as a fixed upsampling ratio. These strong constraints can potentially lead to poor generalization ability in mismatched real-world cases. In this paper, we propose a neural vocoder based speech super-resolution method (NVSR) that can handle a variety of input resolution and upsampling ratios. NVSR consists of a mel-bandwidth extension module, a neural vocoder module, and a post-processing module. Our proposed system achieves state-of-the-art results on the VCTK multi-speaker benchmark. On 44.1 kHz target resolution, NVSR outperforms WSRGlow and Nu-wave by 8% and 37% respectively on log spectral distance and achieves a significantly better perceptual quality. We also demonstrate that prior knowledge in the pre-trained vocoder is crucial for speech SR by performing mel-bandwidth extension with a simple replication-padding method. Samples can be found in https://haoheliu.github.io/nvsr.","url_abs":"https://arxiv.org/abs/2203.14941v1","url_pdf":"https://arxiv.org/pdf/2203.14941v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"neural-vocoder-is-all-you-need-for-speech","repo_url":"https://github.com/haoheliu/ssr_eval","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"all","task_name":"All"},{"task_slug":"audio-super-resolution","task_name":"Audio Super-Resolution"},{"task_slug":"bandwidth-extension","task_name":"Bandwidth Extension"},{"task_slug":"super-resolution","task_name":"Super-Resolution"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/audio-super-resolution-on-vctk-multi-speaker-1","task":"Audio Super-Resolution","dataset":"VCTK Multi-Speaker","model":"NVSR","rank_in_archive_order":2,"of":7,"metrics":{"Log-Spectral Distance":"0.78"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2203.14941","atlas_url":"https://app.syntology.ai/?focus=2203.14941","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}