{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/schrodinger-bridge-for-generative-speech","title":"Schrödinger Bridge for Generative Speech Enhancement","arxiv_id":"2407.16074","date":"2024-07-22","proceeding":null,"authors":["Ante Jukić","Roman Korostik","Jagadeesh Balam","Boris Ginsburg"],"abstract":"This paper proposes a generative speech enhancement model based on Schr\\\"odinger bridge (SB). The proposed model is employing a tractable SB to formulate a data-to-data process between the clean speech distribution and the observed noisy speech distribution. The model is trained with a data prediction loss, aiming to recover the complex-valued clean speech coefficients, and an auxiliary time-domain loss is used to improve training of the model. The effectiveness of the proposed SB-based model is evaluated in two different speech enhancement tasks: speech denoising and speech dereverberation. The experimental results demonstrate that the proposed SB-based outperforms diffusion-based models in terms of speech quality metrics and ASR performance, e.g., resulting in relative word error rate reduction of 20% for denoising and 6% for dereverberation compared to the best baseline model. The proposed model also demonstrates improved efficiency, achieving better quality than the baselines for the same number of sampling steps and with a reduced computational cost.","url_abs":"https://arxiv.org/abs/2407.16074v1","url_pdf":"https://arxiv.org/pdf/2407.16074v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"denoising","task_name":"Denoising"},{"task_slug":"speech-denoising","task_name":"Speech Denoising"},{"task_slug":"speech-dereverberation","task_name":"Speech Dereverberation"},{"task_slug":"speech-enhancement","task_name":"Speech Enhancement"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/speech-enhancement-on-ears-wham","task":"Speech Enhancement","dataset":"EARS-WHAM","model":"Schrödinger Bridge","rank_in_archive_order":4,"of":6,"metrics":{"DNSMOS":"3.83","ESTOI":"0.73","PESQ-WB":"2.33","POLQA":"3.46","SI-SDR":"17.85","SIGMOS":"3.44"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2407.16074","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}