Papers › Schrödinger Bridge for Generative Speech Enhancement
Schrödinger Bridge for Generative Speech Enhancement
Ante Jukić, Roman Korostik, Jagadeesh Balam, Boris Ginsburg
This paper proposes a generative speech enhancement model based on Schr\"odinger bridge (SB). The proposed model is employing a tractable SB to formulate a data-to-data process between the clean speech distribution and the observed noisy speech distribution. The model is trained with a data prediction loss, aiming to recover the complex-valued clean speech coefficients, and an auxiliary time-domain loss is used to improve training of the model. The effectiveness of the proposed SB-based model is evaluated in two different speech enhancement tasks: speech denoising and speech dereverberation. The experimental results demonstrate that the proposed SB-based outperforms diffusion-based models in terms of speech quality metrics and ASR performance, e.g., resulting in relative word error rate reduction of 20% for denoising and 6% for dereverberation compared to the best baseline model. The proposed model also demonstrates improved efficiency, achieving better quality than the baselines for the same number of sampling steps and with a reduced computational cost.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Speech Enhancement | EARS-WHAM | Schrödinger Bridge | DNSMOS | 3.83 | #4 of 6 | Archive leaderboard | report |
| Speech Enhancement | EARS-WHAM | Schrödinger Bridge | ESTOI | 0.73 | #4 of 6 | Archive leaderboard | report |
| Speech Enhancement | EARS-WHAM | Schrödinger Bridge | PESQ-WB | 2.33 | #4 of 6 | Archive leaderboard | report |
| Speech Enhancement | EARS-WHAM | Schrödinger Bridge | POLQA | 3.46 | #4 of 6 | Archive leaderboard | report |
| Speech Enhancement | EARS-WHAM | Schrödinger Bridge | SI-SDR | 17.85 | #4 of 6 | Archive leaderboard | report |
| Speech Enhancement | EARS-WHAM | Schrödinger Bridge | SIGMOS | 3.44 | #4 of 6 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections