{"url":"/sota/speech-enhancement-on-deep-noise-suppression","task":{"name":"Speech Enhancement","url":"/task/speech-enhancement","note":null},"dataset":{"name":"Deep Noise Suppression (DNS) Challenge","url":"/dataset/deep-noise-suppression-2020"},"category":"Audio","categories":["Audio","Speech"],"category_note":null,"description":"**Speech Enhancement** is a signal processing task that involves improving the quality of speech signals captured under noisy or degraded conditions. The goal of speech enhancement is to make speech signals clearer, more intelligible, and more pleasant to listen to, which can be used for various applications such as voice recognition, teleconferencing, and hearing aids. A representative Github project with online demo : [ClearerVoice-Studio](https://github.com/modelscope/ClearerVoice-Studio).\r\n\r\n<span style=\"color:grey; opacity: 0.6\">( Image credit: [A Fully Convolutional Neural Network For Speech Enhancement](https://arxiv.org/pdf/1609.07132v1.pdf) )</span>","description_from":"task","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","rank":"the archive's row order at snapshot; not re-ranked","rows_end_at":"2025-07-28","rows_withheld_as_spam":0,"metric_values":"the archive's strings, untouched"},"metrics":["PESQ-WB","SI-SDR-WB","STOI","PESQ-NB","SI-SDR-NB","Number of parameters (M)","FLOPS (G)","ESTOI","SSNR"],"metric_direction":{"note":"inferred from the metric name only (the archive records no direction); null = not inferred, chart draws points only","by_metric":{"PESQ-WB":null,"SI-SDR-WB":null,"STOI":null,"PESQ-NB":null,"SI-SDR-NB":null,"Number of parameters (M)":null,"FLOPS (G)":"lower","ESTOI":null,"SSNR":null}},"counts":{"rows":36,"rows_with_code":24,"rows_with_paper_page":33,"rows_dated":33,"rows_using_additional_data":2},"rows":[{"rank_in_archive_order":1,"model":"ZipEnhancer (M)","metrics":{"FLOPS (G)":"266.96","Number of parameters (M)":"11.34","PESQ-NB":"4.08","PESQ-WB":"3.81","SI-SDR-WB":"22.22","STOI":"98.65"},"uses_additional_data":false,"paper_date":null,"paper":null,"paper_url":null,"paper_title":"","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":2,"model":"TF-Locoformer (M)","metrics":{"FLOPS (G)":"497.24","Number of parameters (M)":"15","PESQ-WB":"3.72","SI-SDR-WB":"23.3","STOI":"98.8"},"uses_additional_data":false,"paper_date":"2024-08-06","paper":"/paper/tf-locoformer-transformer-with-local-modeling","paper_url":"https://arxiv.org/abs/2408.03440v1","paper_title":"TF-Locoformer: Transformer with Local Modeling by Convolution for Speech Separation and Enhancement","code":"https://github.com/merlresearch/tf-locoformer","n_code_links":1,"syntology":null},{"rank_in_archive_order":3,"model":"ZipEnhancer (S)","metrics":{"FLOPS (G)":"62.85","Number of parameters (M)":"2.04","PESQ-NB":"3.99","PESQ-WB":"3.69","SI-SDR-WB":"21.15","STOI":"98.32"},"uses_additional_data":false,"paper_date":null,"paper":null,"paper_url":null,"paper_title":"","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":4,"model":"MambAttention","metrics":{"ESTOI":"95.9","Number of parameters (M)":"2.33","PESQ-WB":"3.671","SI-SDR-WB":"21.234","SSNR":"15.116"},"uses_additional_data":false,"paper_date":"2025-07-01","paper":"/paper/mambattention-mamba-with-multi-head-attention","paper_url":"https://arxiv.org/abs/2507.00966v1","paper_title":"MambAttention: Mamba with Multi-Head Attention for Generalizable Single-Channel Speech Enhancement","code":"https://github.com/nikolaikyhne/xlstm-senet","n_code_links":2,"syntology":null},{"rank_in_archive_order":5,"model":"MP-SENet","metrics":{"PESQ-NB":"3.92","PESQ-WB":"3.62","SI-SDR-WB":"21.03"},"uses_additional_data":false,"paper_date":"2023-08-17","paper":"/paper/explicit-estimation-of-magnitude-and-phase","paper_url":"https://arxiv.org/abs/2308.08926v2","paper_title":"Explicit Estimation of Magnitude and Phase Spectra in Parallel for High-Quality Speech Enhancement","code":"https://github.com/yxlu-0102/MP-SENet","n_code_links":1,"syntology":{"n_ran":13,"n_unverified":1,"n_samples":14,"n_pointer_only_licence":0}},{"rank_in_archive_order":6,"model":"xLSTM-SENet","metrics":{"ESTOI":"95.4","Number of parameters (M)":"2.20","PESQ-WB":"3.588","SI-SDR-WB":"20.854","SSNR":"14.526"},"uses_additional_data":false,"paper_date":"2025-07-01","paper":"/paper/mambattention-mamba-with-multi-head-attention","paper_url":"https://arxiv.org/abs/2507.00966v1","paper_title":"MambAttention: Mamba with Multi-Head Attention for Generalizable Single-Channel Speech Enhancement","code":"https://github.com/nikolaikyhne/xlstm-senet","n_code_links":2,"syntology":null},{"rank_in_archive_order":7,"model":"BSRNN-S + MRSD","metrics":{"PESQ-NB":"3.89","PESQ-WB":"3.53","SI-SDR-WB":"21.4","STOI":"98.4"},"uses_additional_data":false,"paper_date":"2022-12-01","paper":"/paper/high-fidelity-speech-enhancement-with-band","paper_url":"https://arxiv.org/abs/2212.00406v2","paper_title":"High Fidelity Speech Enhancement with Band-split RNN","code":"https://github.com/sungwon23/bsrnn","n_code_links":1,"syntology":null},{"rank_in_archive_order":8,"model":"BSRNN-16k","metrics":{"PESQ-NB":"3.87","PESQ-WB":"3.45","SI-SDR-WB":"21.1","STOI":"98.3"},"uses_additional_data":false,"paper_date":"2022-12-01","paper":"/paper/high-fidelity-speech-enhancement-with-band","paper_url":"https://arxiv.org/abs/2212.00406v2","paper_title":"High Fidelity Speech Enhancement with Band-split RNN","code":"https://github.com/sungwon23/bsrnn","n_code_links":1,"syntology":null},{"rank_in_archive_order":9,"model":"MFNET","metrics":{"PESQ-NB":"3.74","PESQ-WB":"3.43","SI-SDR-WB":"20.31"},"uses_additional_data":false,"paper_date":"2023-06-07","paper":"/paper/a-mask-free-neural-network-for-monaural","paper_url":"https://arxiv.org/abs/2306.04286v1","paper_title":"A Mask Free Neural Network for Monaural Speech Enhancement","code":"https://github.com/ioyy900205/mfnet","n_code_links":2,"syntology":{"n_ran":2,"n_unverified":0,"n_samples":2,"n_pointer_only_licence":2}},{"rank_in_archive_order":10,"model":"BSRNN-S","metrics":{"PESQ-WB":"3.42","SI-SDR-WB":"21.3"},"uses_additional_data":false,"paper_date":"2022-12-01","paper":"/paper/high-fidelity-speech-enhancement-with-band","paper_url":"https://arxiv.org/abs/2212.00406v2","paper_title":"High Fidelity Speech Enhancement with Band-split RNN","code":"https://github.com/sungwon23/bsrnn","n_code_links":1,"syntology":null},{"rank_in_archive_order":11,"model":"BSRNN","metrics":{"PESQ-NB":"3.79","PESQ-WB":"3.32","STOI":"98"},"uses_additional_data":false,"paper_date":"2022-12-01","paper":"/paper/high-fidelity-speech-enhancement-with-band","paper_url":"https://arxiv.org/abs/2212.00406v2","paper_title":"High Fidelity Speech Enhancement with Band-split RNN","code":"https://github.com/sungwon23/bsrnn","n_code_links":1,"syntology":null},{"rank_in_archive_order":12,"model":"CleanUNet-2","metrics":{"PESQ-NB":"3.658","PESQ-WB":"3.262"},"uses_additional_data":false,"paper_date":"2023-09-12","paper":"/paper/cleanunet-2-a-hybrid-speech-denoising-model","paper_url":"https://arxiv.org/abs/2309.05975v1","paper_title":"CleanUNet 2: A Hybrid Speech Denoising Model on Waveform and Spectrogram","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":13,"model":"FRCRN","metrics":{"PESQ-WB":"3.23"},"uses_additional_data":false,"paper_date":"2021-02-03","paper":"/paper/monaural-speech-enhancement-with-complex","paper_url":"https://arxiv.org/abs/2102.01993v2","paper_title":"Monaural Speech Enhancement with Complex Convolutional Block Attention Module and Joint Time Frequency Losses","code":"https://github.com/modelscope/ClearerVoice-Studio","n_code_links":2,"syntology":null},{"rank_in_archive_order":14,"model":"FullSubNet+","metrics":{"PESQ-NB":"3.666","PESQ-WB":"3.218","SI-SDR-WB":"16.81"},"uses_additional_data":false,"paper_date":"2022-03-23","paper":"/paper/fullsubnet-channel-attention-fullsubnet-with","paper_url":"https://arxiv.org/abs/2203.12188v2","paper_title":"FullSubNet+: Channel Attention FullSubNet with Complex Spectrograms for Speech Enhancement","code":"https://github.com/thuhcsi/fullsubnet-plus","n_code_links":2,"syntology":{"n_ran":1,"n_unverified":0,"n_samples":1,"n_pointer_only_licence":0}},{"rank_in_archive_order":15,"model":"CleanUNet","metrics":{"PESQ-NB":"3.551","PESQ-WB":"3.146"},"uses_additional_data":false,"paper_date":"2022-02-15","paper":"/paper/speech-denoising-in-the-waveform-domain-with","paper_url":"https://arxiv.org/abs/2202.07790v3","paper_title":"Speech Denoising in the Waveform Domain with Self-Attention","code":"https://github.com/nvidia/cleanunet","n_code_links":1,"syntology":null},{"rank_in_archive_order":16,"model":"aTENNuate","metrics":{"PESQ-WB":"2.98"},"uses_additional_data":false,"paper_date":"2024-09-05","paper":"/paper/raw-speech-enhancement-with-deep-state-space","paper_url":"https://arxiv.org/abs/2409.03377v4","paper_title":"aTENNuate: Optimized Real-time Speech Enhancement with Deep SSMs on Raw Audio","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":17,"model":"Sudo rm -rf (U=32)","metrics":{"PESQ-WB":"2.95","SI-SDR-WB":"19.7"},"uses_additional_data":false,"paper_date":"2022-02-17","paper":"/paper/remixit-continual-self-training-of-speech","paper_url":"https://arxiv.org/abs/2202.08862v3","paper_title":"RemixIT: Continual self-training of speech enhancement models via bootstrapped remixing","code":"https://github.com/etzinis/unsup_speech_enh_adaptation","n_code_links":2,"syntology":{"n_ran":4,"n_unverified":0,"n_samples":4,"n_pointer_only_licence":0}},{"rank_in_archive_order":18,"model":"DCTCRN-P","metrics":{"PESQ-WB":"2.82"},"uses_additional_data":false,"paper_date":"2021-02-09","paper":"/paper/real-time-monaural-speech-enhancement-with","paper_url":"https://arxiv.org/abs/2102.04629","paper_title":"Real-time Monaural Speech Enhancement With Short-time Discrete Cosine Transform","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":19,"model":"DCTCRN-T","metrics":{"PESQ-WB":"2.82"},"uses_additional_data":false,"paper_date":"2021-02-09","paper":"/paper/real-time-monaural-speech-enhancement-with","paper_url":"https://arxiv.org/abs/2102.04629","paper_title":"Real-time Monaural Speech Enhancement With Short-time Discrete Cosine Transform","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":20,"model":"DCCRN-E","metrics":{"PESQ-WB":"2.79"},"uses_additional_data":false,"paper_date":"2021-02-09","paper":"/paper/real-time-monaural-speech-enhancement-with","paper_url":"https://arxiv.org/abs/2102.04629","paper_title":"Real-time Monaural Speech Enhancement With Short-time Discrete Cosine Transform","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":21,"model":"PoCoNet","metrics":{"PESQ-WB":"2.7885"},"uses_additional_data":false,"paper_date":"2020-08-11","paper":"/paper/poconet-better-speech-enhancement-with","paper_url":"https://arxiv.org/abs/2008.04470v1","paper_title":"PoCoNet: Better Speech Enhancement with Frequency-Positional Embeddings, Semi-Supervised Conversational Data, and Biased Loss","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":22,"model":"FullSubNet","metrics":{"PESQ-NB":"3.305","PESQ-WB":"2.777","SI-SDR-WB":"17.29"},"uses_additional_data":false,"paper_date":"2020-10-29","paper":"/paper/fullsubnet-a-full-band-and-sub-band-fusion","paper_url":"https://arxiv.org/abs/2010.15508v2","paper_title":"FullSubNet: A Full-Band and Sub-Band Fusion Model for Real-Time Single-Channel Speech Enhancement","code":"https://github.com/audio-westlakeu/fullsubnet","n_code_links":6,"syntology":{"n_ran":17,"n_unverified":4,"n_samples":21,"n_pointer_only_licence":0}},{"rank_in_archive_order":23,"model":"DCTCRN-S","metrics":{"PESQ-WB":"2.77"},"uses_additional_data":false,"paper_date":"2021-02-09","paper":"/paper/real-time-monaural-speech-enhancement-with","paper_url":"https://arxiv.org/abs/2102.04629","paper_title":"Real-time Monaural Speech Enhancement With Short-time Discrete Cosine Transform","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":24,"model":"RNN-Modulation","metrics":{"PESQ-WB":"2.75"},"uses_additional_data":false,"paper_date":"2021-02-15","paper":"/paper/a-modulation-domain-loss-for-neural-network-1","paper_url":"https://arxiv.org/abs/2102.07330v1","paper_title":"A Modulation-Domain Loss for Neural-Network-based Real-time Speech Enhancement","code":"https://github.com/tvuong123/ModulationDomainLoss","n_code_links":1,"syntology":null},{"rank_in_archive_order":25,"model":"Conv-TasNet-SNR","metrics":{"PESQ-WB":"2.73"},"uses_additional_data":false,"paper_date":"2020-08-20","paper":"/paper/exploring-the-best-loss-function-for-dnn-1","paper_url":"http://arxiv.org/abs/2005.11611v3","paper_title":"Exploring the Best Loss Function for DNN-Based Low-latency Speech Enhancement with Temporal Convolutional Networks","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":26,"model":"Sudo rm-rf (U=8)","metrics":{"PESQ-WB":"2.69","SI-SDR-WB":"18.6"},"uses_additional_data":false,"paper_date":"2021-10-19","paper":"/paper/continual-self-training-with-bootstrapped","paper_url":"https://arxiv.org/abs/2110.10103v2","paper_title":"Continual self-training with bootstrapped remixing for speech enhancement","code":"https://github.com/etzinis/unsup_speech_enh_adaptation","n_code_links":1,"syntology":null},{"rank_in_archive_order":27,"model":"Proposed (0.35)","metrics":{"PESQ-NB":"2.65","PESQ-WB":"2.65"},"uses_additional_data":false,"paper_date":"2020-02-12","paper":"/paper/weighted-speech-distortion-losses-for-neural-1","paper_url":"http://arxiv.org/abs/2001.10601v2","paper_title":"Weighted Speech Distortion Losses for Neural-network-based Real-time Speech Enhancement","code":"https://github.com/microsoft/DNS-Challenge","n_code_links":3,"syntology":null},{"rank_in_archive_order":28,"model":"RemixIT (w Sudo U=32)","metrics":{"PESQ-WB":"2.60","SI-SDR-WB":"18.0"},"uses_additional_data":true,"paper_date":"2021-10-19","paper":"/paper/continual-self-training-with-bootstrapped","paper_url":"https://arxiv.org/abs/2110.10103v2","paper_title":"Continual self-training with bootstrapped remixing for speech enhancement","code":"https://github.com/etzinis/unsup_speech_enh_adaptation","n_code_links":1,"syntology":null},{"rank_in_archive_order":29,"model":"RemixIT (w Sudo U=32)","metrics":{"PESQ-WB":"2.34","SI-SDR-WB":"16.0"},"uses_additional_data":true,"paper_date":"2022-02-17","paper":"/paper/remixit-continual-self-training-of-speech","paper_url":"https://arxiv.org/abs/2202.08862v3","paper_title":"RemixIT: Continual self-training of speech enhancement models via bootstrapped remixing","code":"https://github.com/etzinis/unsup_speech_enh_adaptation","n_code_links":2,"syntology":{"n_ran":4,"n_unverified":0,"n_samples":4,"n_pointer_only_licence":0}},{"rank_in_archive_order":30,"model":"Noisy","metrics":{"PESQ-WB":"1.58","SI-SDR-WB":"9.1","STOI":"91.5"},"uses_additional_data":false,"paper_date":null,"paper":null,"paper_url":null,"paper_title":"","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":31,"model":"BSRNN-S + MGD","metrics":{"PESQ-NB":"3.85","SI-SDR-WB":"21.4","STOI":"98.4"},"uses_additional_data":false,"paper_date":"2022-12-01","paper":"/paper/high-fidelity-speech-enhancement-with-band","paper_url":"https://arxiv.org/abs/2212.00406v2","paper_title":"High Fidelity Speech Enhancement with Band-split RNN","code":"https://github.com/sungwon23/bsrnn","n_code_links":1,"syntology":null},{"rank_in_archive_order":32,"model":"DTLN","metrics":{"PESQ-NB":"3.04","SI-SDR-WB":"16.34"},"uses_additional_data":false,"paper_date":"2020-05-15","paper":"/paper/dual-signal-transformation-lstm-network-for","paper_url":"https://arxiv.org/abs/2005.07551v1","paper_title":"Dual-Signal Transformation LSTM Network for Real-Time Noise Suppression","code":"https://github.com/breizhn/DTLN","n_code_links":2,"syntology":{"n_ran":3,"n_unverified":1,"n_samples":4,"n_pointer_only_licence":0}},{"rank_in_archive_order":33,"model":"Non-Real-Time MultiScale+","metrics":{"PESQ-NB":"3.01","SI-SDR-WB":"16.22"},"uses_additional_data":false,"paper_date":"2020-06-01","paper":"/paper/phase-aware-single-stage-speech-denoising-and-1","paper_url":"https://arxiv.org/abs/2006.00687v1","paper_title":"Phase-aware Single-stage Speech Denoising and Dereverberation with U-Net","code":"https://github.com/jonashaag/paperswithcode-speech-enhancement-audiosamples/tree/master/Phase-aware%20Single-stage%20Speech%20Denoising%20and%20Dereverberation%20with%20U-Net","n_code_links":1,"syntology":null},{"rank_in_archive_order":34,"model":"SN-Net","metrics":{"PESQ-NB":"3.39","SI-SDR-NB":"19.52"},"uses_additional_data":false,"paper_date":"2020-12-17","paper":"/paper/interactive-speech-and-noise-modeling-for-1","paper_url":"https://arxiv.org/abs/2012.09408v2","paper_title":"Interactive Speech and Noise Modeling for Speech Enhancement","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":35,"model":"DCCRN-E-Aug","metrics":{"PESQ-NB":"3.214"},"uses_additional_data":false,"paper_date":"2020-08-01","paper":"/paper/dccrn-deep-complex-convolution-recurrent-1","paper_url":"https://arxiv.org/abs/2008.00264v4","paper_title":"DCCRN: Deep Complex Convolution Recurrent Network for Phase-Aware Speech Enhancement","code":"https://github.com/mpariente/asteroid","n_code_links":7,"syntology":{"n_ran":8,"n_unverified":7,"n_samples":15,"n_pointer_only_licence":0}},{"rank_in_archive_order":36,"model":"DCCRN-E","metrics":{"PESQ-NB":"3.04"},"uses_additional_data":false,"paper_date":"2020-08-01","paper":"/paper/dccrn-deep-complex-convolution-recurrent-1","paper_url":"https://arxiv.org/abs/2008.00264v4","paper_title":"DCCRN: Deep Complex Convolution Recurrent Network for Phase-Aware Speech Enhancement","code":"https://github.com/mpariente/asteroid","n_code_links":7,"syntology":{"n_ran":8,"n_unverified":7,"n_samples":15,"n_pointer_only_licence":0}}],"since_archive":{"claim":"Results that newer papers report for their own method, placed here by Syntology. A model pointed at the cell in the paper's own table; the number was read from that cell and checked against this leaderboard's metric, dataset, split and scale; an independent check that saw this leaderboard's other rows and every other leaderboard on the same dataset accepted it. Not reviewed by the paper's authors or by the archive's editors, and not ranked against the archive rows.","extraction_file_present":true,"measurement":{"test_papers":883,"papers_with_output":881,"judged_true":108,"judged":110,"wilson95_lower":0.9361,"measured_on":"2026-09-24","frozen_commit":"0e3de0df94"},"measurement_note":"blind adjudication of accepted entries on a held-out split of archive papers, rules frozen before the test","coverage":{"sentence":"Syntology has checked 6,821 of the 9,623 papers on this site that are newer than the archive; results from the others appear after they are checked.","complete":false,"papers_newer_than_archive":9623,"papers_checked":6821,"papers_extracted_not_yet_verified":65,"boards_without_verdict":29,"papers_not_yet_extracted":2737},"order":"newest first by month (arXiv date, else the arXiv-id month), then arXiv id descending","columns":[],"entries":[]},"syntology":{"read_at":"2026-09-25T09:33:49+00:00","claim":"Per row: N of M harvested code samples from that row's paper executed on a synthesized fixture; the other M-N are unverified. Not a reproduction of the row's number; not a correctness claim. n_pointer_only_licence counts samples the site points at rather than redistributes (a licence axis, independent of ran/unverified).","rows_with_graph_line":9,"rows_with_any_sample_ran":9,"distinct_papers_with_graph_line":7,"distinct_papers_with_any_sample_ran":7,"samples_over_distinct_papers":{"n_ran":48,"n_unverified":13,"n_samples":61,"n_pointer_only_licence":2,"note":"each paper (arXiv id) counted once, however many rows it is behind; this is the page-level figure"},"samples_row_weighted":{"n_ran":60,"n_unverified":20,"n_samples":80,"n_pointer_only_licence":2,"note":"row-weighted: a paper behind several rows is counted once per row; inflated relative to samples_over_distinct_papers by design, kept for readers summing the per-row syntology blocks"}}}