Browse State-of-the-Art › Speech Enhancement
Speech Enhancement
280 papers with code · 17 benchmarks · 24 datasets archive 2025-07-28
Speech Enhancement is a signal processing task that involves improving the quality of speech signals captured under noisy or degraded conditions. The goal of speech enhancement is to make speech signals clearer, more intelligible, and more pleasant to listen to, which can be used for various applications such as voice recognition, teleconferencing, and hearing aids. A representative Github project with online demo : ClearerVoice-Studio.
( Image credit: A Fully Convolutional Neural Network For Speech Enhancement )
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
18 leaderboard tables shown for this task, 17 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 18 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
24 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
4 subtasks in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 280 papers with code (982 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
20 Jul 2017 188 repositories listed Syntology ran 99 of 176 samples · 77 unverified · 94 pointer-only (licence)We propose a new family of policy gradient methods for reinforcement learning, which alternate between sampling data through interaction with the environment, and optimizing a "surrogate" objective function using…
-
27 Mar 2016 80 repositories listed Syntology ran 13 of 46 samples · 33 unverified · 5 pointer-only (licence)We consider image transformation problems, where an input image is transformed into an output image.
-
28 Mar 2017 21 repositories listed Syntology ran 3 of 39 samples · 36 unverified · 4 pointer-only (licence)In contrast to current techniques, we operate at the waveform level, training the model end-to-end, and incorporate 28 speakers and 40 different noise conditions into the same model, such that model parameters are…
-
20 Sep 2018 17 repositories listed Syntology ran 7 of 36 samples · 29 unverified · 16 pointer-only (licence)The majority of the previous methods have formulated the separation problem through the time-frequency representation of the mixed signal, which has several drawbacks, including the decoupling of the phase and magnitude…
-
19 Jan 2020 13 repositories listed Syntology ran 1 of 16 samples · 15 unverifiedThe paper presents a novel method, Zero-Reference Deep Curve Estimation (Zero-DCE), which formulates light enhancement as a task of image-specific curve estimation with a deep network.
-
7 Mar 2019 9 repositories listed Syntology ran 3 of 10 samples · 7 unverified · 2 pointer-only (licence)Most deep learning-based models for speech enhancement have mainly focused on estimating the magnitude of spectrogram while reusing the phase from noisy speech for reconstruction.
-
14 Oct 2021 6 repositories listed Syntology ran 2 of 3 samples · 1 unverified · 3 pointer-only (licence)Motivated by the success of T5 (Text-To-Text Transfer Transformer) in pre-trained natural language processing models, we propose a unified-modal SpeechT5 framework that explores the encoder-decoder pre-training for…
-
7 Jul 2021 6 repositories listed Syntology ran 4 of 4 samples · 0 unverified · 3 pointer-only (licence)We present SoundStream, a novel neural audio codec that can efficiently compress speech, music and general audio at bitrates normally targeted by speech-tailored codecs.
-
29 Oct 2020 6 repositories listed Syntology ran 1 of 21 samples · 20 unverifiedIn our proposed FullSubNet, we connect a pure full-band model and a pure sub-band model sequentially and use practical joint training to integrate these two types of models' advantages.
-
22 Sep 2016 6 repositories listedIn hearing aids, the presence of babble noise degrades hearing intelligibility of human speech greatly.
-
6 Apr 2020 5 repositories listedIn WaveCRN, the speech locality feature is captured by a convolutional neural network (CNN), while the temporal sequential property of the locality feature is modeled by stacked simple recurrent units (SRU).
-
13 May 2019 5 repositories listedAdversarial loss in a conditional generative adversarial network (GAN) is not designed to directly optimize evaluation metrics of a target task, and thus, may not always guide the generator in a GAN to generate data…
-
11 Oct 2018 5 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)In this paper, we present a novel system that separates the voice of a target speaker from multi-speaker signals, by making use of a reference signal from the target speaker.
-
25 May 2021 4 repositories listedRecent interest in exploiting Deep Learning techniques for Noise Suppression, has led to the creation of Hybrid Denoising Systems that combine classic Signal Processing with Deep Learning.
-
31 Aug 2018 4 repositories listedMost methods of voice restoration for patients suffering from aphonia either produce whispered or monotone speech.
-
24 Mar 2022 3 repositories listed Syntology ran 9 of 12 samples · 3 unverified · 1 pointer-only (licence)Generative adversarial networks have recently demonstrated outstanding performance in neural vocoding outperforming best autoregressive and flow-based models.
-
8 Apr 2021 3 repositories listed Syntology ran 0 of 11 samples · 11 unverifiedThe discrepancy between the cost function used for training a speech enhancement model and human auditory perception usually makes the quality of enhanced speech unsatisfactory.
-
23 Jun 2020 3 repositories listed Syntology ran 1 of 2 samples · 1 unverified · 2 pointer-only (licence)The proposed model matches state-of-the-art performance of both causal and non causal methods while working directly on the raw waveform.
-
12 Feb 2020 3 repositories listedThis paper investigates several aspects of training a RNN (recurrent neural network) that impact the objective and subjective quality of enhanced speech for real-time single-channel speech enhancement.
-
4 Nov 2019 3 repositories listedWe present and release a new tool for music source separation with pre-trained models called Spleeter.
-
9 Jun 2019 3 repositories listedIn the end, a posteriori SNR weighted energy difference is applied to the extended pitch segments of the denoised speech signal for detecting voice activity.
-
8 Mar 2019 3 repositories listedA popular approach to multichannel source separation is to integrate a spatial model with a source model for estimating the spatial covariance matrices (SCMs) and power spectral densities (PSDs) of each sound source in…
-
27 Nov 2018 3 repositories listedWe study the use of the Wave-U-Net architecture for speech enhancement, a model introduced by Stoller et al for the separation of music vocals and accompaniment.
-
18 Dec 2017 3 repositories listedIn this work, we present the results of adapting a speech enhancement generative adversarial network by finetuning the generator with small amounts of data.
-
1 Jul 2025 2 repositories listedNevertheless, neither the concept of hybrid Mamba and time-frequency attention models nor their generalization performance have been explored for speech enhancement.
-
2 Oct 2024 2 repositories listed Syntology ran 8 of 11 samples · 3 unverified · 11 pointer-only (licence)Additionally, to investigate the differences between synthetic and real-world data, we selected 5 hours of raw, non-reverberant data from the SonicSet validation set and recorded a real-world speech separation dataset,…
-
30 Aug 2024 2 repositories listedConvolutional layers with 1-D filters are often used as frontend to encode audio signals.
-
10 Jun 2024 2 repositories listed Syntology ran 7 of 8 samples · 1 unverifiedWe release the EARS (Expressive Anechoic Recordings of Speech) dataset, a high-quality speech dataset comprising 107 speakers from diverse backgrounds, totaling in 100 hours of clean, anechoic speech data.
-
7 Oct 2023 2 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedPrevious mainstream audio-and-text LLMs use discrete audio tokens to represent both input and output audio; however, they suffer from performance degradation on tasks such as automatic speech recognition, speech-to-text…
-
14 Sep 2023 2 repositories listedThe commonly used standard ITU-T Rec.
Syntology lines on 16 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections