Browse State-of-the-Art › Speech Separation
Speech Separation
120 papers with code · 19 benchmarks · 16 datasets archive 2025-07-28
The task of extracting all overlapping speech sources in a given mixed speech signal refers to the Speech Separation. Speech Separation is a special scenario of source separation problem, where the focus is only on the overlapping speech signal sources and other interferences such as music or noise signals are not the main concern of the study. A recent representative Github project can be referred to ClearerVoice-Studio.
Source: A Unified Framework for Speech Separation
Image credit: Speech Separation of A Target Speaker Based on Deep Neural Networks
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
19 leaderboard tables shown for this task, 19 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 19 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
16 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 120 papers with code (359 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
20 Sep 2018 17 repositories listed Syntology ran 7 of 36 samples · 29 unverified · 16 pointer-only (licence)The majority of the previous methods have formulated the separation problem through the time-frequency representation of the mixed signal, which has several drawbacks, including the decoupling of the phase and magnitude…
-
14 Oct 2019 8 repositories listed Syntology ran 4 of 22 samples · 18 unverifiedRecent studies in deep learning-based speech separation have proven the superiority of time-domain approaches to conventional time-frequency-based methods.
-
18 Aug 2015 8 repositories listed Syntology ran 2 of 3 samples · 1 unverified · 3 pointer-only (licence)The framework can be used without class labels, and therefore has the potential to be trained on a diverse set of sound types, and to generalize to novel sources.
-
6 Apr 2020 5 repositories listedIn WaveCRN, the speech locality feature is captured by a convolutional neural network (CNN), while the temporal sequential property of the locality feature is modeled by stacked simple recurrent units (SRU).
-
13 Jul 2024 4 repositories listedIt is too early to conclude that Mamba is a better alternative to transformers for speech before comparing Mamba with transformers in terms of both performance and efficiency in multiple speech-related tasks.
-
25 Oct 2020 4 repositories listedTransformers are emerging as a natural alternative to standard RNNs, replacing recurrent computations with a multi-head attention mechanism.
-
14 Jul 2020 4 repositories listed Syntology ran 1 of 10 samples · 9 unverifiedIn this paper, we present an efficient neural network for end-to-end general purpose audio source separation.
-
29 Feb 2020 4 repositories listed Syntology ran 10 of 13 samples · 3 unverified · 9 pointer-only (licence)We present a new method for separating a mixed audio sequence, in which multiple voices speak simultaneously.
-
10 Apr 2018 4 repositories listed Syntology ran 0 of 23 samples · 23 unverifiedSolving this task using only audio as input is extremely challenging and does not provide an association of the separated speech signals with speakers in the video.
-
3 Mar 2021 3 repositories listedRecent progress in audio source separation lead by deep learning has enabled many neural network models to provide robust solutions to this fundamental estimation problem.
-
1 Nov 2017 3 repositories listedWe directly model the signal in the time-domain using an encoder-decoder framework and perform the source separation on nonnegative encoder outputs.
-
18 Mar 2017 3 repositories listed Syntology ran 1 of 3 samples · 2 unverified · 3 pointer-only (licence)We evaluated uPIT on the WSJ0 and Danish two- and three-talker mixed-speech separation tasks and found that uPIT outperforms techniques based on Non-negative Matrix Factorization (NMF) and Computational Auditory Scene…
-
2 Oct 2024 2 repositories listed Syntology ran 8 of 11 samples · 3 unverified · 11 pointer-only (licence)Additionally, to investigate the differences between synthetic and real-world data, we selected 5 hours of raw, non-reverberant data from the SonicSet validation set and recorded a real-world speech separation dataset,…
-
23 Feb 2023 2 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)To effectively solve the indirect elemental interactions across chunks in the dual-path architecture, MossFormer employs a joint local and global self-attention architecture that simultaneously performs a…
-
21 Dec 2022 2 repositories listed Syntology ran 2 of 3 samples · 1 unverifiedThen, inspired by the large number of connections between cortical regions and the thalamus, the model fuses the auditory and visual information in a thalamic subnetwork through top-down connections.
-
27 Oct 2022 2 repositories listedIn this work deformable convolution is proposed as a solution to allow TCN models to have dynamic RFs that can adapt to various reverberation times for reverberant speech separation.
-
22 Sep 2022 2 repositories listedRather than focusing exclusively on the speech denoising task, we extend this work to address the dereverberation and super-resolution tasks.
-
18 May 2022 2 repositories listed Syntology ran 6 of 6 samples · 0 unverified · 6 pointer-only (licence)We show that the acoustic metrics of the IRs predicted from our MESH2IR match the ground truth with less than 10% error.
-
20 Apr 2021 2 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)The task of isolating a target singing voice in music videos has useful applications.
-
1 Mar 2021 2 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)One of the leading single-channel speech separation (SS) models is based on a TasNet with a dual-path segmentation technique, where the size of each segment remains unchanged throughout all layers.
-
30 Jan 2021 2 repositories listedIn blind source separation of speech signals, the inherent imbalance in the source spectrum poses a challenge for methods that rely on single-source dominance for the estimation of the mixing matrix.
-
13 Jan 2021 2 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Recent research on the time-domain audio separation networks (TasNets) has brought great success to speech separation.
-
24 Nov 2020 2 repositories listedBeyond the model, we also propose a metric on how to evaluate source separation with variable number of speakers.
-
4 Oct 2020 2 repositories listed Syntology ran 1 of 9 samples · 8 unverifiedAlthough our system is trained on simulated room impulse responses (RIR) based on a fixed number of microphones arranged in a given geometry, it generalizes well to a real array with the same geometry.
-
30 Oct 2019 2 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)An important problem in ad-hoc microphone speech separation is how to guarantee the robustness of a system with respect to the locations and numbers of microphones.
-
23 Oct 2019 2 repositories listed Syntology ran 1 of 8 samples · 7 unverifiedAlso, we validate the use of parameterized filterbanks and show that complex-valued representations and masks are beneficial in all conditions.
-
22 Oct 2019 2 repositories listedIn the first step we learn a transform (and it's inverse) to a latent space where masking-based separation performance using oracles is optimal.
-
7 Jul 2016 2 repositories listedIn this paper we extend the baseline system with an end-to-end signal approximation objective that greatly improves performance on a challenging speech separation.
-
13 Feb 2015 2 repositories listedIn this paper, we explore joint optimization of masking functions and deep recurrent neural networks for monaural source separation tasks, including monaural speech separation, monaural singing voice separation, and…
-
25 May 2025 1 repository listedOn the other hand, generative models for TSE lag in perceptual quality and intelligibility.
Syntology lines on 17 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections