Browse State-of-the-Art › Audio Source Separation
Audio Source Separation
54 papers with code · 2 benchmarks · 15 datasets archive 2025-07-28
Audio Source Separation is the process of separating a mixture (e.g. a pop band recording) into isolated sounds from individual sources (e.g. just the lead vocals).
Source: Model selection for deep audio source separation via clustering analysis
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
2 leaderboard tables shown for this task, 2 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| AudioSet (2 rows) | ST-SED-SEP | Zero-shot Audio Source Separation through Query-based Learning... | code | Syntology ran 3 of 10 samples · 7 unverified | Compare |
| MUSIC (multi-source) (1 row) | Co-Separation | Co-Separating Sounds of Visual Objects | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
15 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
3 subtasks in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 54 papers with code (112 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
8 Jun 2018 10 repositories listed Syntology ran 1 of 19 samples · 18 unverified · 1 pointer-only (licence)Models for audio source separation usually operate on the magnitude spectrum, which ignores phase information and makes separation performance dependant on hyper-parameters for the spectral front-end.
-
29 Jun 2017 5 repositories listedThis paper deals with the problem of audio source separation.
-
14 Jul 2020 4 repositories listed Syntology ran 1 of 10 samples · 9 unverifiedIn this paper, we present an efficient neural network for end-to-end general purpose audio source separation.
-
19 Oct 2021 3 repositories listed Syntology ran 0 of 4 samples · 4 unverifiedThe cocktail party problem aims at isolating any source of interest within a complex acoustic scene, and has long inspired audio source separation research.
-
3 Mar 2021 3 repositories listedRecent progress in audio source separation lead by deep learning has enabled many neural network models to provide robust solutions to this fundamental estimation problem.
-
16 Apr 2019 3 repositories listedLearning how objects sound from video is challenging, since they often heavily overlap in a single audio channel.
-
27 Nov 2018 3 repositories listedWe study the use of the Wave-U-Net architecture for speech enhancement, a model introduced by Stoller et al for the separation of music vocals and accompaniment.
-
31 Oct 2017 3 repositories listed Syntology ran 0 of 13 samples · 13 unverifiedBased on this idea, we drive the separator towards outputs deemed as realistic by discriminator networks that are trained to tell apart real from separator samples.
-
24 Jan 2022 2 repositories listedIntegrating domain knowledge in the form of source models into a data-driven method leads to high data efficiency: the proposed approach achieves good separation quality even when trained on less than three minutes of…
-
30 Jan 2021 2 repositories listedIn blind source separation of speech signals, the inherent imbalance in the source spectrum poses a challenge for methods that rely on single-source dominance for the estimation of the mixing matrix.
-
2 Jul 2019 2 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)The input vector is embedded to obtain the parameters that control Feature-wise Linear Modulation (FiLM) layers.
-
5 Apr 2018 2 repositories listedOur work is the first to learn audio source separation from large-scale "in the wild" videos containing multiple audio sources per video.
-
15 Jul 2025 1 repository listedTraditional Blind Source Separation Evaluation (BSS-Eval) metrics were originally designed to evaluate linear audio source separation models based on methods such as time-frequency masking.
-
26 May 2025 1 repository listedOur empirical results demonstrate that our multi-step separation approach consistently outperforms one-step inference across both speech enhancement and music source separation tasks, and can achieve scaling performance…
-
2 Nov 2024 1 repository listedSound Event Detection (SED) is challenging in noisy environments where overlapping sounds obscure target events.
-
20 Sep 2024 1 repository listed Syntology ran 8 of 8 samples · 0 unverifiedTo address this issue, we propose the text-queried SED (TQ-SED) framework.
-
19 Aug 2024 1 repository listedWe demonstrate that our framework, used with diffusion models, naturally addresses the task of unsupervised audio source separation, showing that our model is able to perform high-quality separation.
-
7 Aug 2024 1 repository listedIn this work, we demonstrate a very straightforward extension of the dedicated-decoder Bandit and query-based single-decoder Banquet models to a four-stem problem, treating non-musical dialogue, instrumental music,…
-
9 Jul 2024 1 repository listedCinematic audio source separation (CASS), as a problem of extracting the dialogue, music, and effects stems from their mixture, is a relatively new subtask of audio source separation.
-
26 Jun 2024 1 repository listed Syntology ran 4 of 4 samples · 0 unverifiedOf the very few current systems that support source separation beyond this setup, most continue to rely on an inflexible decoder setup that can only support a fixed pre-defined set of stems.
-
30 May 2024 1 repository listedThe combination of frequency-axis normalization with Min/Max scaling and the Mean Absolute Error (MAE) loss function achieved the highest Source-to-Distortion Ratio (SDR) of 7.
-
5 Sep 2023 1 repository listedCinematic audio source separation is a relatively new subtask of audio source separation, with the aim of extracting the dialogue, music, and effects stems from their mixture.
-
9 Aug 2023 1 repository listed Syntology ran 8 of 9 samples · 1 unverifiedIn this work, we introduce AudioSep, a foundation model for open-domain audio source separation with natural language queries.
-
29 Oct 2022 1 repository listed Syntology ran 1 of 2 samples · 1 unverifiedIn this paper, we propose to use this connection between audio and visual dynamics for solving two challenging tasks simultaneously, namely: (i) separating audio sources from a mixture using visual cues, and (ii)…
-
21 Jul 2022 1 repository listedA network with relevant deep priors is likely to generate a cleaner version of the signal before converging on the corrupted signal.
-
28 Mar 2022 1 repository listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)In this paper, we introduce the task of language-queried audio source separation (LASS), which aims to separate a target source from an audio mixture based on a natural language query of the target source (e.
-
15 Dec 2021 1 repository listed Syntology ran 3 of 10 samples · 7 unverifiedOur approach uses a single model for source separation of multiple sound types, and relies solely on weakly-labeled data for training.
-
11 Dec 2021 1 repository listedOn-device directional hearing requires audio source separation from a given direction while achieving stringent human-imperceptible latency requirements.
-
28 Nov 2021 1 repository listedIn this work, we demonstrate how a publicly available, pre-trained Jukebox model can be adapted for the problem of audio source separation from a single mixed audio channel.
-
25 Oct 2021 1 repository listedWe showcase an unsupervised method that repurposes deep models trained for music generation and music tagging for audio source separation, without any retraining.
Syntology lines on 11 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections