Browse State-of-the-Art › Target Sound Extraction
Target Sound Extraction
8 papers with code · 3 benchmarks · 3 datasets archive 2025-07-28
Target Sound Extraction is the task of extracting a sound corresponding to a given class from an audio mixture. The audio mixture may contain background noise with a relatively low amplitude compared to the foreground mixture components. The choice of the sound class is provided as input to the model in form of a string, integer, or a one-hot encoding of the sound class.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
3 leaderboard tables shown for this task, 3 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| AudioCaps (1 row) | CLAPSep | CLAPSep: Leveraging Contrastive Pre-trained Model for Multi-Modal... | code | Syntology ran 4 of 4 samples · 0 unverified | Compare |
| AudioSet (1 row) | CLAPSep | CLAPSep: Leveraging Contrastive Pre-trained Model for Multi-Modal... | code | Syntology ran 4 of 4 samples · 0 unverified | Compare |
| FSDSoundScapes (1 row) | Waveformer | Real-Time Target Sound Extraction | code | Syntology ran 3 of 7 samples · 4 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
3 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
8 shown of 8 papers with code (16 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
6 Oct 2023 2 repositories listedCommon target sound extraction (TSE) approaches primarily relied on discriminative approaches in order to separate the target sound while minimizing interference from the unwanted sources, with varying success in…
-
12 Sep 2024 1 repository listedIn this paper, we introduce SoloAudio, a novel diffusion-based generative model for target sound extraction (TSE).
-
7 Sep 2024 1 repository listed Syntology ran 2 of 4 samples · 2 unverified · 4 pointer-only (licence)In this paper, we propose CrossMamba for target sound extraction, which leverages the hidden attention mechanism of Mamba to compute dependencies between the given clues and the audio mixture.
-
22 Jul 2024 1 repository listedThe experimental results using the CHiME-4 dataset suggested that 1) all 12 variations can achieve the theoretical upper-bound performance, and 2) mask-based scaling can behave like IS, even when constraining the…
-
27 Feb 2024 1 repository listed Syntology ran 4 of 4 samples · 0 unverifiedUniversal sound separation (USS) aims to extract arbitrary types of sounds from real-world recordings.
-
1 Nov 2023 1 repository listedTo achieve this, we make two technical contributions: 1) we present the first neural network that can achieve binaural target sound extraction in the presence of interfering sounds and background noise, and 2) we design…
-
15 Mar 2023 1 repository listedAutomatic target sound extraction (TSE) is a machine learning approach to mimic the human auditory perception capability of attending to a sound source of interest from a mixture of sources.
-
4 Nov 2022 1 repository listed Syntology ran 3 of 7 samples · 4 unverifiedWe present the first neural network model to achieve real-time and streaming target sound extraction.
Syntology lines on 3 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections