Browse State-of-the-Art › Sound Event Detection
Sound Event Detection
92 papers with code · 5 benchmarks · 20 datasets archive 2025-07-28
Sound Event Detection (SED) is the task of recognizing the sound events and their respective temporal start and end time in a recording. Sound events in real life do not always occur in isolation, but tend to considerably overlap with each other. Recognizing such overlapping sound events is referred as polyphonic SED.
Source: A report on sound event detection with different binaural features
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
5 leaderboard tables shown for this task, 5 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| DESED (13 rows) | ATST-SED | Fine-tune the pretrained ATST model for sound event detection | code | Syntology ran 10 of 10 samples · 0 unverified | Compare |
| L3DAS21 (5 rows) | PHC SEDnet n=2 | PHNNs: Lightweight Neural Networks via Parameterized Hypercomplex... | code | Syntology ran 5 of 34 samples · 29 unverified | Compare |
| WildDESED (5 rows) | CRNN (with BEATs + Separation) | Leveraging LLM and Text-Queried Separation for Noise-Robust Sound... | code | — | Compare |
| Mivia Audio Events (1 row) | DENet | DENet: a deep architecture for audio surveillance applications | code | — | Compare |
| Mivia Road Events (1 row) | DENet | DENet: a deep architecture for audio surveillance applications | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
20 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 92 papers with code (194 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
19 Jun 2017 59 repositories listed Syntology ran 9 of 17 samples · 8 unverified · 13 pointer-only (licence)Its principled nature also enables us to identify methods for both training and attacking neural networks that are reliable and, in a certain sense, universal.
-
8 Oct 2021 4 repositories listed Syntology ran 5 of 34 samples · 29 unverifiedIn this paper, we define the parameterization of hypercomplex convolutional layers and introduce the family of parameterized hypercomplex neural networks (PHNNs) that are lightweight and efficient large-scale models.
-
30 Mar 2023 3 repositories listed Syntology ran 0 of 12 samples · 12 unverifiedTo address this data scarcity issue, we introduce WavCaps, the first large-scale weakly-labelled audio captioning dataset, comprising approximately 400k audio clips with paired captions.
-
14 Sep 2024 2 repositories listedFor five transformers, we obtain a substantial performance improvement over previously available checkpoints both on AudioSet frame-level predictions and on frame-level sound event detection downstream tasks, confirming…
-
27 Mar 2024 2 repositories listedA recent development in the field is the introduction of the task known as few-shot bioacoustic sound event detection, which aims to train a versatile animal sound detector using only a small set of audio samples.
-
5 Oct 2023 2 repositories listedIn recent years, deep learning systems have shown a concerning trend toward increased complexity and higher energy consumption.
-
7 Jun 2023 2 repositories listedIn order to tackle both clip-level and frame-level tasks, this paper proposes Audio Teacher-Student Transformer (ATST), with a clip-level version (named ATST-Clip) and a frame-level version (named ATST-Frame),…
-
5 Sep 2022 2 repositories listedOur system submitted to the DCASE 2022 Task 3 is based on our previous proposed Event-Independent Network V2 (EINV2) with a novel data augmentation method.
-
21 Oct 2021 2 repositories listedSound event detection (SED), as a core module of acoustic environmental analysis, suffers from the problem of data deficiency.
-
12 Oct 2021 2 repositories listedThe recently proposed Mean Teacher method, which exploits large-scale unlabeled data in a self-ensembling manner, has achieved state-of-the-art results in several semi-supervised learning benchmarks.
-
29 Oct 2020 2 repositories listedConventional NN-based methods use two branches for a sound event detection (SED) target and a direction-of-arrival (DOA) target.
-
3 Mar 2020 2 repositories listedThe understanding of the surrounding environment plays a critical role in autonomous robotic systems, such as self-driving cars.
-
4 Jan 2019 2 repositories listed Syntology ran 0 of 9 samples · 9 unverifiedTo foster the investigation of label noise in sound event classification we present FSDnoisy18k, a dataset containing 42.
-
26 Apr 2018 2 repositories listedIn this work, we treat SED as a multiple instance learning (MIL) problem, where training labels are static over a short excerpt, indicating the presence or absence of sound sources but not their temporal locality.
-
4 Apr 2016 2 repositories listedIn this paper we present an approach to polyphonic sound event detection in real life recordings based on bi-directional long short term memory (BLSTM) recurrent neural networks (RNNs).
-
27 May 2025 1 repository listedBioacoustic sound event detection (BioSED) is crucial for biodiversity conservation but faces practical challenges during model development and training: limited amounts of annotated data, sparse events, species…
-
17 Apr 2025 1 repository listedTo address this limitation, we propose temporal attention pooling frequency dynamic convolution (TFD conv) to replace temporal average pooling with temporal attention pooling (TAP).
-
14 Mar 2025 1 repository listedWe target the problem of developing new low-complexity networks for the sound event detection task.
-
28 Feb 2025 1 repository listedSound event detection (SED) has significantly benefited from self-supervised learning (SSL) approaches, particularly masked audio transformer for SED (MAT-SED), which leverages masked block prediction to reconstruct…
-
2 Nov 2024 1 repository listedSound Event Detection (SED) is challenging in noisy environments where overlapping sounds obscure target events.
-
26 Sep 2024 1 repository listedA significant challenge in sound event detection (SED) is the effective utilization of unlabeled data, given the limited availability of labeled data due to high annotation costs.
-
20 Sep 2024 1 repository listed Syntology ran 8 of 8 samples · 0 unverifiedTo address this issue, we propose the text-queried SED (TQ-SED) framework.
-
17 Sep 2024 1 repository listedThis paper presents a residential audio dataset to support sound event detection research for smart home applications aimed at promoting wellbeing for older adults.
-
10 Sep 2024 1 repository listedSound Event Detection (SED) plays a vital role in comprehending and perceiving acoustic scenes.
-
16 Aug 2024 1 repository listedSound event detection (SED) methods that leverage a large pre-trained Transformer encoder network have shown promising performance in recent DCASE challenges.
-
17 Jul 2024 1 repository listedA single model and an ensemble, both based on our proposed training procedure, ranked first in Task 4 of the DCASE Challenge 2024.
-
17 Jul 2024 1 repository listed Syntology ran 5 of 5 samples · 0 unverified · 5 pointer-only (licence)We fine-tune three large Audio Spectrogram Transformers, PaSST, BEATs, and ATST, on the joint DESED and MAESTRO datasets in a two-stage training procedure.
-
4 Jul 2024 1 repository listedThis work aims to advance sound event detection (SED) research by presenting a new large language model (LLM)-powered dataset namely wild domestic environment sound event detection (WildDESED).
-
22 Jun 2024 1 repository listedUsing best ensemble model, we applied self training to obtain pseudo label from DESED weak set, unlabeled set and AudioSet.
-
19 Jun 2024 1 repository listedFrom the results of extensive ablation studies, we discovered that not only multiple dynamic branches but also specific proportion of static branch helps SED.
Syntology lines on 6 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections