Browse State-of-the-Art › Sound Event Localization and Detection
Sound Event Localization and Detection
35 papers with code · 5 benchmarks · 9 datasets archive 2025-07-28
Given multichannel audio input, a sound event detection and localization (SELD) system outputs a temporal activation track for each of the target sound classes, along with one or more corresponding spatial trajectories when the track indicates activity. This results in a spatio-temporal characterization of the acoustic scene that can be used in a wide range of machine cognition tasks, such as inference on the type of environment, self-localization, navigation without visual input or with occluded targets, tracking of specific types of sound sources, smart-home applications, scene visualization systems, and audio surveillance, among others.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
5 leaderboard tables shown for this task, 5 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| STARSS22 (2 rows) | Baseline (FOA) | STARSS22: A dataset of spatial recordings of real scenes with... | code | — | Compare |
| PodcastFillers (2 rows) | AVC-FillerNet | Filler Word Detection and Classification: A Dataset and Benchmark | code | — | Compare |
| L3DAS21 (1 row) | DualQSELD-TCN (parallel) | Dual Quaternion Ambisonics Array for Six-Degree-of-Freedom... | code | — | Compare |
| RWCP Sound Scene Database (1 row) | STL-SNN | A Synapse-Threshold Synergistic Learning Approach for Spiking... | code | — | Compare |
| TAU-NIGENS Spatial Sound Events 2021 (1 row) | SALSA-FOA | SALSA: Spatial Cue-Augmented Log-Spectrogram Features for... | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
9 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 35 papers with code (65 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
16 Nov 2021 4 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedIn this work, we introduce SALSA-Lite, a fast and effective feature for polyphonic SELD using microphone array inputs.
-
6 Sep 2020 4 repositories listedA large-scale realistic dataset of spatialized sound events was generated for the challenge, to be used for training of learning-based approaches, and for evaluation of the submissions in an unlabeled subset.
-
10 Nov 2024 2 repositories listedRecently, deep neural networks trained on large-scale datasets have achieved remarkable success in the sound event classification (SEC) field, prompting an open question of whether these advancements can be extended to…
-
5 Sep 2022 2 repositories listedOur system submitted to the DCASE 2022 Task 3 is based on our previous proposed Event-Independent Network V2 (EINV2) with a novel data augmentation method.
-
4 Jun 2022 2 repositories listedAdditionally, the report presents the baseline system that accompanies the dataset in the challenge with emphasis on the differences with the baseline of the previous iterations; namely, introduction of the multi-ACCDOA…
-
14 Oct 2021 2 repositories listedThe multi- ACCDOA format (a class- and track-wise output format) enables the model to solve the cases with overlaps from the same class.
-
29 Oct 2020 2 repositories listedConventional NN-based methods use two branches for a sound event detection (SED) target and a direction-of-arrival (DOA) target.
-
2 Jun 2020 2 repositories listed Syntology ran 0 of 1 samples · 1 unverified · 1 pointer-only (licence)This report presents the dataset and the evaluation setup of the Sound Event Localization & Detection (SELD) task for the DCASE 2020 Challenge.
-
3 Mar 2020 2 repositories listedThe understanding of the surrounding environment plays a critical role in autonomous robotic systems, such as self-driving cars.
-
11 Apr 2025 1 repository listedOnly recently, SELD was extended to 3D by incorporating distance estimation, enabling the prediction of sound event positions in 3D space (3D SELD).
-
21 Nov 2024 1 repository listedWe propose a novel output representation that combines the DOA with distance of sound sources by calculating the real Cartesian coordinates to address the newly introduced source distance estimation (SDE) task in the…
-
18 Sep 2024 1 repository listedSound event localization and detection (SELD) is critical for various real-world applications, including smart monitoring and Internet of Things (IoT) systems.
-
30 Aug 2024 1 repository listedSound event localization and detection (SELD) systems using audio recordings from a microphone array rely on spatial cues for determining the location of sound events.
-
13 Jun 2024 1 repository listed Syntology ran 3 of 4 samples · 1 unverifiedThis paper proposes a three-stage network structure named Multi-scale Feature Fusion (MFF) module to fully extract multi-scale features across spectral, spatial, and temporal domains.
-
29 Jan 2024 1 repository listedThis technical report details our work towards building an enhanced audio-visual sound event localization and detection (SELD) network.
-
19 Jan 2024 1 repository listed Syntology ran 11 of 12 samples · 1 unverified · 12 pointer-only (licence)Major advancements rely on simulated data with sound events in specific rooms and strong spatio-temporal labels.
-
27 Dec 2023 1 repository listedIn addition, we introduce environment representations to characterize different acoustic settings, enhancing the adaptability of our attenuation approach to various environments.
-
14 Dec 2023 1 repository listedSound event localization and detection (SELD) combines two subtasks: sound event detection (SED) and direction of arrival (DOA) estimation.
-
12 Dec 2023 1 repository listedBy applying this approach to SELD, we can leverage a substantial amount of unlabeled 3D audio data to learn robust representations of sound events and their locations.
-
6 Sep 2023 1 repository listedAs deeper and more complex models are developed for the task of sound event localization and detection (SELD), the demand for annotated spatial audio data continues to increase.
-
15 Jun 2023 1 repository listed Syntology ran 5 of 11 samples · 6 unverifiedWhile direction of arrival (DOA) of sound events is generally estimated from multichannel audio data recorded in a microphone array, sound events usually derive from visually perceptible source objects, e.
-
23 May 2023 1 repository listedWe propose a novel multimodal video benchmark - the Perception Test - to evaluate the perception and reasoning skills of pre-trained multimodal models (e.
-
28 Mar 2023 1 repository listed Syntology ran 1 of 8 samples · 7 unverifiedHence, the format enables the model to handle the polyphony problem, regardless of the number of sound overlaps.
-
10 Jun 2022 1 repository listedMost existing methods for training SNNs are based on the concept of synaptic plasticity; however, learning in the realistic brain also utilizes intrinsic non-synaptic mechanisms of neurons.
-
4 Apr 2022 1 repository listedWe show that our dual quaternion SELD model with temporal convolution blocks (DualQSELD-TCN) achieves better results with respect to real and quaternion-valued baselines thanks to our augmented representation of the…
-
28 Mar 2022 1 repository listedIn this work, we present a novel speech dataset, PodcastFillers, with 35K annotated filler words and 50K annotations of other sounds that commonly occur in podcasts such as breaths, laughter, and word repetitions.
-
21 Feb 2022 1 repository listed Syntology ran 0 of 1 samples · 1 unverified · 1 pointer-only (licence)The L3DAS22 Challenge is aimed at encouraging the development of machine learning strategies for 3D speech enhancement and 3D sound localization and detection in office-like environments.
-
18 Feb 2022 1 repository listedOur goal is to develop a sound event localization and detection (SELD) system that works robustly in unknown environments.
-
17 Feb 2022 1 repository listedSound event localization and detection (SELD) is a combined task of identifying the sound event and its direction.
-
12 Oct 2021 1 repository listedData augmentation methods have shown great importance in diverse supervised learning problems where labeled data is scarce or costly to obtain.
Syntology lines on 7 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections