Browse State-of-the-Art › Sound Source Localization
Sound Source Localization
39 papers with code · 1 benchmark · 0 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
1 leaderboard table shown for this task, 1 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| ^(#!@#)(()))****** (1 row) | YOLO | ODAS: Open embeddeD Audition System | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
No dataset record in the archive lists this task.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 39 papers with code (104 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
24 Jun 2022 2 repositories listedSound source localization aims to seek the direction of arrival (DOA) of all sound sources from the observed multi-channel audio.
-
29 Oct 2021 2 repositories listedData-based and learning-based sound source localization (SSL) has shown promising results in challenging conditions, and is commonly set as a classification or a regression problem.
-
8 May 2025 1 repository listedWe introduce a framework that maps audios into tokens compatible with CLIP's text encoder, producing audio-driven embeddings.
-
1 Jan 2025 1 repository listedAudio-visual sound source localization task aims to spatially localize sound-making objects within visual scenes by integrating visual and audio cues.
-
9 Dec 2024 1 repository listed Syntology ran 5 of 18 samples · 13 unverifiedWe address this challenge by designing a model that aligns audio-visual modalities by enriching audio features with visual information and translating them into the visual latent space.
-
1 Oct 2024 1 repository listedThe task of Visual Sound Source Localization (VSSL) involves identifying the location of sound sources in visual scenes, integrating audio-visual data for enhanced scene understanding.
-
29 Aug 2024 1 repository listedTo address this issue, we propose a novel audio-visual learning framework which is instantiated with two individual learning schemes: self-supervised predictive learning (SSPL) and semantic-aware contrastive learning…
-
28 Aug 2024 1 repository listedWe present a novel approach to the 3D sound source localization task for distributed ad-hoc microphone arrays by formulating it as a set-to-set regression problem.
-
18 Jul 2024 1 repository listedSecond, we introduce new evaluation metrics to rigorously assess sound source localization methods, focusing on accurately evaluating both localization performance and cross-modal interaction ability.
-
11 May 2024 1 repository listedSecond, a new multi-track DP-IPD learning target is proposed for the localization of flexible number of sound sources.
-
30 Apr 2024 1 repository listedIn recent years, Event Sound Source Localization has been widely applied in various fields.
-
26 Mar 2024 1 repository listedIn this paper, to overcome this limitation, we present a novel multi-sound source localization method that can perform localization without prior knowledge of the number of sound sources.
-
26 Mar 2024 1 repository listedDistance estimation from audio plays a crucial role in various applications, such as acoustic scene analysis, sound source localization, and room modeling.
-
21 Nov 2023 1 repository listedTo address this, we propose an Unbiased Label Distribution (ULD) to eliminate quantization error in training targets.
-
7 Nov 2023 1 repository listed Syntology ran 10 of 10 samples · 0 unverified · 10 pointer-only (licence)Large-scale pre-trained image-text models demonstrate remarkable versatility across diverse tasks, benefiting from their robust representational capabilities and effective multimodal alignment.
-
2 Nov 2023 1 repository listed Syntology ran 6 of 6 samples · 0 unverifiedMultiple Choice Learning is a simple framework to tackle multimodal density estimation, using the Winner-Takes-All (WTA) loss for a set of hypotheses.
-
28 Oct 2023 1 repository listedIn this paper, we propose a new multi-modal task, termed audio-visual instance segmentation (AVIS), which aims to simultaneously identify, segment and track individual sounding object instances in audible videos.
-
23 Oct 2023 1 repository listedGiven audio recordings from 2-4 microphones and the 3D geometry and material of a scene containing multiple unknown sound sources, we estimate the sound anywhere in the scene.
-
11 Aug 2023 1 repository listedWhile the audio modality provides spatial cues to locate the sound source, existing approaches only use audio as an auxiliary role to compare spatial regions of the visual modality.
-
9 Aug 2023 1 repository listedBy decoupling the gradients of visual and audio modalities, the discriminative visual representations of sound sources can be learned with the designed Induction Vector in a bootstrap manner, which also enables the…
-
8 Aug 2023 1 repository listedIn many signal processing applications, metadata may be advantageously used in conjunction with a high dimensional signal to produce a desired output.
-
31 May 2023 1 repository listedExtracting direct-path spatial features is critical for sound source localization in adverse acoustic environments.
-
29 Mar 2023 1 repository listed Syntology ran 6 of 11 samples · 5 unverifiedSound source localization is a typical and challenging task that predicts the location of sound sources in a video.
-
1 Jan 2023 1 repository listedUnderstanding and analyzing human behaviors (actions and interactions of people), voices, and sounds in chaotic events is crucial in many applications, e.
-
15 Nov 2022 1 repository listedMost recent work in visual sound source localization relies on semantic audio-visual representations learned in a self-supervised manner, and by design excludes temporal information present in videos.
-
6 Nov 2022 1 repository listedExisting work in this area focuses on creating attention maps to capture the correlation between the two modalities to localize the source of the sound.
-
30 Aug 2022 1 repository listed Syntology ran 4 of 6 samples · 2 unverifiedWe also propose a new approach for visual sound source localization that addresses both these problems.
-
25 Mar 2022 1 repository listedSound source localization in visual scenes aims to localize objects emitting the sound in a given image.
-
7 Mar 2022 1 repository listedThe origin transfer function at the head center was obtained by the proposed measurement scheme using a 0 degree on-axis microphone to ensure accurate spectral cue pattern of HRTFs, whereas in the previous measurements…
-
13 Feb 2022 1 repository listedSpecifically, we observe that the previous practice of learning only a single audio representation is insufficient due to the additive nature of audio signals.
Syntology lines on 5 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections