Papers › Leveraging LLM and Text-Queried Separation for Noise-Robust Sound Event Detection

Leveraging LLM and Text-Queried Separation for Noise-Robust Sound Event Detection

2 Nov 2024arXiv:2411.01174archive 2025-07-28

Han Yin, Yang Xiao, Jisheng Bai, Rohan Kumar Das

Sound Event Detection (SED) is challenging in noisy environments where overlapping sounds obscure target events. Language-queried audio source separation (LASS) aims to isolate the target sound events from a noisy clip. However, this approach can fail when the exact target sound is unknown, particularly in noisy test sets, leading to reduced performance. To address this issue, we leverage the capabilities of large language models (LLMs) to analyze and summarize acoustic data. By using LLMs to identify and select specific noise types, we implement a noise augmentation method for noise-robust fine-tuning. The fine-tuned model is applied to predict clip-wise event predictions as text queries for the LASS model. Our studies demonstrate that the proposed method improves SED performance in noisy environments. This work represents an early application of LLMs in noise-robust SED and suggests a promising direction for handling overlapping events in SED. Codes and pretrained models are available at https://github.com/apple-yinhan/Noise-robust-SED.

PaperPDFCode

Code

apple-yinhan/noise-robust-sed officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Audio Source SeparationEvent DetectionSound Event Detection

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Sound Event Detection WildDESED CRNN (with BEATs + Separation) PSDS1 (-5dB) 0.134 #1 of 5 Archive leaderboard report
Sound Event Detection WildDESED CRNN (with BEATs + Separation) PSDS1 (0dB) 0.219 #1 of 5 Archive leaderboard report
Sound Event Detection WildDESED CRNN (with BEATs + Separation) PSDS1 (10dB) 0.356 #1 of 5 Archive leaderboard report
Sound Event Detection WildDESED CRNN (with BEATs + Separation) PSDS1 (5dB) 0.291 #1 of 5 Archive leaderboard report
Sound Event Detection WildDESED CRNN (with BEATs + Separation) PSDS1 (Clean) 0.440 #1 of 5 Archive leaderboard report
Sound Event Detection WildDESED CRNN (with BEATs) PSDS1 (-5dB) 0.065 #2 of 5 Archive leaderboard report
Sound Event Detection WildDESED CRNN (with BEATs) PSDS1 (0dB) 0.138 #2 of 5 Archive leaderboard report
Sound Event Detection WildDESED CRNN (with BEATs) PSDS1 (10dB) 0.329 #2 of 5 Archive leaderboard report
Sound Event Detection WildDESED CRNN (with BEATs) PSDS1 (5dB) 0.236 #2 of 5 Archive leaderboard report
Sound Event Detection WildDESED CRNN (with BEATs) PSDS1 (Clean) 0.500 #2 of 5 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections