{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/wilddesed-an-llm-powered-dataset-for-wild","title":"WildDESED: An LLM-Powered Dataset for Wild Domestic Environment Sound Event Detection System","arxiv_id":"2407.03656","date":"2024-07-04","proceeding":null,"authors":["Yang Xiao","Rohan Kumar Das"],"abstract":"This work aims to advance sound event detection (SED) research by presenting a new large language model (LLM)-powered dataset namely wild domestic environment sound event detection (WildDESED). It is crafted as an extension to the original DESED dataset to reflect diverse acoustic variability and complex noises in home settings. We leveraged LLMs to generate eight different domestic scenarios based on target sound categories of the DESED dataset. Then we enriched the scenarios with a carefully tailored mixture of noises selected from AudioSet and ensured no overlap with target sound. We consider widely popular convolutional neural recurrent network to study WildDESED dataset, which depicts its challenging nature. We then apply curriculum learning by gradually increasing noise complexity to enhance the model's generalization capabilities across various noise levels. Our results with this approach show improvements within the noisy environment, validating the effectiveness on the WildDESED dataset promoting noise-robust SED advancements.","url_abs":"https://arxiv.org/abs/2407.03656v3","url_pdf":"https://arxiv.org/pdf/2407.03656v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"wilddesed-an-llm-powered-dataset-for-wild","repo_url":"https://github.com/swagshaw/wilddesed","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"event-detection","task_name":"Event Detection"},{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"large-language-model","task_name":"Large Language Model"},{"task_slug":"sound-event-detection","task_name":"Sound Event Detection"}],"methods":[],"datasets_introduced":[{"slug":"wilddesed","name":"WildDESED","full_name":"Wild Domestic Environment Sound Event Detection"}],"methods_introduced":[],"results":[{"leaderboard":"/sota/sound-event-detection-on-wilddesed","task":"Sound Event Detection","dataset":"WildDESED","model":"CRNN (WildDESED + Curriculrm learning)","rank_in_archive_order":3,"of":5,"metrics":{"PSDS1 (-5dB)":"0.049","PSDS1 (0dB)":"0.114","PSDS1 (10dB)":"0.212","PSDS1 (5dB)":"0.175","PSDS1 (Clean)":"0.265"},"uses_additional_data":false},{"leaderboard":"/sota/sound-event-detection-on-wilddesed","task":"Sound Event Detection","dataset":"WildDESED","model":"CRNN (WildDESED)","rank_in_archive_order":4,"of":5,"metrics":{"PSDS1 (-5dB)":"0.048","PSDS1 (0dB)":"0.087","PSDS1 (10dB)":"0.175","PSDS1 (5dB)":"0.135","PSDS1 (Clean)":"0.200"},"uses_additional_data":false},{"leaderboard":"/sota/sound-event-detection-on-wilddesed","task":"Sound Event Detection","dataset":"WildDESED","model":"CRNN","rank_in_archive_order":5,"of":5,"metrics":{"PSDS1 (-5dB)":"0.017","PSDS1 (0dB)":"0.064","PSDS1 (10dB)":"0.222","PSDS1 (5dB)":"0.148","PSDS1 (Clean)":"0.348"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}