{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/acoustic-scene-classification-by-implicitly","title":"Acoustic Scene Classification by Implicitly Identifying Distinct Sound Events","arxiv_id":"1904.05204","date":"2019-04-10","proceeding":null,"authors":["Hongwei Song","Jiqing Han","Shiwen Deng","Zhihao Du"],"abstract":"In this paper, we propose a new strategy for acoustic scene classification\n(ASC) , namely recognizing acoustic scenes through identifying distinct sound\nevents. This differs from existing strategies, which focus on characterizing\nglobal acoustical distributions of audio or the temporal evolution of\nshort-term audio features, without analysis down to the level of sound events.\nTo identify distinct sound events for each scene, we formulate ASC in a\nmulti-instance learning (MIL) framework, where each audio recording is mapped\ninto a bag-of-instances representation. Here, instances can be seen as\nhigh-level representations for sound events inside a scene. We also propose a\nMIL neural networks model, which implicitly identifies distinct instances\n(i.e., sound events). Furthermore, we propose two specially designed modules\nthat model the multi-temporal scale and multi-modal natures of the sound events\nrespectively. The experiments were conducted on the official development set of\nthe DCASE2018 Task1 Subtask B, and our best-performing model improves over the\nofficial baseline by 9.4% (68.3% vs 58.9%) in terms of classification accuracy.\nThis study indicates that recognizing acoustic scenes by identifying distinct\nsound events is effective and paves the way for future studies that combine\nthis strategy with previous ones.","url_abs":"http://arxiv.org/abs/1904.05204v2","url_pdf":"http://arxiv.org/pdf/1904.05204v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"acoustic-scene-classification-by-implicitly","repo_url":"https://github.com/hackerekcah/distinct-events-asc","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"acoustic-scene-classification","task_name":"Acoustic Scene Classification"},{"task_slug":"classification-1","task_name":"Classification"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"scene-classification","task_name":"Scene Classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}