{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/device-robust-acoustic-scene-classification","title":"Device-Robust Acoustic Scene Classification Based on Two-Stage Categorization and Data Augmentation","arxiv_id":"2007.08389","date":"2020-07-16","proceeding":null,"authors":["Hu Hu","Chao-Han Huck Yang","Xianjun Xia","Xue Bai","Xin Tang","Yajian Wang","Shutong Niu","Li Chai","Juanjuan Li","Hongning Zhu","Feng Bao","Yuanjun Zhao","Sabato Marco Siniscalchi","Yannan Wang","Jun Du","Chin-Hui Lee"],"abstract":"In this technical report, we present a joint effort of four groups, namely GT, USTC, Tencent, and UKE, to tackle Task 1 - Acoustic Scene Classification (ASC) in the DCASE 2020 Challenge. Task 1 comprises two different sub-tasks: (i) Task 1a focuses on ASC of audio signals recorded with multiple (real and simulated) devices into ten different fine-grained classes, and (ii) Task 1b concerns with classification of data into three higher-level classes using low-complexity solutions. For Task 1a, we propose a novel two-stage ASC system leveraging upon ad-hoc score combination of two convolutional neural networks (CNNs), classifying the acoustic input according to three classes, and then ten classes, respectively. Four different CNN-based architectures are explored to implement the two-stage classifiers, and several data augmentation techniques are also investigated. For Task 1b, we leverage upon a quantization method to reduce the complexity of two of our top-accuracy three-classes CNN-based architectures. On Task 1a development data set, an ASC accuracy of 76.9\\% is attained using our best single classifier and data augmentation. An accuracy of 81.9\\% is then attained by a final model fusion of our two-stage ASC classifiers. On Task 1b development data set, we achieve an accuracy of 96.7\\% with a model size smaller than 500KB. Code is available: https://github.com/MihawkHu/DCASE2020_task1.","url_abs":"https://arxiv.org/abs/2007.08389v2","url_pdf":"https://arxiv.org/pdf/2007.08389v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"device-robust-acoustic-scene-classification","repo_url":"https://github.com/MihawkHu/DCASE2020_task1","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"acoustic-scene-classification","task_name":"Acoustic Scene Classification"},{"task_slug":"data-augmentation","task_name":"Data Augmentation"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"quantization","task_name":"Quantization"},{"task_slug":"scene-classification","task_name":"Scene Classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2007.08389","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}