{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/general-audio-tagging-with-ensembling","title":"General audio tagging with ensembling convolutional neural network and statistical features","arxiv_id":"1810.12832","date":"2018-10-30","proceeding":null,"authors":["Kele Xu","Boqing Zhu","Qiuqiang Kong","Haibo Mi","Bo Ding","Dezhi Wang","Huaimin Wang"],"abstract":"Audio tagging aims to infer descriptive labels from audio clips. Audio\ntagging is challenging due to the limited size of data and noisy labels. In\nthis paper, we describe our solution for the DCASE 2018 Task 2 general audio\ntagging challenge. The contributions of our solution include: We investigated a\nvariety of convolutional neural network architectures to solve the audio\ntagging task. Statistical features are applied to capture statistical patterns\nof audio features to improve the classification performance. Ensemble learning\nis applied to ensemble the outputs from the deep classifiers to utilize\ncomplementary information. a sample re-weight strategy is employed for ensemble\ntraining to address the noisy label problem. Our system achieves a mean average\nprecision (mAP@3) of 0.958, outperforming the baseline system of 0.704. Our\nsystem ranked the 1st and 4th out of 558 submissions in the public and private\nleaderboard of DCASE 2018 Task 2 challenge. Our codes are available at\nhttps://github.com/Cocoxili/DCASE2018Task2/.","url_abs":"http://arxiv.org/abs/1810.12832v1","url_pdf":"http://arxiv.org/pdf/1810.12832v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"general-audio-tagging-with-ensembling","repo_url":"https://github.com/Cocoxili/DCASE2018Task2","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"general-audio-tagging-with-ensembling","repo_url":"https://github.com/r0mer0m/learning_audio_modeling","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"audio-tagging","task_name":"Audio Tagging"},{"task_slug":"descriptive","task_name":"Descriptive"},{"task_slug":"ensemble-learning","task_name":"Ensemble Learning"},{"task_slug":"task-2","task_name":"Task 2"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}