{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/unsupervised-feature-learning-based-on-deep","title":"Unsupervised Feature Learning Based on Deep Models for Environmental Audio Tagging","arxiv_id":"1607.03681","date":"2016-07-13","proceeding":null,"authors":["Yong Xu","Qiang Huang","Wenwu Wang","Peter Foster","Siddharth Sigtia","Philip J. B. Jackson","Mark D. Plumbley"],"abstract":"Environmental audio tagging aims to predict only the presence or absence of\ncertain acoustic events in the interested acoustic scene. In this paper we make\ncontributions to audio tagging in two parts, respectively, acoustic modeling\nand feature learning. We propose to use a shrinking deep neural network (DNN)\nframework incorporating unsupervised feature learning to handle the multi-label\nclassification task. For the acoustic modeling, a large set of contextual\nframes of the chunk are fed into the DNN to perform a multi-label\nclassification for the expected tags, considering that only chunk (or\nutterance) level rather than frame-level labels are available. Dropout and\nbackground noise aware training are also adopted to improve the generalization\ncapability of the DNNs. For the unsupervised feature learning, we propose to\nuse a symmetric or asymmetric deep de-noising auto-encoder (sDAE or aDAE) to\ngenerate new data-driven features from the Mel-Filter Banks (MFBs) features.\nThe new features, which are smoothed against background noise and more compact\nwith contextual information, can further improve the performance of the DNN\nbaseline. Compared with the standard Gaussian Mixture Model (GMM) baseline of\nthe DCASE 2016 audio tagging challenge, our proposed method obtains a\nsignificant equal error rate (EER) reduction from 0.21 to 0.13 on the\ndevelopment set. The proposed aDAE system can get a relative 6.7% EER reduction\ncompared with the strong DNN baseline on the development set. Finally, the\nresults also show that our approach obtains the state-of-the-art performance\nwith 0.15 EER on the evaluation set of the DCASE 2016 audio tagging task while\nEER of the first prize of this challenge is 0.17.","url_abs":"http://arxiv.org/abs/1607.03681v2","url_pdf":"http://arxiv.org/pdf/1607.03681v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"unsupervised-feature-learning-based-on-deep","repo_url":"https://github.com/yongxuUSTC/aDAE_DNN_audio_tagging","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null},{"paper_slug":"unsupervised-feature-learning-based-on-deep","repo_url":"https://github.com/lgerrets/asait18-tagging","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"audio-tagging","task_name":"Audio Tagging"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"multi-label-classification-2","task_name":"MUlTI-LABEL-ClASSIFICATION"},{"task_slug":"multi-label-classification","task_name":"Multi-Label Classification"}],"methods":[{"method_slug":"dropout","method_name":"Dropout"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}