{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/convolutional-gated-recurrent-neural-network","title":"Convolutional Gated Recurrent Neural Network Incorporating Spatial Features for Audio Tagging","arxiv_id":"1702.07787","date":"2017-02-24","proceeding":null,"authors":["Yong Xu","Qiuqiang Kong","Qiang Huang","Wenwu Wang","Mark D. Plumbley"],"abstract":"Environmental audio tagging is a newly proposed task to predict the presence\nor absence of a specific audio event in a chunk. Deep neural network (DNN)\nbased methods have been successfully adopted for predicting the audio tags in\nthe domestic audio scene. In this paper, we propose to use a convolutional\nneural network (CNN) to extract robust features from mel-filter banks (MFBs),\nspectrograms or even raw waveforms for audio tagging. Gated recurrent unit\n(GRU) based recurrent neural networks (RNNs) are then cascaded to model the\nlong-term temporal structure of the audio signal. To complement the input\ninformation, an auxiliary CNN is designed to learn on the spatial features of\nstereo recordings. We evaluate our proposed methods on Task 4 (audio tagging)\nof the Detection and Classification of Acoustic Scenes and Events 2016 (DCASE\n2016) challenge. Compared with our recent DNN-based method, the proposed\nstructure can reduce the equal error rate (EER) from 0.13 to 0.11 on the\ndevelopment set. The spatial features can further reduce the EER to 0.10. The\nperformance of the end-to-end learning on raw waveforms is also comparable.\nFinally, on the evaluation set, we get the state-of-the-art performance with\n0.12 EER while the performance of the best existing system is 0.15 EER.","url_abs":"http://arxiv.org/abs/1702.07787v1","url_pdf":"http://arxiv.org/pdf/1702.07787v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"convolutional-gated-recurrent-neural-network","repo_url":"https://github.com/yongxuUSTC/cnn_rnn_spatial_audio_tagging","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null},{"paper_slug":"convolutional-gated-recurrent-neural-network","repo_url":"https://github.com/mariyashcheg/kaggle-freesound-2019","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"audio-tagging","task_name":"Audio Tagging"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}