{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-comparison-of-deep-learning-methods-for","title":"A Comparison of deep learning methods for environmental sound","arxiv_id":"1703.06902","date":"2017-03-20","proceeding":null,"authors":["Juncheng Li","Wei Dai","Florian Metze","Shuhui Qu","Samarjit Das"],"abstract":"Environmental sound detection is a challenging application of machine\nlearning because of the noisy nature of the signal, and the small amount of\n(labeled) data that is typically available. This work thus presents a\ncomparison of several state-of-the-art Deep Learning models on the IEEE\nchallenge on Detection and Classification of Acoustic Scenes and Events (DCASE)\n2016 challenge task and data, classifying sounds into one of fifteen common\nindoor and outdoor acoustic scenes, such as bus, cafe, car, city center, forest\npath, library, train, etc. In total, 13 hours of stereo audio recordings are\navailable, making this one of the largest datasets available. We perform\nexperiments on six sets of features, including standard Mel-frequency cepstral\ncoefficients (MFCC), Binaural MFCC, log Mel-spectrum and two different large-\nscale temporal pooling features extracted using OpenSMILE. On these features,\nwe apply five models: Gaussian Mixture Model (GMM), Deep Neural Network (DNN),\nRecurrent Neural Network (RNN), Convolutional Deep Neural Net- work (CNN) and\ni-vector. Using the late-fusion approach, we improve the performance of the\nbaseline 72.5% by 15.6% in 4-fold Cross Validation (CV) avg. accuracy and 11%\nin test accuracy, which matches the best result of the DCASE 2016 challenge.\nWith large feature sets, deep neural network models out- perform traditional\nmethods and achieve the best performance among all the studied methods.\nConsistent with other work, the best performing single model is the\nnon-temporal DNN model, which we take as evidence that sounds in the DCASE\nchallenge do not exhibit strong temporal dynamics.","url_abs":"http://arxiv.org/abs/1703.06902v1","url_pdf":"http://arxiv.org/pdf/1703.06902v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-comparison-of-deep-learning-methods-for","repo_url":"https://github.com/lijuncheng16/AudioTaggingDoneRight","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":null,"task_name":"Avg"},{"task_slug":"deep-learning","task_name":"Deep Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}