{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-simple-fusion-of-deep-and-shallow-learning","title":"A Simple Fusion of Deep and Shallow Learning for Acoustic Scene Classification","arxiv_id":"1806.07506","date":"2018-06-19","proceeding":null,"authors":["Eduardo Fonseca","Rong Gong","Xavier Serra"],"abstract":"In the past, Acoustic Scene Classification systems have been based on hand\ncrafting audio features that are input to a classifier. Nowadays, the common\ntrend is to adopt data driven techniques, e.g., deep learning, where audio\nrepresentations are learned from data. In this paper, we propose a system that\nconsists of a simple fusion of two methods of the aforementioned types: a deep\nlearning approach where log-scaled mel-spectrograms are input to a\nconvolutional neural network, and a feature engineering approach, where a\ncollection of hand-crafted features is input to a gradient boosting machine. We\nfirst show that both methods provide complementary information to some extent.\nThen, we use a simple late fusion strategy to combine both methods. We report\nclassification accuracy of each method individually and the combined system on\nthe TUT Acoustic Scenes 2017 dataset. The proposed fused system outperforms\neach of the individual methods and attains a classification accuracy of 72.8%\non the evaluation set, improving the baseline system by 11.8%.","url_abs":"http://arxiv.org/abs/1806.07506v2","url_pdf":"http://arxiv.org/pdf/1806.07506v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-simple-fusion-of-deep-and-shallow-learning","repo_url":"https://github.com/edufonseca/icassp19","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"a-simple-fusion-of-deep-and-shallow-learning","repo_url":"https://github.com/tbayetird/audio_environment_description","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"acoustic-scene-classification","task_name":"Acoustic Scene Classification"},{"task_slug":"classification-1","task_name":"Classification"},{"task_slug":"feature-engineering","task_name":"Feature Engineering"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"scene-classification","task_name":"Scene Classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}