{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/end-to-end-environmental-sound-classification","title":"End-to-End Environmental Sound Classification using a 1D Convolutional Neural Network","arxiv_id":"1904.08990","date":"2019-04-18","proceeding":null,"authors":["Sajjad Abdoli","Patrick Cardinal","Alessandro Lameiras Koerich"],"abstract":"In this paper, we present an end-to-end approach for environmental sound\nclassification based on a 1D Convolution Neural Network (CNN) that learns a\nrepresentation directly from the audio signal. Several convolutional layers are\nused to capture the signal's fine time structure and learn diverse filters that\nare relevant to the classification task. The proposed approach can deal with\naudio signals of any length as it splits the signal into overlapped frames\nusing a sliding window. Different architectures considering several input sizes\nare evaluated, including the initialization of the first convolutional layer\nwith a Gammatone filterbank that models the human auditory filter response in\nthe cochlea. The performance of the proposed end-to-end approach in classifying\nenvironmental sounds was assessed on the UrbanSound8k dataset and the\nexperimental results have shown that it achieves 89% of mean accuracy.\nTherefore, the propose approach outperforms most of the state-of-the-art\napproaches that use handcrafted features or 2D representations as input.\nFurthermore, the proposed approach has a small number of parameters compared to\nother architectures found in the literature, which reduces the amount of data\nrequired for training.","url_abs":"http://arxiv.org/abs/1904.08990v1","url_pdf":"http://arxiv.org/pdf/1904.08990v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"end-to-end-environmental-sound-classification","repo_url":"https://github.com/sajabdoli/Environmental_sound_classification","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":null},{"paper_slug":"end-to-end-environmental-sound-classification","repo_url":"https://github.com/Logan97117/environmental_sound_classification_1DCNN","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"environmental-sound-classification","task_name":"Environmental Sound Classification"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"sound-classification","task_name":"Sound Classification"}],"methods":[{"method_slug":"convolution","method_name":"Convolution"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/environmental-sound-classification-on","task":"Environmental Sound Classification","dataset":"UrbanSound8K","model":"1DCNN","rank_in_archive_order":3,"of":3,"metrics":{"Accuracy":"89"},"uses_additional_data":true}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}