{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/convolutional-recurrent-neural-networks-for-6","title":"Convolutional Recurrent Neural Networks for Polyphonic Sound Event Detection","arxiv_id":"1702.06286","date":"2017-02-21","proceeding":null,"authors":["Emre Çakır","Giambattista Parascandolo","Toni Heittola","Heikki Huttunen","Tuomas Virtanen"],"abstract":"Sound events often occur in unstructured environments where they exhibit wide\nvariations in their frequency content and temporal structure. Convolutional\nneural networks (CNN) are able to extract higher level features that are\ninvariant to local spectral and temporal variations. Recurrent neural networks\n(RNNs) are powerful in learning the longer term temporal context in the audio\nsignals. CNNs and RNNs as classifiers have recently shown improved performances\nover established methods in various sound recognition tasks. We combine these\ntwo approaches in a Convolutional Recurrent Neural Network (CRNN) and apply it\non a polyphonic sound event detection task. We compare the performance of the\nproposed CRNN method with CNN, RNN, and other established methods, and observe\na considerable improvement for four different datasets consisting of everyday\nsound events.","url_abs":"http://arxiv.org/abs/1702.06286v1","url_pdf":"http://arxiv.org/pdf/1702.06286v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"convolutional-recurrent-neural-networks-for-6","repo_url":"https://github.com/cchinchristopherj/Right-Whale-Unsupervised-Model","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"event-detection","task_name":"Event Detection"},{"task_slug":"sound-event-detection","task_name":"Sound Event Detection"}],"methods":[],"datasets_introduced":[{"slug":"tut-sed-synthetic-2016","name":"TUT-SED Synthetic 2016","full_name":"TUT-SED Synthetic 2016"}],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1702.06286","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}