{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/180606773","title":"Towards an efficient deep learning model for musical onset detection","arxiv_id":"1806.06773","date":"2018-06-18","proceeding":null,"authors":["Rong Gong","Xavier Serra"],"abstract":"In this paper, we propose an efficient and reproducible deep learning model\nfor musical onset detection (MOD). We first review the state-of-the-art deep\nlearning models for MOD, and identify their shortcomings and challenges: (i)\nthe lack of hyper-parameter tuning details, (ii) the non-availability of code\nfor training models on other datasets, and (iii) ignoring the network\ncapability when comparing different architectures. Taking the above issues into\naccount, we experiment with seven deep learning architectures. The most\nefficient one achieves equivalent performance to our implementation of the\nstate-of-the-art architecture. However, it has only 28.3% of the total number\nof trainable parameters compared to the state-of-the-art. Our experiments are\nconducted using two different datasets: one mainly consists of instrumental\nmusic excerpts, and another developed by ourselves includes only solo singing\nvoice excerpts. Further, inter-dataset transfer learning experiments are\nconducted. The results show that the model pre-trained on one dataset fails to\ndetect onsets on another dataset, which denotes the importance of providing the\nimplementation code to enable re-training the model for a different dataset.\nDatasets, code and a Jupyter notebook running on Google Colab are publicly\navailable to make this research understandable and easy to reproduce.","url_abs":"http://arxiv.org/abs/1806.06773v2","url_pdf":"http://arxiv.org/pdf/1806.06773v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"180606773","repo_url":"https://github.com/ronggong/musical-onset-efficient","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null},{"paper_slug":"180606773","repo_url":"https://github.com/YingjingLu/Music_Onset","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"deep-learning","task_name":"Deep Learning"},{"task_slug":"onset-detection","task_name":"Onset Detection"},{"task_slug":"transfer-learning","task_name":"Transfer Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}