{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-deep-bag-of-features-model-for-music-auto","title":"A Deep Bag-of-Features Model for Music Auto-Tagging","arxiv_id":"1508.04999","date":"2015-08-20","proceeding":null,"authors":["Juhan Nam","Jorge Herrera","Kyogu Lee"],"abstract":"Feature learning and deep learning have drawn great attention in recent years\nas a way of transforming input data into more effective representations using\nlearning algorithms. Such interest has grown in the area of music information\nretrieval (MIR) as well, particularly in music audio classification tasks such\nas auto-tagging. In this paper, we present a two-stage learning model to\neffectively predict multiple labels from music audio. The first stage learns to\nproject local spectral patterns of an audio track onto a high-dimensional\nsparse space in an unsupervised manner and summarizes the audio track as a\nbag-of-features. The second stage successively performs the unsupervised\nlearning on the bag-of-features in a layer-by-layer manner to initialize a deep\nneural network and finally fine-tunes it with the tag labels. Through the\nexperiment, we rigorously examine training choices and tuning parameters, and\nshow that the model achieves high performance on Magnatagatune, a popularly\nused dataset in music auto-tagging.","url_abs":"http://arxiv.org/abs/1508.04999v3","url_pdf":"http://arxiv.org/pdf/1508.04999v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-deep-bag-of-features-model-for-music-auto","repo_url":"https://github.com/juhannam/deepbof","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"audio-classification","task_name":"Audio Classification"},{"task_slug":"information-retrieval","task_name":"Information Retrieval"},{"task_slug":"music-auto-tagging","task_name":"Music Auto-Tagging"},{"task_slug":"music-information-retrieval","task_name":"Music Information Retrieval"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"tag","task_name":"TAG"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}