{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/one-deep-music-representation-to-rule-them","title":"One Deep Music Representation to Rule Them All? : A comparative analysis of different representation learning strategies","arxiv_id":"1802.04051","date":"2018-02-12","proceeding":null,"authors":["Jaehun Kim","Julián Urbano","Cynthia C. S. Liem","Alan Hanjalic"],"abstract":"Inspired by the success of deploying deep learning in the fields of Computer\nVision and Natural Language Processing, this learning paradigm has also found\nits way into the field of Music Information Retrieval. In order to benefit from\ndeep learning in an effective, but also efficient manner, deep transfer\nlearning has become a common approach. In this approach, it is possible to\nreuse the output of a pre-trained neural network as the basis for a new\nlearning task. The underlying hypothesis is that if the initial and new\nlearning tasks show commonalities and are applied to the same type of input\ndata (e.g. music audio), the generated deep representation of the data is also\ninformative for the new task. Since, however, most of the networks used to\ngenerate deep representations are trained using a single initial learning\nsource, their representation is unlikely to be informative for all possible\nfuture tasks. In this paper, we present the results of our investigation of\nwhat are the most important factors to generate deep representations for the\ndata and learning tasks in the music domain. We conducted this investigation\nvia an extensive empirical study that involves multiple learning sources, as\nwell as multiple deep learning architectures with varying levels of information\nsharing between sources, in order to learn music representations. We then\nvalidate these representations considering multiple target datasets for\nevaluation. The results of our experiments yield several insights on how to\napproach the design of methods for learning widely deployable deep data\nrepresentations in the music domain.","url_abs":"http://arxiv.org/abs/1802.04051v4","url_pdf":"http://arxiv.org/pdf/1802.04051v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"one-deep-music-representation-to-rule-them","repo_url":"https://github.com/eldrin/MTLMusicRepresentation-PyTorch","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"all","task_name":"All"},{"task_slug":"deep-learning","task_name":"Deep Learning"},{"task_slug":"information-retrieval","task_name":"Information Retrieval"},{"task_slug":"music-information-retrieval","task_name":"Music Information Retrieval"},{"task_slug":"representation-learning","task_name":"Representation Learning"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"transfer-learning","task_name":"Transfer Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}