{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/investigating-personalization-methods-in-text","title":"Investigating Personalization Methods in Text to Music Generation","arxiv_id":"2309.11140","date":"2023-09-20","proceeding":null,"authors":["Manos Plitsis","Theodoros Kouzelis","Georgios Paraskevopoulos","Vassilis Katsouros","Yannis Panagakis"],"abstract":"In this work, we investigate the personalization of text-to-music diffusion models in a few-shot setting. Motivated by recent advances in the computer vision domain, we are the first to explore the combination of pre-trained text-to-audio diffusers with two established personalization methods. We experiment with the effect of audio-specific data augmentation on the overall system performance and assess different training strategies. For evaluation, we construct a novel dataset with prompts and music clips. We consider both embedding-based and music-specific metrics for quantitative evaluation, as well as a user study for qualitative evaluation. Our analysis shows that similarity metrics are in accordance with user preferences and that current personalization approaches tend to learn rhythmic music constructs more easily than melody. The code, dataset, and example material of this study are open to the research community.","url_abs":"https://arxiv.org/abs/2309.11140v1","url_pdf":"https://arxiv.org/pdf/2309.11140v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"investigating-personalization-methods-in-text","repo_url":"https://github.com/zelaki/DreamSound","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"data-augmentation","task_name":"Data Augmentation"},{"task_slug":"music-generation","task_name":"Music Generation"},{"task_slug":"text-to-music-generation","task_name":"Text-to-Music Generation"}],"methods":[{"method_slug":"diffusion","method_name":"Diffusion"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2309.11140","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2309.11140"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/zelaki/DreamSound","reach":{"status":"ok"}}],"summary":{"ran":2,"unverified":1},"by_repo_kind":{"official":{"samples":3,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"cc28ab1bf6563032","entry":"create_mixture","repo":"zelaki/DreamSound","repo_kind":"official","path":"dreambooth_audioldm.py","file_url":"https://github.com/zelaki/DreamSound/blob/HEAD/dreambooth_audioldm.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"cc28ab1bf6563032"}},{"code_sha256_prefix":"d5a09007aef6b7ff","entry":"prepare_inputs_for_generation","repo":"zelaki/DreamSound","repo_kind":"official","path":"pipeline/pipeline_audioldm2.py","file_url":"https://github.com/zelaki/DreamSound/blob/HEAD/pipeline/pipeline_audioldm2.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"d5a09007aef6b7ff"}},{"code_sha256_prefix":"cec37d613e438bcc","entry":"add_special_tokens","repo":"zelaki/DreamSound","repo_kind":"official","path":"pipeline/modeling_audioldm2.py","file_url":"https://github.com/zelaki/DreamSound/blob/HEAD/pipeline/modeling_audioldm2.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"cec37d613e438bcc"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}