{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/dance-to-music-generation-with-encoder-based","title":"Dance-to-Music Generation with Encoder-based Textual Inversion","arxiv_id":"2401.17800","date":"2024-01-31","proceeding":null,"authors":["Sifei Li","Weiming Dong","Yuxin Zhang","Fan Tang","Chongyang Ma","Oliver Deussen","Tong-Yee Lee","Changsheng Xu"],"abstract":"The seamless integration of music with dance movements is essential for communicating the artistic intent of a dance piece. This alignment also significantly improves the immersive quality of gaming experiences and animation productions. Although there has been remarkable advancement in creating high-fidelity music from textual descriptions, current methodologies mainly focus on modulating overall characteristics such as genre and emotional tone. They often overlook the nuanced management of temporal rhythm, which is indispensable in crafting music for dance, since it intricately aligns the musical beats with the dancers' movements. Recognizing this gap, we propose an encoder-based textual inversion technique to augment text-to-music models with visual control, facilitating personalized music generation. Specifically, we develop dual-path rhythm-genre inversion to effectively integrate the rhythm and genre of a dance motion sequence into the textual space of a text-to-music model. Contrary to traditional textual inversion methods, which directly update text embeddings to reconstruct a single target object, our approach utilizes separate rhythm and genre encoders to obtain text embeddings for two pseudo-words, adapting to the varying rhythms and genres. We collect a new dataset called In-the-wild Dance Videos (InDV) and demonstrate that our approach outperforms state-of-the-art methods across multiple evaluation metrics. Furthermore, our method is able to adapt to changes in tempo and effectively integrates with the inherent text-guided generation capability of the pre-trained model. Our source code and demo videos are available at \\url{https://github.com/lsfhuihuiff/Dance-to-music_Siggraph_Asia_2024}","url_abs":"https://arxiv.org/abs/2401.17800v2","url_pdf":"https://arxiv.org/pdf/2401.17800v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"links_only","authors_date_abstract":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license), from the Kaggle arXiv metadata snapshot of 2026-09-12"},"code_links":[{"paper_slug":"dance-to-music-generation-with-encoder-based","repo_url":"https://github.com/lsfhuihuiff/dance-to-music_siggraph_asia_2024","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2401.17800","atlas_url":"https://app.syntology.ai/?focus=2401.17800","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2401.17800"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/lsfhuihuiff/dance-to-music_siggraph_asia_2024","reach":{"status":"ok"}}],"summary":{"ran":4},"by_repo_kind":{"official":{"samples":4,"ran":4,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":4,"samples":[{"code_sha256_prefix":"b94d3dcb2a918be3","entry":"get_adv_criterion","repo":"lsfhuihuiff/dance-to-music_siggraph_asia_2024","repo_kind":"official","path":"audiocraft/adversarial/losses.py","file_url":"https://github.com/lsfhuihuiff/dance-to-music_siggraph_asia_2024/blob/HEAD/audiocraft/adversarial/losses.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"b94d3dcb2a918be3"}},{"code_sha256_prefix":"88b0022a7cb492a1","entry":"get_fake_criterion","repo":"lsfhuihuiff/dance-to-music_siggraph_asia_2024","repo_kind":"official","path":"audiocraft/adversarial/losses.py","file_url":"https://github.com/lsfhuihuiff/dance-to-music_siggraph_asia_2024/blob/HEAD/audiocraft/adversarial/losses.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"88b0022a7cb492a1"}},{"code_sha256_prefix":"1e08a3af21dad9b1","entry":"get_init_fn","repo":"lsfhuihuiff/dance-to-music_siggraph_asia_2024","repo_kind":"official","path":"audiocraft/models/lm.py","file_url":"https://github.com/lsfhuihuiff/dance-to-music_siggraph_asia_2024/blob/HEAD/audiocraft/models/lm.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"1e08a3af21dad9b1"}},{"code_sha256_prefix":"31e35a10d885d5a5","entry":"get_real_criterion","repo":"lsfhuihuiff/dance-to-music_siggraph_asia_2024","repo_kind":"official","path":"audiocraft/adversarial/losses.py","file_url":"https://github.com/lsfhuihuiff/dance-to-music_siggraph_asia_2024/blob/HEAD/audiocraft/adversarial/losses.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"31e35a10d885d5a5"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}