{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/tm2d-bimodality-driven-3d-dance-generation","title":"TM2D: Bimodality Driven 3D Dance Generation via Music-Text Integration","arxiv_id":"2304.02419","date":"2023-04-05","proceeding":"ICCV 2023 1","authors":["Kehong Gong","Dongze Lian","Heng Chang","Chuan Guo","Zihang Jiang","Xinxin Zuo","Michael Bi Mi","Xinchao Wang"],"abstract":"We propose a novel task for generating 3D dance movements that simultaneously incorporate both text and music modalities. Unlike existing works that generate dance movements using a single modality such as music, our goal is to produce richer dance movements guided by the instructive information provided by the text. However, the lack of paired motion data with both music and text modalities limits the ability to generate dance movements that integrate both. To alleviate this challenge, we propose to utilize a 3D human motion VQ-VAE to project the motions of the two datasets into a latent space consisting of quantized vectors, which effectively mix the motion tokens from the two datasets with different distributions for training. Additionally, we propose a cross-modal transformer to integrate text instructions into motion generation architecture for generating 3D dance movements without degrading the performance of music-conditioned dance generation. To better evaluate the quality of the generated motion, we introduce two novel metrics, namely Motion Prediction Distance (MPD) and Freezing Score (FS), to measure the coherence and freezing percentage of the generated motion. Extensive experiments show that our approach can generate realistic and coherent dance movements conditioned on both text and music while maintaining comparable performance with the two single modalities. Code is available at https://garfield-kh.github.io/TM2D/.","url_abs":"https://arxiv.org/abs/2304.02419v2","url_pdf":"https://arxiv.org/pdf/2304.02419v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"tm2d-bimodality-driven-3d-dance-generation","repo_url":"https://github.com/Garfield-kh/TM2D","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"motion-generation","task_name":"Motion Generation"},{"task_slug":"motion-synthesis","task_name":"Motion Synthesis"},{"task_slug":"motion-prediction","task_name":"motion prediction"}],"methods":[{"method_slug":"vq-vae","method_name":"VQ-VAE"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/motion-synthesis-on-aist","task":"Motion Synthesis","dataset":"AIST++","model":"TM2D","rank_in_archive_order":3,"of":12,"metrics":{"Beat alignment score":"0.2049","FID":"19.01"},"uses_additional_data":false},{"leaderboard":"/sota/motion-synthesis-on-aist","task":"Motion Synthesis","dataset":"AIST++","model":"TM2D (only motion data)","rank_in_archive_order":4,"of":12,"metrics":{"Beat alignment score":"0.2127","FID":"23.94"},"uses_additional_data":false},{"leaderboard":"/sota/motion-synthesis-on-humanml3d","task":"Motion Synthesis","dataset":"HumanML3D","model":"TM2D (t2m)","rank_in_archive_order":32,"of":37,"metrics":{"Diversity":"9.513","FID":"1.021","Multimodality":"4.139"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2304.02419","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2304.02419"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Garfield-kh/TM2D","reach":null}],"summary":{"ran":1},"by_repo_kind":{"official":{"samples":1,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":1,"samples":[{"code_sha256_prefix":"5fa64dfd4c3d7e22","entry":"TransformerX3","repo":"Garfield-kh/TM2D","repo_kind":"official","path":"tm2d_60fps/networks/transformer_x_lf3.py","file_url":"https://github.com/Garfield-kh/TM2D/blob/HEAD/tm2d_60fps/networks/transformer_x_lf3.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"5fa64dfd4c3d7e22"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}