{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/mmm-generative-masked-motion-model","title":"MMM: Generative Masked Motion Model","arxiv_id":"2312.03596","date":"2023-12-06","proceeding":"CVPR 2024 1","authors":["Ekkasit Pinyoanuntapong","Pu Wang","Minwoo Lee","Chen Chen"],"abstract":"Recent advances in text-to-motion generation using diffusion and autoregressive models have shown promising results. However, these models often suffer from a trade-off between real-time performance, high fidelity, and motion editability. To address this gap, we introduce MMM, a novel yet simple motion generation paradigm based on Masked Motion Model. MMM consists of two key components: (1) a motion tokenizer that transforms 3D human motion into a sequence of discrete tokens in latent space, and (2) a conditional masked motion transformer that learns to predict randomly masked motion tokens, conditioned on the pre-computed text tokens. By attending to motion and text tokens in all directions, MMM explicitly captures inherent dependency among motion tokens and semantic mapping between motion and text tokens. During inference, this allows parallel and iterative decoding of multiple motion tokens that are highly consistent with fine-grained text descriptions, therefore simultaneously achieving high-fidelity and high-speed motion generation. In addition, MMM has innate motion editability. By simply placing mask tokens in the place that needs editing, MMM automatically fills the gaps while guaranteeing smooth transitions between editing and non-editing parts. Extensive experiments on the HumanML3D and KIT-ML datasets demonstrate that MMM surpasses current leading methods in generating high-quality motion (evidenced by superior FID scores of 0.08 and 0.429), while offering advanced editing features such as body-part modification, motion in-betweening, and the synthesis of long motion sequences. In addition, MMM is two orders of magnitude faster on a single mid-range GPU than editable motion diffusion models. Our project page is available at \\url{https://exitudio.github.io/MMM-page}.","url_abs":"https://arxiv.org/abs/2312.03596v2","url_pdf":"https://arxiv.org/pdf/2312.03596v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"mmm-generative-masked-motion-model","repo_url":"https://github.com/exitudio/MMM","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":null,"task_name":"GPU"},{"task_slug":"motion-generation","task_name":"Motion Generation"},{"task_slug":"motion-synthesis","task_name":"Motion Synthesis"},{"task_slug":"model","task_name":"model"},{"task_slug":"motion-in-betweening","task_name":"motion in-betweening"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/motion-synthesis-on-humanml3d","task":"Motion Synthesis","dataset":"HumanML3D","model":"MMM (predict length)","rank_in_archive_order":13,"of":37,"metrics":{"Diversity":"9.411","FID":"0.080","Multimodality":"1.164","R Precision Top3":"0.794"},"uses_additional_data":false},{"leaderboard":"/sota/motion-synthesis-on-humanml3d","task":"Motion Synthesis","dataset":"HumanML3D","model":"MMM (gt length)","rank_in_archive_order":14,"of":37,"metrics":{"Diversity":"9.577","FID":"0.089","Multimodality":"1.226","R Precision Top3":"0.804"},"uses_additional_data":false},{"leaderboard":"/sota/motion-synthesis-on-kit-motion-language","task":"Motion Synthesis","dataset":"KIT Motion-Language","model":"MMM (gt length)","rank_in_archive_order":14,"of":31,"metrics":{"Diversity":"10.910","FID":"0.316","Multimodality":"1.232","R Precision Top3":"0.744"},"uses_additional_data":false},{"leaderboard":"/sota/motion-synthesis-on-kit-motion-language","task":"Motion Synthesis","dataset":"KIT Motion-Language","model":"MMM (predict length)","rank_in_archive_order":17,"of":31,"metrics":{"Diversity":"10.633","FID":"0.429","Multimodality":"1.105","R Precision Top3":"0.718"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2312.03596","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2312.03596"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/exitudio/MMM","reach":null}],"summary":{"ran_fixture":1},"by_repo_kind":{"official":{"samples":1,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":1,"samples":[{"code_sha256_prefix":"2697201a279492a9","entry":"get_acc","repo":"exitudio/MMM","repo_kind":"official","path":"train_t2m_trans.py","file_url":"https://github.com/exitudio/MMM/blob/HEAD/train_t2m_trans.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"2697201a279492a9"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}