{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/executing-your-commands-via-motion-diffusion","title":"Executing your Commands via Motion Diffusion in Latent Space","arxiv_id":"2212.04048","date":"2022-12-08","proceeding":"CVPR 2023 1","authors":["Xin Chen","Biao Jiang","Wen Liu","Zilong Huang","Bin Fu","Tao Chen","Jingyi Yu","Gang Yu"],"abstract":"We study a challenging task, conditional human motion generation, which produces plausible human motion sequences according to various conditional inputs, such as action classes or textual descriptors. Since human motions are highly diverse and have a property of quite different distribution from conditional modalities, such as textual descriptors in natural languages, it is hard to learn a probabilistic mapping from the desired conditional modality to the human motion sequences. Besides, the raw motion data from the motion capture system might be redundant in sequences and contain noises; directly modeling the joint distribution over the raw motion sequences and conditional modalities would need a heavy computational overhead and might result in artifacts introduced by the captured noises. To learn a better representation of the various human motion sequences, we first design a powerful Variational AutoEncoder (VAE) and arrive at a representative and low-dimensional latent code for a human motion sequence. Then, instead of using a diffusion model to establish the connections between the raw motion sequences and the conditional inputs, we perform a diffusion process on the motion latent space. Our proposed Motion Latent-based Diffusion model (MLD) could produce vivid motion sequences conforming to the given conditional inputs and substantially reduce the computational overhead in both the training and inference stages. Extensive experiments on various human motion generation tasks demonstrate that our MLD achieves significant improvements over the state-of-the-art methods among extensive human motion generation tasks, with two orders of magnitude faster than previous diffusion models on raw motion sequences.","url_abs":"https://arxiv.org/abs/2212.04048v3","url_pdf":"https://arxiv.org/pdf/2212.04048v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"executing-your-commands-via-motion-diffusion","repo_url":"https://github.com/chenfengye/motion-latent-diffusion","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"motion-generation","task_name":"Motion Generation"},{"task_slug":"motion-synthesis","task_name":"Motion Synthesis"}],"methods":[{"method_slug":"diffusion","method_name":"Diffusion"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/motion-synthesis-on-humanact12","task":"Motion Synthesis","dataset":"HumanAct12","model":"MLD","rank_in_archive_order":2,"of":2,"metrics":{"Accuracy":"0.964","FID":"0.077","Multimodality":"2.824"},"uses_additional_data":false},{"leaderboard":"/sota/motion-synthesis-on-humanml3d","task":"Motion Synthesis","dataset":"HumanML3D","model":"MLD","rank_in_archive_order":28,"of":37,"metrics":{"Diversity":"9.724","FID":"0.473","Multimodality":"2.413","R Precision Top3":"0.772"},"uses_additional_data":false},{"leaderboard":"/sota/motion-synthesis-on-kit-motion-language","task":"Motion Synthesis","dataset":"KIT Motion-Language","model":"MLD","rank_in_archive_order":16,"of":31,"metrics":{"Diversity":"10.80","FID":"0.404","Multimodality":"2.192","R Precision Top3":"0.734"},"uses_additional_data":false},{"leaderboard":"/sota/motion-synthesis-on-kit-motion-language","task":"Motion Synthesis","dataset":"KIT Motion-Language","model":"TEMOS","rank_in_archive_order":29,"of":31,"metrics":{"Diversity":"10.84","FID":"3.717","Multimodality":"0.532","R Precision Top3":"0.687"},"uses_additional_data":false},{"leaderboard":"/sota/motion-synthesis-on-motion-x","task":"Motion Synthesis","dataset":"Motion-X","model":"MLD","rank_in_archive_order":3,"of":4,"metrics":{"Diversity":"10.420","FID":"3.407","MModality":"2.448","TMR-Matching Score":"0.883","TMR-R-Precision Top3":"0.683"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2212.04048","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2212.04048"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/chenfengye/motion-latent-diffusion","reach":null}],"summary":{"ran_draft_wrong":1},"by_repo_kind":{"official":{"samples":1,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"1782ab820abb4184","entry":"extend_paths","repo":"chenfengye/motion-latent-diffusion","repo_kind":"official","path":"render.py","file_url":"https://github.com/chenfengye/motion-latent-diffusion/blob/HEAD/render.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"1782ab820abb4184"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}