Papers › Generating Diverse and Natural 3D Human Motions From Text

Generating Diverse and Natural 3D Human Motions From Text

1 Jan 2022CVPR 2022 1archive 2025-07-28

Chuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang, Wei Ji, Xingyu Li, Li Cheng

Automated generation of 3D human motions from text is a challenging problem. The generated motions are expected to be sufficiently diverse to explore the text-grounded motion space, and more importantly, accurately depicting the content in prescribed text descriptions. Here we tackle this problem with a two-stage approach: text2length sampling and text2motion generation. Text2length involves sampling from the learned distribution function of motion lengths conditioned on the input text. This is followed by our text2motion module using temporal variational autoencoder to synthesize a diverse set of human motions of the sampled lengths. Instead of directly engaging with pose sequences, we propose motion snippet code as our internal motion representation, which captures local semantic motion contexts and is empirically shown to facilitate the generation of plausible motions faithful to the input text. Moreover, a large-scale dataset of scripted 3D Human motions, HumanML3D, is constructed, consisting of 14,616 motion clips and 44,970 text descriptions. Extensive empirical experiments demonstrate the effectiveness of our approach. Project webpage: https://ericguo5513.github.io/text-to-motion/.

PaperPDFCode

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Motion Synthesis

Datasets

Introduced by this paper, per the archive.

HumanML3D

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Motion Synthesis HumanML3D T2M Diversity 9.175 #33 of 37 Archive leaderboard report
Motion Synthesis HumanML3D T2M FID 1.087 #33 of 37 Archive leaderboard report
Motion Synthesis HumanML3D T2M Multimodality 2.219 #33 of 37 Archive leaderboard report
Motion Synthesis HumanML3D T2M R Precision Top3 0.736 #33 of 37 Archive leaderboard report
Motion Synthesis Inter-X T2M FID 5.481 #3 of 6 Archive leaderboard report
Motion Synthesis Inter-X T2M MMDist 9.576 #3 of 6 Archive leaderboard report
Motion Synthesis Inter-X T2M MModality 2.761 #3 of 6 Archive leaderboard report
Motion Synthesis Inter-X T2M R-Precision Top3 0.396 #3 of 6 Archive leaderboard report
Motion Synthesis InterHuman T2M FID 13.769 #9 of 10 Archive leaderboard report
Motion Synthesis InterHuman T2M MMDist 5.731 #9 of 10 Archive leaderboard report
Motion Synthesis InterHuman T2M MModality 1.387 #9 of 10 Archive leaderboard report
Motion Synthesis InterHuman T2M R-Precision Top3 0.464 #9 of 10 Archive leaderboard report
Motion Synthesis KIT Motion-Language T2M Diversity 10.72 #27 of 31 Archive leaderboard report
Motion Synthesis KIT Motion-Language T2M FID 3.022 #27 of 31 Archive leaderboard report
Motion Synthesis KIT Motion-Language T2M Multimodality 2.052 #27 of 31 Archive leaderboard report
Motion Synthesis KIT Motion-Language T2M R Precision Top3 0.681 #27 of 31 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections