Papers › Fg-T2M: Fine-Grained Text-Driven Human Motion Generation via Diffusion Model
Fg-T2M: Fine-Grained Text-Driven Human Motion Generation via Diffusion Model
Yin Wang, Zhiying Leng, Frederick W. B. Li, Shun-Cheng Wu, Xiaohui Liang
Text-driven human motion generation in computer vision is both significant and challenging. However, current methods are limited to producing either deterministic or imprecise motion sequences, failing to effectively control the temporal and spatial relationships required to conform to a given text description. In this work, we propose a fine-grained method for generating high-quality, conditional human motion sequences supporting precise text description. Our approach consists of two key components: 1) a linguistics-structure assisted module that constructs accurate and complete language feature to fully utilize text information; and 2) a context-aware progressive reasoning module that learns neighborhood and overall semantic linguistics features from shallow and deep graph neural networks to achieve a multi-step inference. Experiments show that our approach outperforms text-driven motion generation methods on HumanML3D and KIT test sets and generates better visually confirmed motion to the text conditions.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Motion Synthesis | HumanML3D | Fg-T2M | Diversity | 9.278 | #25 of 37 | Archive leaderboard | report |
| Motion Synthesis | HumanML3D | Fg-T2M | FID | 0.243 | #25 of 37 | Archive leaderboard | report |
| Motion Synthesis | HumanML3D | Fg-T2M | Multimodality | 1.614 | #25 of 37 | Archive leaderboard | report |
| Motion Synthesis | HumanML3D | Fg-T2M | R Precision Top3 | 0.783 | #25 of 37 | Archive leaderboard | report |
| Motion Synthesis | KIT Motion-Language | Fg-T2M | Diversity | 10.93 | #22 of 31 | Archive leaderboard | report |
| Motion Synthesis | KIT Motion-Language | Fg-T2M | FID | 0.571 | #22 of 31 | Archive leaderboard | report |
| Motion Synthesis | KIT Motion-Language | Fg-T2M | Multimodality | 1.019 | #22 of 31 | Archive leaderboard | report |
| Motion Synthesis | KIT Motion-Language | Fg-T2M | R Precision Top3 | 0.745 | #22 of 31 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections