Papers › Act As You Wish: Fine-Grained Control of Motion Diffusion Model with Hierarchical...
Act As You Wish: Fine-Grained Control of Motion Diffusion Model with Hierarchical Semantic Graphs
Most text-driven human motion generation methods employ sequential modeling approaches, e.g., transformer, to extract sentence-level text representations automatically and implicitly for human motion synthesis. However, these compact text representations may overemphasize the action names at the expense of other important properties and lack fine-grained details to guide the synthesis of subtly distinct motion. In this paper, we propose hierarchical semantic graphs for fine-grained control over motion generation. Specifically, we disentangle motion descriptions into hierarchical semantic graphs including three levels of motions, actions, and specifics. Such global-to-local structures facilitate a comprehensive understanding of motion description and fine-grained control of motion generation. Correspondingly, to leverage the coarse-to-fine topology of hierarchical semantic graphs, we decompose the text-to-motion diffusion process into three semantic levels, which correspond to capturing the overall motion, local actions, and action specifics. Extensive experiments on two benchmark human motion datasets, including HumanML3D and KIT, with superior performances, justify the efficacy of our method. More encouragingly, by modifying the edge weights of hierarchical semantic graphs, our method can continuously refine the generated motion, which may have a far-reaching impact on the community. Code and pre-training weights are available at https://github.com/jpthu17/GraphMotion.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Motion Synthesis | HumanML3D | GraphMotion | Diversity | 9.692 | #20 of 37 | Archive leaderboard | report |
| Motion Synthesis | HumanML3D | GraphMotion | FID | 0.116 | #20 of 37 | Archive leaderboard | report |
| Motion Synthesis | HumanML3D | GraphMotion | Multimodality | 2.766 | #20 of 37 | Archive leaderboard | report |
| Motion Synthesis | HumanML3D | GraphMotion | R Precision Top3 | 0.785 | #20 of 37 | Archive leaderboard | report |
| Motion Synthesis | KIT Motion-Language | GraphMotion | Diversity | 11.12 | #13 of 31 | Archive leaderboard | report |
| Motion Synthesis | KIT Motion-Language | GraphMotion | FID | 0.313 | #13 of 31 | Archive leaderboard | report |
| Motion Synthesis | KIT Motion-Language | GraphMotion | Multimodality | 3.627 | #13 of 31 | Archive leaderboard | report |
| Motion Synthesis | KIT Motion-Language | GraphMotion | R Precision Top3 | 0.769 | #13 of 31 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections